Advancing Harmful Content Detection in Organizational Research: Integrating Large Language Models with Elo Rating System

📅 2025-06-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
In organizational research, built-in LLM moderation mechanisms often over-censor responses to microaggressions and hate speech—either refusing to answer or diluting harmful content—thereby compromising analytical validity. To address this, we propose an Elo-based dynamic evaluation framework, the first to adapt the Elo rating system for LLM-based harmful content detection. Our method comprises dual-stage annotation validation, task-specific fine-tuning, and real-time response quality calibration—overcoming key limitations of static classification models and conventional prompt engineering. Evaluated on benchmark datasets for microaggressions and hate speech, our approach significantly improves accuracy, precision, and F1-score while substantially reducing false positive rates. It enables high-fidelity, scalable toxicity analysis and real-time assessment of workplace textual data.

Technology Category

Application Category

📝 Abstract
Large language models (LLMs) offer promising opportunities for organizational research. However, their built-in moderation systems can create problems when researchers try to analyze harmful content, often refusing to follow certain instructions or producing overly cautious responses that undermine validity of the results. This is particularly problematic when analyzing organizational conflicts such as microaggressions or hate speech. This paper introduces an Elo rating-based method that significantly improves LLM performance for harmful content analysis In two datasets, one focused on microaggression detection and the other on hate speech, we find that our method outperforms traditional LLM prompting techniques and conventional machine learning models on key measures such as accuracy, precision, and F1 scores. Advantages include better reliability when analyzing harmful content, fewer false positives, and greater scalability for large-scale datasets. This approach supports organizational applications, including detecting workplace harassment, assessing toxic communication, and fostering safer and more inclusive work environments.
Problem

Research questions and friction points this paper is trying to address.

Improving harmful content detection in organizational research using LLMs
Reducing false positives and cautious responses in hate speech analysis
Enhancing accuracy and scalability for microaggression and toxicity detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Integrates LLMs with Elo rating system
Improves harmful content detection accuracy
Reduces false positives in analysis
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.