Institution profile

Pindrop

Industry researchnorthamerica · us
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Spotlights and Blindspots: Evaluation Machine-Generated Text Detection

Apr 17, 2026

This study addresses the lack of standardized evaluation protocols in machine-generated text detection, which hinders fair comparison of model performance. The authors systematically evaluate 15 detection methods across diverse datasets comprising both human-written and machine-generated English texts, covering six detector families and seven generative models. Employing a multi-dataset cross-validation framework and multiple evaluation metrics, they find that no single detector consistently outperforms others across all scenarios—most excel only in specific settings—and overall performance degrades significantly on novel, human-authored texts from high-stakes domains. The work highlights the strong dependence of detector efficacy on training and evaluation data as well as metric choice, exposing critical blind spots in current evaluation paradigms and underscoring the decisive role of methodological decisions in shaping empirical conclusions.

0 citationsRead paper

Identifying Bias in Machine-generated Text Detection

Dec 09, 2025

This study systematically evaluates the fairness of 16 mainstream English AI text detectors across four sociodemographic attributes—gender, race/ethnicity, English Language Learner (ELL) status, and socioeconomic status. Method: Using a student writing dataset, we employ regression modeling, subgroup difference testing, human annotation comparison, and a multi-model bias auditing framework. Contribution/Results: We uncover previously undocumented systematic disparities: essays by ELL and non-White students are significantly over-classified as AI-generated, whereas those by economically disadvantaged students are disproportionately misclassified as human-written. In contrast, human annotators—though less accurate overall—exhibit no statistically significant attribute-based bias. These findings reveal severe, asymmetric societal biases embedded in current detectors, indicating an urgent need for algorithmic fairness interventions at the model design and deployment levels.

0 citationsRead paper

Pindrop it! Audio and Visual Deepfake Countermeasures for Robust Detection and Fine Grained-Localization

Aug 11, 2025

To address the degraded robustness and localization accuracy in deepfake video detection caused by fine-grained local manipulations—particularly audio-visual inconsistency—this paper proposes a joint audio-visual multimodal framework for detection and precise spatiotemporal localization. Methodologically, it fuses spectrogram-based audio features with frame-level visual features, incorporates a spatiotemporal attention mechanism to model cross-modal temporal dependencies, and employs contrastive learning to enhance discriminability for subtle forgeries. The framework supports both temporal localization of forged segments and semantic classification of manipulation types, significantly improving sensitivity to minute synthetic artifacts and providing interpretable, pixel- and frame-level evidence. Evaluated on the ACM 1M Deepfakes Detection Challenge, our method ranks first in the temporal localization track and places among the top four in the classification track on the TestA set, demonstrating state-of-the-art performance and practical applicability.

0 citationsRead paper

Open-Set Source Tracing of Audio Deepfake Systems

Jul 08, 2025

Open-set provenance attribution for audio deepfake systems—robustly identifying a large number of previously unseen forgery systems—remains a critical challenge. Method: We propose a softmax energy (SME)-based out-of-distribution detection framework tailored for provenance attribution. Our approach innovatively integrates SME into the attribution task via an SME-guided training mechanism and synergistically combines multiple data augmentation strategies: temperature scaling, copy-synthesis, codec-induced distortion, and reverberation enhancement. Contribution/Results: Evaluated under the Interspeech 2025 benchmark protocol, our method significantly improves open-set generalization, achieving an FPR95 of 8.3%—a 31% relative reduction over the prior state of the art. This establishes a novel paradigm for tracing previously unknown audio deepfake systems in real-world scenarios.

0 citationsRead paper
Recent publications

Latest Papers

Spotlights and Blindspots: Evaluation Machine-Generated Text Detection

Apr 17, 2026

This study addresses the lack of standardized evaluation protocols in machine-generated text detection, which hinders fair comparison of model performance. The authors systematically evaluate 15 detection methods across diverse datasets comprising both human-written and machine-generated English texts, covering six detector families and seven generative models. Employing a multi-dataset cross-validation framework and multiple evaluation metrics, they find that no single detector consistently outperforms others across all scenarios—most excel only in specific settings—and overall performance degrades significantly on novel, human-authored texts from high-stakes domains. The work highlights the strong dependence of detector efficacy on training and evaluation data as well as metric choice, exposing critical blind spots in current evaluation paradigms and underscoring the decisive role of methodological decisions in shaping empirical conclusions.

0 citationsRead paper

Identifying Bias in Machine-generated Text Detection

Dec 09, 2025

This study systematically evaluates the fairness of 16 mainstream English AI text detectors across four sociodemographic attributes—gender, race/ethnicity, English Language Learner (ELL) status, and socioeconomic status. Method: Using a student writing dataset, we employ regression modeling, subgroup difference testing, human annotation comparison, and a multi-model bias auditing framework. Contribution/Results: We uncover previously undocumented systematic disparities: essays by ELL and non-White students are significantly over-classified as AI-generated, whereas those by economically disadvantaged students are disproportionately misclassified as human-written. In contrast, human annotators—though less accurate overall—exhibit no statistically significant attribute-based bias. These findings reveal severe, asymmetric societal biases embedded in current detectors, indicating an urgent need for algorithmic fairness interventions at the model design and deployment levels.

0 citationsRead paper

Pindrop it! Audio and Visual Deepfake Countermeasures for Robust Detection and Fine Grained-Localization

Aug 11, 2025

To address the degraded robustness and localization accuracy in deepfake video detection caused by fine-grained local manipulations—particularly audio-visual inconsistency—this paper proposes a joint audio-visual multimodal framework for detection and precise spatiotemporal localization. Methodologically, it fuses spectrogram-based audio features with frame-level visual features, incorporates a spatiotemporal attention mechanism to model cross-modal temporal dependencies, and employs contrastive learning to enhance discriminability for subtle forgeries. The framework supports both temporal localization of forged segments and semantic classification of manipulation types, significantly improving sensitivity to minute synthetic artifacts and providing interpretable, pixel- and frame-level evidence. Evaluated on the ACM 1M Deepfakes Detection Challenge, our method ranks first in the temporal localization track and places among the top four in the classification track on the TestA set, demonstrating state-of-the-art performance and practical applicability.

0 citationsRead paper

Open-Set Source Tracing of Audio Deepfake Systems

Jul 08, 2025

Open-set provenance attribution for audio deepfake systems—robustly identifying a large number of previously unseen forgery systems—remains a critical challenge. Method: We propose a softmax energy (SME)-based out-of-distribution detection framework tailored for provenance attribution. Our approach innovatively integrates SME into the attribution task via an SME-guided training mechanism and synergistically combines multiple data augmentation strategies: temperature scaling, copy-synthesis, codec-induced distortion, and reverberation enhancement. Contribution/Results: Evaluated under the Interspeech 2025 benchmark protocol, our method significantly improves open-set generalization, achieving an FPR95 of 8.3%—a 31% relative reduction over the prior state of the art. This establishes a novel paradigm for tracing previously unknown audio deepfake systems in real-world scenarios.

0 citationsRead paper