Institution profile

Reality Defender

Industry researchnorthamerica · us
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

FoeGlass: Simple In-Context Learning Is Enough for Red Teaming Audio Deepfake Detectors

Jun 03, 2026

This work addresses the limitations of current audio deepfake detection models, whose evaluation relies on manually curated data and thus fails to efficiently uncover critical blind spots. To overcome this, the authors propose FoeGlass—the first black-box, automated red-teaming framework tailored for text-to-speech systems. FoeGlass leverages the in-context learning capability of large language models to explore the input space and generate adversarial audio samples capable of evading detection. By incorporating diversity-guided prompt design to mitigate mode collapse, FoeGlass substantially enhances attack efficacy in black-box settings, reducing detector false negative rates by up to 94% while demonstrating strong cross-model transferability. Furthermore, fine-tuning detectors on FoeGlass-generated data improves their robustness by as much as 41%.

0 citationsRead paper

Post-hoc Selective Classification for Reliable Synthetic Image Detection

May 08, 2026

Existing deepfake image detectors often exhibit insufficient reliability under distribution shifts, leading to frequent misclassifications. To address this issue, this work proposes a training-free, post-hoc selective classification framework that enhances confidence estimation by extending the logit concept to intermediate network layers. The approach aggregates multi-layer features and aligns intermediate-layer representations with class-specific centroids to construct a more robust uncertainty measure. Furthermore, it introduces a preference optimization algorithm based on an upper bound of the Area Under the Risk-Coverage curve (AURC) to enable efficient rejection of unreliable predictions. Evaluated under common covariate shifts, the method substantially improves selective classification performance, achieving up to a 69.55% reduction in AURC.

0 citationsRead paper

PolyJuice Makes It Real: Black-Box, Universal Red Teaming for Synthetic Image Detectors

Sep 18, 2025

Existing synthetic image detectors (SIDs) rely on white-box access and image-level online optimization for red-teaming, rendering them impractical against black-box, state-of-the-art text-to-image (T2I) models and computationally prohibitive. This paper introduces PolyJuice—the first black-box, universal red-teaming attack against SIDs. PolyJuice identifies transferable distribution shift directions in the latent space of T2I models via black-box queries, enabling image-agnostic universal adversarial perturbations. It supports low-resolution direction estimation and high-resolution transfer, incorporating direction interpolation and enhanced data fine-tuning. Experiments demonstrate that PolyJuice achieves up to 84% evasion success against mainstream SIDs under black-box settings. Furthermore, fine-tuning SIDs with PolyJuice-generated adversarial data improves their robustness by up to 30%. These results significantly advance the practicality and realism of synthetic image detection research, shifting the adversarial paradigm toward deployable, real-world evaluation.

0 citationsRead paper

X-Edit: Detecting and Localizing Edits in Images Altered by Text-Guided Diffusion Models

May 16, 2025

This work addresses the detection of subtle, text-guided image manipulations generated by diffusion models—a critical challenge in AI-generated content forensics. We propose the first localization-aware deepfake detection method capable of precisely identifying tampered regions. Our approach leverages pre-trained diffusion models for image inversion to extract editing-sensitive features, and introduces an attention-driven semantic segmentation network integrating multi-scale reconstruction with a low-frequency-prior supervision strategy. A novel dual-objective loss function jointly optimizes pixel-level segmentation accuracy and low-frequency perceptual correlation. Key contributions include: (1) the first localized detection framework specifically designed for diffusion-based edits; (2) the first benchmark dataset comprising paired original–edited images; and (3) state-of-the-art performance—significantly surpassing existing baselines in PSNR and SSIM—while achieving high-precision pixel-level localization of fine-grained edits, demonstrating strong efficacy and robustness for forensic analysis of AI-generated imagery.

0 citationsRead paper
Recent publications

Latest Papers

FoeGlass: Simple In-Context Learning Is Enough for Red Teaming Audio Deepfake Detectors

Jun 03, 2026

This work addresses the limitations of current audio deepfake detection models, whose evaluation relies on manually curated data and thus fails to efficiently uncover critical blind spots. To overcome this, the authors propose FoeGlass—the first black-box, automated red-teaming framework tailored for text-to-speech systems. FoeGlass leverages the in-context learning capability of large language models to explore the input space and generate adversarial audio samples capable of evading detection. By incorporating diversity-guided prompt design to mitigate mode collapse, FoeGlass substantially enhances attack efficacy in black-box settings, reducing detector false negative rates by up to 94% while demonstrating strong cross-model transferability. Furthermore, fine-tuning detectors on FoeGlass-generated data improves their robustness by as much as 41%.

0 citationsRead paper

Post-hoc Selective Classification for Reliable Synthetic Image Detection

May 08, 2026

Existing deepfake image detectors often exhibit insufficient reliability under distribution shifts, leading to frequent misclassifications. To address this issue, this work proposes a training-free, post-hoc selective classification framework that enhances confidence estimation by extending the logit concept to intermediate network layers. The approach aggregates multi-layer features and aligns intermediate-layer representations with class-specific centroids to construct a more robust uncertainty measure. Furthermore, it introduces a preference optimization algorithm based on an upper bound of the Area Under the Risk-Coverage curve (AURC) to enable efficient rejection of unreliable predictions. Evaluated under common covariate shifts, the method substantially improves selective classification performance, achieving up to a 69.55% reduction in AURC.

0 citationsRead paper

PolyJuice Makes It Real: Black-Box, Universal Red Teaming for Synthetic Image Detectors

Sep 18, 2025

Existing synthetic image detectors (SIDs) rely on white-box access and image-level online optimization for red-teaming, rendering them impractical against black-box, state-of-the-art text-to-image (T2I) models and computationally prohibitive. This paper introduces PolyJuice—the first black-box, universal red-teaming attack against SIDs. PolyJuice identifies transferable distribution shift directions in the latent space of T2I models via black-box queries, enabling image-agnostic universal adversarial perturbations. It supports low-resolution direction estimation and high-resolution transfer, incorporating direction interpolation and enhanced data fine-tuning. Experiments demonstrate that PolyJuice achieves up to 84% evasion success against mainstream SIDs under black-box settings. Furthermore, fine-tuning SIDs with PolyJuice-generated adversarial data improves their robustness by up to 30%. These results significantly advance the practicality and realism of synthetic image detection research, shifting the adversarial paradigm toward deployable, real-world evaluation.

0 citationsRead paper

X-Edit: Detecting and Localizing Edits in Images Altered by Text-Guided Diffusion Models

May 16, 2025

This work addresses the detection of subtle, text-guided image manipulations generated by diffusion models—a critical challenge in AI-generated content forensics. We propose the first localization-aware deepfake detection method capable of precisely identifying tampered regions. Our approach leverages pre-trained diffusion models for image inversion to extract editing-sensitive features, and introduces an attention-driven semantic segmentation network integrating multi-scale reconstruction with a low-frequency-prior supervision strategy. A novel dual-objective loss function jointly optimizes pixel-level segmentation accuracy and low-frequency perceptual correlation. Key contributions include: (1) the first localized detection framework specifically designed for diffusion-based edits; (2) the first benchmark dataset comprising paired original–edited images; and (3) state-of-the-art performance—significantly surpassing existing baselines in PSNR and SSIM—while achieving high-precision pixel-level localization of fine-grained edits, demonstrating strong efficacy and robustness for forensic analysis of AI-generated imagery.

0 citationsRead paper