Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation

📅 2026-07-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了自生成文本识别能力对AI安全的影响,通过不同实验设计选择评估多种模型,发现模型性能受评估格式、对话格式和任务领域影响,并提出需在关键AI应用中监控此能力。
📝 Abstract
Self-Generated Text Recognition (SGTR)--the ability of an LLM to identify its own outputs--poses risks to AI safeguards that rely on LLMs as evaluators or monitors. Specifically, an LLM may recognize outputs from other copies of the same model and make biased judgments or collude outright. Prior work has drawn conflicting conclusions about whether current models possess significant SGTR capabilities. We reconcile these findings by identifying key experimental design choices--which we term operationalizations--that drive divergent results. Evaluating 13-21 models across six operationalizations, we find that accuracy varies substantially with evaluation format (pairwise vs. individual assessments of text), conversation structure (presenting candidate text in user tags vs. assistant tags), and the domain of the task used to generate candidate text (e.g., coding vs. summarization). We corroborate previous observations that a quality heuristic--models attributing authorship to text they perceive as higher quality--is a dominant confound. We also find that improving a model's SGTR performance via SFT in one evaluation configuration can generalize to others. Training for SGTR additionally causes models to prefer their own outputs when acting as a judge in the AlpacaEval framework. Finally, we discuss the implications of our evaluations for the safety of future AI systems: our work suggests that, despite confounds, some models possess practical SGTR capabilities, and that training a model for SGTR in one setting can affect its self-recognition and self-preference more generally. We conclude that SGTR should be monitored and considered in the design of safety-critical AI applications.
Problem

Research questions and friction points this paper is trying to address.

Self-Generated Text Recognition
Large Language Models
AI Safeguards
Biased Judgments
Operationalizations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Generated Text Recognition
Operationalizations
Supervised Fine-Tuning
AlpacaEval
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jesse St. Amand
MARS
C
Callum Canavan
Independent
S
Sohaib Imran
MARS
J
Joseph Hewson
Independent
A
Aaron Lutz
Independent
Shi Feng
Shi Feng
The George Washington University
AI safetyalignment
Puria Radmard
Puria Radmard
Univeristy of Cambridge
Lennie Wells
Lennie Wells
University of Cambridge
statisticsmachine learning