Competence, Not Accuracy: A Diagnostic for Reference-Free Judge Gates in Skill Optimization

๐Ÿ“… 2026-08-19
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
็ ”็ฉถ้€š่ฟ‡ๅปบ็ซ‹ๆ— ๅ‚่€ƒ่ฏ„ๅˆคๆจกๅž‹๏ผŒ่ฏ„ไผฐ่ฏญ่จ€ๆจกๅž‹ไฝœไธบ้ชŒ่ฏ่€…็š„่ƒฝๅŠ›๏ผŒ่งฃๅ†ณๆŠ€่ƒฝไผ˜ๅŒ–ไธญ่‡ชๅŠจ้ชŒ่ฏ้™ๅˆถ็š„้—ฎ้ข˜ใ€‚
๐Ÿ“ Abstract
Text-space skill optimization adapts a frozen agent by evolving a natural-language skill document, accepting each candidate through a validation gate. Existing gates rely on verifiable rewards, confining these methods to tasks with an automatic verifier. Replacing the verifier with an LLM-judge gate would lift that restriction, but whether such a gate carries usable signal is untested. We ask a prior question: can we tell, before placing a judge in the loop, whether its scores separate correct from incorrect answers at all? We formalize a reference-free judge as a latent solver -- its verdict rests on agreement with whatever it would itself conclude, so its capacity to evaluate is bounded by its capacity to solve. The model yields a closed-form bound on discriminability (ROC-AUC) in the judge's competence $c$ and answer-space size $k$, a necessary condition $c > 1/k$, and the result that the marginal AUC is confounded by item difficulty while a within-question estimator is not. A non-intervening probe records judge scores on genuine optimization runs without altering any decision. We find discriminability at chance where competence sits near the floor and usable above it; that a judge's benchmark accuracy overstates the competence that matters; and, in a closed-loop study, that the screen predicts which kind of gating error occurs. The result is a cheap pre-deployment diagnostic for judge gates.
Problem

Research questions and friction points this paper is trying to address.

text-space skill optimization
validation gate
LLM-judge
discriminability
Innovation

Methods, ideas, or system contributions that make the work stand out.

reference-free judge
latent solver
discriminability bound
non-intervening probe
gating error prediction
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
C
Chenle Chen
University of California, Los Angeles
Y
Yangbo Wei
Shanghai Jiao Tong University
Chao Yao
Chao Yao
Northwestern polytechnical university
S
Shaoqiang Lu
Shanghai Jiao Tong University
J
Junhong Qian
Shanghai Jiao Tong University
Chen Wu
Chen Wu
Institute of Computing Technology, Chinese Academy of Sciences
Information RetrievalNatural Language ProcessingAdversarial Attack
L
Lei He
Eastern Institute of Technology, Ningbo