Stopping and Routing LLM Judge Panels

📅 2026-08-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
论文解决了LLM评估管道中选择和调度评委的问题,通过角色条件分配方法估计评委角色并制定策略,以优化评委组合。
📝 Abstract
LLM evaluation pipelines often have many candidate judges: general LLM-as-a-judge prompts, reward models, safety classifiers, confidence variants, and task-specific verifiers. The deployment question is not only which judge is best, but which judges should be called, on which examples, and when panel construction should stop. We formulate judge-panel design as a role-conditioned allocation problem. From a small labeled audit set, declared slices, and judge costs, the method estimates target-relative roles: copies add no conditional information, complements improve the global panel, and specialists help only on slices. These roles induce a policy: drop copies, add complements globally, route specialists conditionally, and stop when validation gain falls below a threshold. Across reasoning, code, safety, preference, reward-model, summarization, and math audits, the method is compared with single judges, flat panels, matched diversity heuristics, full-call stacking, reliability juries, and frugal cascades. The result is a regime map for judge calls: route specialists on deployable slices, stop in saturated verifier regimes, keep broad ensembles when their risk benefit is worth the cost, and ignore conditional copies. The output is a reusable, auditable call plan for the next evaluation batch.
Problem

Research questions and friction points this paper is trying to address.

LLM evaluation
judge panel
allocation problem
conditional information
validation gain
Innovation

Methods, ideas, or system contributions that make the work stand out.

role-conditioned allocation
target-relative roles
conditional routing
validation gain threshold
judge panel design
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
B
Bin Zhu
School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China
Y
Yi Xie
School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China
Yanghui Rao
Yanghui Rao
Sun Yat-sen University
Text MiningTopic ModelingRepresentation Learning