Beyond Correctness: Validity-Oriented Evaluation of Biomedical LLM Judges

📅 2026-08-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出一种可扩展的有效性评估方法,用于评估生物医学大模型裁判,并通过四种训练方式测试其性能,发现SFT→RL组合在正确性、合规性和鲁棒性上表现最佳。
📝 Abstract
We propose a scalable, validity-oriented pipeline for evaluating biomedical LLM judges when high-quality human judgments are scarce. First, we augment existing human-labelled biomedical benchmarks with deterministic, metric-grounded mutations that produce auditable preference pairs. Second, we evaluate judges beyond aggregate correctness using three deployment-relevant dimensions: correctness against metric-derived gold labels, robustness under repeated stochastic sampling, and compliance with the requested output format. We use this pipeline to assess Llama-3.1-8B-Instruct under four regimes: (1) base, using the instruct model as is; (2) SFT, distillation-based supervised fine-tuning only; (3) RL, GRPO-based reinforcement learning only; and (4) SFT$\rightarrow$RL, SFT followed by RL. The base and single-stage regimes struggle on structured medical discrimination such as PICO extraction and clinical calculations, whereas SFT$\rightarrow$RL performs best across correctness, compliance, and robustness; gains concentrate on decomposable tasks (PICO, MedCalc), at times matching or outperforming frontier models.
Problem

Research questions and friction points this paper is trying to address.

biomedical LLM judges
validity-oriented evaluation
high-quality human judgments
Innovation

Methods, ideas, or system contributions that make the work stand out.

validity-oriented evaluation
metric-grounded mutations
robustness under repeated stochastic sampling
SFT→RL
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Rodrigo de Oliveira
Independent Researcher, London, UK
F
Federico Pittino
IQVIA, Barcelona, Spain
J
James Gwinnutt
IQVIA, Reading, UK
J
Jay Nanavati
IQVIA, Cambridge, UK