RadMatch: Auditable Radiology Report Evaluation via Finding-Level Matching

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决AI生成的放射学报告临床质量评估难题,提出RadMatch方法,通过结构化发现匹配和多维度评分来提供可解释和可审计的评估结果。
📝 Abstract
As AI systems are increasingly used to draft radiology reports, reliably evaluating their clinical quality remains a critical challenge. Large language model (LLM)-based metrics are now the best-correlated with radiologist judgment, yet they output a single opaque score that neither a clinician nor a model builder can easily interpret or audit. We introduce RadMatch, a multi-stage, LLM-based metric that decomposes report comparison into a structured finding-level matching with significance-aware scoring and error characterization across seven clinical attribute dimensions (status, location, severity, morphology, certainty, longitudinal comparison, and measurement). The main score is the actionable-error count, both interpretable and auditable. Candidate findings are graded correct, partial, or incorrect, and unmatched findings are counted as missed or hallucinated. Triage and actionable safety recall/precision and per-subset views add complementary, deployment-oriented lenses. Across two expert benchmarks, RadMatch is the most clinically aligned metric, matching inter-radiologist agreement on ReXVal and more than doubling the best prior metric on the harder RadEvalExpert. Relying only on few-shot prompting, it is designed to extend to other modalities and anatomies. We will release RadMatch as open-source code with an interactive dashboard for inspecting results.
Problem

Research questions and friction points this paper is trying to address.

Radiology Report
Clinical Quality
Large Language Model
Finding-Level Matching
Auditable
Innovation

Methods, ideas, or system contributions that make the work stand out.

Finding-Level Matching
Significance-Aware Scoring
Actionable-Error Count
Auditable Metric
Multi-Stage LLM-based
🔎 Similar Papers
No similar papers found.