What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks
该研究通过适应收敛和区分效度的方法,评估了56个AI基准测试是否准确衡量其声称的概念,如推理、拒绝等,并发现许多基准测试在概念一致性上存在问题。
该研究通过适应收敛和区分效度的方法,评估了56个AI基准测试是否准确衡量其声称的概念,如推理、拒绝等,并发现许多基准测试在概念一致性上存在问题。
This work proposes a collaborative framework in which large language models (LLMs) serve as “empathy editors” to refine clinician-authored messages, enhancing empathetic expression while preserving medical accuracy. Recognizing that clinicians under pressure often struggle to balance factual precision with emotional attunement—leading to perceived empathy deficits in patient communication—the approach emphasizes human–AI collaboration rather than end-to-end generation. To evaluate performance, the study introduces two novel metrics: the Empathy Ranking Score, which assesses perceived empathy, and the MedFactChecking Score, which verifies clinical fidelity. Experimental results demonstrate that LLM-edited texts significantly improve patients’ perception of empathy without compromising medical correctness, outperforming fully LLM-generated responses.
Existing AI-based clinical note evaluation methods suffer from misalignment between automated metrics and physician preferences, alongside the subjectivity and scalability limitations of expert reviews. Method: We propose Feedback2Checklist—a novel framework that distills large-scale, de-identified real-world clinical feedback into structured, interpretable, and actionable evaluation checklists, and builds an LLM-powered automated evaluator. Contribution/Results: Our approach significantly improves agreement between automated scores and physician preferences (+32.7% Spearman correlation), demonstrating high coverage, diversity, and robustness to quality degradation. Offline experiments show superior performance in identifying low-quality clinical notes compared to baselines, while ensuring strong clinical alignment and practical deployability.
Large language models (LLMs) struggle to balance information value against acquisition cost in data-scarce scenarios, leading to suboptimal decision-making. Method: This paper proposes a test-time zero-shot active information acquisition framework. Its core is CuriosiTree—a novel, cost-aware greedy tree search strategy that integrates heuristic search, zero-shot information gain estimation, and fusion of heterogeneous multi-source information—requiring no fine-tuning. Crucially, it formalizes information acquisition as a dynamic sequential decision problem, optimizing query actions in real time within a single inference pass. Contribution/Results: Empirical evaluation on clinical diagnosis simulation demonstrates that CuriosiTree achieves significantly higher diagnostic accuracy at lower total acquisition cost, outperforming baselines including random sampling and naive greedy strategies across all metrics.
This paper investigates the statistical performance of Prediction-Driven Inference (PPI++) under limited-sample regimes, specifically characterizing when its estimation error degrades relative to inference using only ground-truth labels. Through non-asymptotic analysis and a correlation-driven error decomposition, we derive the first exact finite-sample error bound for PPI++. We establish a “no-free-lunch” theorem: PPI++ improves estimation accuracy if and only if the correlation between pseudo-labels and ground-truth labels exceeds $1/sqrt{n-2}$ in the Gaussian setting—a threshold rigorously derived for both binary and Gaussian cases. We further analyze trade-offs between single-sample and sample-splitting variants of PPI++. All theoretical findings are empirically validated on real-world datasets, confirming that the predicted correlation threshold accurately governs performance gains.
该研究通过适应收敛和区分效度的方法,评估了56个AI基准测试是否准确衡量其声称的概念,如推理、拒绝等,并发现许多基准测试在概念一致性上存在问题。
This work proposes a collaborative framework in which large language models (LLMs) serve as “empathy editors” to refine clinician-authored messages, enhancing empathetic expression while preserving medical accuracy. Recognizing that clinicians under pressure often struggle to balance factual precision with emotional attunement—leading to perceived empathy deficits in patient communication—the approach emphasizes human–AI collaboration rather than end-to-end generation. To evaluate performance, the study introduces two novel metrics: the Empathy Ranking Score, which assesses perceived empathy, and the MedFactChecking Score, which verifies clinical fidelity. Experimental results demonstrate that LLM-edited texts significantly improve patients’ perception of empathy without compromising medical correctness, outperforming fully LLM-generated responses.
Existing AI-based clinical note evaluation methods suffer from misalignment between automated metrics and physician preferences, alongside the subjectivity and scalability limitations of expert reviews. Method: We propose Feedback2Checklist—a novel framework that distills large-scale, de-identified real-world clinical feedback into structured, interpretable, and actionable evaluation checklists, and builds an LLM-powered automated evaluator. Contribution/Results: Our approach significantly improves agreement between automated scores and physician preferences (+32.7% Spearman correlation), demonstrating high coverage, diversity, and robustness to quality degradation. Offline experiments show superior performance in identifying low-quality clinical notes compared to baselines, while ensuring strong clinical alignment and practical deployability.
Large language models (LLMs) struggle to balance information value against acquisition cost in data-scarce scenarios, leading to suboptimal decision-making. Method: This paper proposes a test-time zero-shot active information acquisition framework. Its core is CuriosiTree—a novel, cost-aware greedy tree search strategy that integrates heuristic search, zero-shot information gain estimation, and fusion of heterogeneous multi-source information—requiring no fine-tuning. Crucially, it formalizes information acquisition as a dynamic sequential decision problem, optimizing query actions in real time within a single inference pass. Contribution/Results: Empirical evaluation on clinical diagnosis simulation demonstrates that CuriosiTree achieves significantly higher diagnostic accuracy at lower total acquisition cost, outperforming baselines including random sampling and naive greedy strategies across all metrics.
This paper investigates the statistical performance of Prediction-Driven Inference (PPI++) under limited-sample regimes, specifically characterizing when its estimation error degrades relative to inference using only ground-truth labels. Through non-asymptotic analysis and a correlation-driven error decomposition, we derive the first exact finite-sample error bound for PPI++. We establish a “no-free-lunch” theorem: PPI++ improves estimation accuracy if and only if the correlation between pseudo-labels and ground-truth labels exceeds $1/sqrt{n-2}$ in the Gaussian setting—a threshold rigorously derived for both binary and Gaussian cases. We further analyze trade-offs between single-sample and sample-splitting variants of PPI++. All theoretical findings are empirically validated on real-world datasets, confirming that the predicted correlation threshold accurately governs performance gains.