Beyond Correctness: Validity-Oriented Evaluation of Biomedical LLM Judges
研究提出一种可扩展的有效性评估方法,用于评估生物医学大模型裁判,并通过四种训练方式测试其性能,发现SFT→RL组合在正确性、合规性和鲁棒性上表现最佳。
研究提出一种可扩展的有效性评估方法,用于评估生物医学大模型裁判,并通过四种训练方式测试其性能,发现SFT→RL组合在正确性、合规性和鲁棒性上表现最佳。
This study addresses the lack of flexible and interpretable probability distribution models suitable for upper-tail quantile analysis in clinical research. The authors propose a novel class of quantile-based effective duration functions, defined as the ratio of the mean to a given quantile, and derive a two-parameter family of non-negative distributions with closed-form expressions by incorporating Möbius transformations and natural boundary conditions. This distributional framework provides a unified characterization of tail behavior in survival data and facilitates quantile-based reliability measures and L-moment analysis. Empirical evaluation on real-world survival datasets demonstrates that the proposed method significantly outperforms existing approaches in both goodness-of-fit and model interpretability.
Traditional efficacy metrics struggle to capture the upper-tail persistence characteristics of high responders in biosimilar assessments. This work proposes a novel quantile-based efficacy persistence function, defined as the ratio of the tail mean to the quantile function, thereby introducing the concept of expected shortfall from risk theory into clinical persistence analysis for the first time. We demonstrate its equivalence to a scaled upper-tail first-order L-moment and develop a corresponding nonparametric estimator along with a two-sample equivalence test calibrated via bootstrap inference. Simulation studies and real-data analyses show that the proposed method effectively detects upper-tail differences undetectable by median- or mean-based approaches, substantially enhancing both sensitivity and specificity in biosimilar efficacy evaluation.
This study investigates the relationship between model scale and downstream task performance in structured healthcare data, an aspect that remains poorly understood. Leveraging insurance claims data from 519 hospitals in Japan, the authors develop and evaluate five Encoder-only Transformer foundation models ranging from 2.2M to 101M parameters for predicting disease onset and medication use. The work reveals, for the first time, a task-dependent performance saturation effect: disease prediction benefits from larger models, whereas medication prediction achieves optimal performance at just 11M parameters—reducing pretraining time by 178 hours. Across all tasks, the best-performing models consistently outperform a LightGBM baseline as measured by PR-AUC, offering empirical guidance for the efficient deployment of foundation models in healthcare settings.
This study addresses the challenge of substantial heterogeneity across patient subgroups when estimating treatment effects using electronic health records, which often undermines covariate balance in key clinical subpopulations under conventional propensity score weighting. To overcome this limitation, the authors propose a stratified propensity score weighting approach that first partitions patients based on clinical indications, reasons for admission, or risk factors, and then constructs separate propensity score models within each stratum to compute weights. This strategy prioritizes comparability within clinically meaningful subgroups and systematically accounts for differences in prognosis, heterogeneity in exposure probabilities, and covariate–subgroup interactions. Empirical analyses demonstrate that the proposed framework markedly improves covariate balance and enhances the accuracy of causal effect estimates within critical subgroups, offering a more robust method for causal inference in complex hospitalized populations.
研究提出一种可扩展的有效性评估方法,用于评估生物医学大模型裁判,并通过四种训练方式测试其性能,发现SFT→RL组合在正确性、合规性和鲁棒性上表现最佳。
This study addresses the lack of flexible and interpretable probability distribution models suitable for upper-tail quantile analysis in clinical research. The authors propose a novel class of quantile-based effective duration functions, defined as the ratio of the mean to a given quantile, and derive a two-parameter family of non-negative distributions with closed-form expressions by incorporating Möbius transformations and natural boundary conditions. This distributional framework provides a unified characterization of tail behavior in survival data and facilitates quantile-based reliability measures and L-moment analysis. Empirical evaluation on real-world survival datasets demonstrates that the proposed method significantly outperforms existing approaches in both goodness-of-fit and model interpretability.
Traditional efficacy metrics struggle to capture the upper-tail persistence characteristics of high responders in biosimilar assessments. This work proposes a novel quantile-based efficacy persistence function, defined as the ratio of the tail mean to the quantile function, thereby introducing the concept of expected shortfall from risk theory into clinical persistence analysis for the first time. We demonstrate its equivalence to a scaled upper-tail first-order L-moment and develop a corresponding nonparametric estimator along with a two-sample equivalence test calibrated via bootstrap inference. Simulation studies and real-data analyses show that the proposed method effectively detects upper-tail differences undetectable by median- or mean-based approaches, substantially enhancing both sensitivity and specificity in biosimilar efficacy evaluation.
This study investigates the relationship between model scale and downstream task performance in structured healthcare data, an aspect that remains poorly understood. Leveraging insurance claims data from 519 hospitals in Japan, the authors develop and evaluate five Encoder-only Transformer foundation models ranging from 2.2M to 101M parameters for predicting disease onset and medication use. The work reveals, for the first time, a task-dependent performance saturation effect: disease prediction benefits from larger models, whereas medication prediction achieves optimal performance at just 11M parameters—reducing pretraining time by 178 hours. Across all tasks, the best-performing models consistently outperform a LightGBM baseline as measured by PR-AUC, offering empirical guidance for the efficient deployment of foundation models in healthcare settings.
This study addresses the challenge of substantial heterogeneity across patient subgroups when estimating treatment effects using electronic health records, which often undermines covariate balance in key clinical subpopulations under conventional propensity score weighting. To overcome this limitation, the authors propose a stratified propensity score weighting approach that first partitions patients based on clinical indications, reasons for admission, or risk factors, and then constructs separate propensity score models within each stratum to compute weights. This strategy prioritizes comparability within clinically meaningful subgroups and systematically accounts for differences in prognosis, heterogeneity in exposure probabilities, and covariate–subgroup interactions. Empirical analyses demonstrate that the proposed framework markedly improves covariate balance and enhances the accuracy of causal effect estimates within critical subgroups, offering a more robust method for causal inference in complex hospitalized populations.