When Text and Numbers Disagree: Evidence Arbitration in Large Language Models
研究通过引入合成基准探讨大型语言模型如何在文本与数值冲突时进行仲裁,发现模型偏好使用特定策略而非随机选择。
研究通过引入合成基准探讨大型语言模型如何在文本与数值冲突时进行仲裁,发现模型偏好使用特定策略而非随机选择。
This study addresses the vulnerability of existing Bayesian dynamic borrowing methods to parametric model misspecification by proposing a nonparametric latent exchangeable prior framework. Integrating Bayesian model averaging with kernel methods, this approach enables individual-level assessment for historical data borrowing without requiring outcome model assumptions, thereby effectively mitigating triple misspecification risks while ensuring posterior consistency. Simulation studies demonstrate that the proposed method outperforms conventional parametric and semiparametric alternatives. Furthermore, its efficacy is successfully validated in a lung cancer clinical trial. Collectively, this work provides a robust nonparametric solution for dynamic information borrowing, offering significant improvements in reliability over traditional approaches when model assumptions are uncertain or violated.
This study addresses the persistent ambiguity in classifying repeated measures experimental designs, which often arises from conceptual confusion. To resolve this issue, the authors systematically clarify the core characteristics of such designs and propose a novel classification framework grounded in experimental units and randomization strategies. For the first time in this context, Hasse diagrams are introduced to visually represent the hierarchical structure of these designs. This approach effectively distinguishes among various types of repeated measures designs, eliminates terminological ambiguities, and substantially enhances both the rigor and interpretability of experimental planning and reporting.
This study systematically evaluates and challenges the scientific validity of the “expected f2” approach for comparing dissolution profiles, particularly its suitability as a replacement for the conventional f2 metric under conditions of high variability. Through comprehensive literature review, statistical analysis, and expert consultation, the EFSPI CMCSNE SIG working group reveals fundamental flaws in the method, including the absence of original theoretical justification, mathematical bias, low statistical power, and ambiguous definition. The findings strongly advise against incorporating “expected f2” into regulatory guidance, thereby providing critical scientific evidence to inform policy decisions and addressing a significant gap in the systematic critique of this methodology.
Existing rule-based approaches struggle to comprehensively capture the multidimensional nature of disease severity in electronic health records. This work proposes MOSAIC, a novel framework that, for the first time, leverages a multi-agent large language model system to assess phenotypic severity in type 2 diabetes by integrating key dimensions such as biomarkers, β-cell function, and social determinants of health, thereby overcoming the rigidity of fixed-rule systems. Trained and validated against established clinical scoring criteria (DCSI, DiSSCo, Cooper) and real-world clinical outcomes, the open-source implementation demonstrates strong agreement with its closed-source counterpart (weighted kappa = 0.773). Stratified severity levels significantly differentiate all-cause mortality and complication risk (log-rank p < 0.001), outperforming rule-based baselines.
研究通过引入合成基准探讨大型语言模型如何在文本与数值冲突时进行仲裁,发现模型偏好使用特定策略而非随机选择。
This study addresses the vulnerability of existing Bayesian dynamic borrowing methods to parametric model misspecification by proposing a nonparametric latent exchangeable prior framework. Integrating Bayesian model averaging with kernel methods, this approach enables individual-level assessment for historical data borrowing without requiring outcome model assumptions, thereby effectively mitigating triple misspecification risks while ensuring posterior consistency. Simulation studies demonstrate that the proposed method outperforms conventional parametric and semiparametric alternatives. Furthermore, its efficacy is successfully validated in a lung cancer clinical trial. Collectively, this work provides a robust nonparametric solution for dynamic information borrowing, offering significant improvements in reliability over traditional approaches when model assumptions are uncertain or violated.
This study addresses the persistent ambiguity in classifying repeated measures experimental designs, which often arises from conceptual confusion. To resolve this issue, the authors systematically clarify the core characteristics of such designs and propose a novel classification framework grounded in experimental units and randomization strategies. For the first time in this context, Hasse diagrams are introduced to visually represent the hierarchical structure of these designs. This approach effectively distinguishes among various types of repeated measures designs, eliminates terminological ambiguities, and substantially enhances both the rigor and interpretability of experimental planning and reporting.
This study systematically evaluates and challenges the scientific validity of the “expected f2” approach for comparing dissolution profiles, particularly its suitability as a replacement for the conventional f2 metric under conditions of high variability. Through comprehensive literature review, statistical analysis, and expert consultation, the EFSPI CMCSNE SIG working group reveals fundamental flaws in the method, including the absence of original theoretical justification, mathematical bias, low statistical power, and ambiguous definition. The findings strongly advise against incorporating “expected f2” into regulatory guidance, thereby providing critical scientific evidence to inform policy decisions and addressing a significant gap in the systematic critique of this methodology.
Existing rule-based approaches struggle to comprehensively capture the multidimensional nature of disease severity in electronic health records. This work proposes MOSAIC, a novel framework that, for the first time, leverages a multi-agent large language model system to assess phenotypic severity in type 2 diabetes by integrating key dimensions such as biomarkers, β-cell function, and social determinants of health, thereby overcoming the rigidity of fixed-rule systems. Trained and validated against established clinical scoring criteria (DCSI, DiSSCo, Cooper) and real-world clinical outcomes, the open-source implementation demonstrates strong agreement with its closed-source counterpart (weighted kappa = 0.773). Stratified severity levels significantly differentiate all-cause mortality and complication risk (log-rank p < 0.001), outperforming rule-based baselines.