Multi-Level Bayesian Calibration of a Multi-Component Dynamic System Model
本文提出了一种多级贝叶斯校准方法,通过融合异构数据来解决时间依赖的多组件系统的建模和测量不确定性问题。
本文提出了一种多级贝叶斯校准方法,通过融合异构数据来解决时间依赖的多组件系统的建模和测量不确定性问题。
Current evaluations of large language models as judges (LLM-as-a-Judge) predominantly rely on single scalar metrics, overlooking their systematic biases and psychometric properties. This work proposes the Judge Datasheet protocol, treating LLM judges as measurement instruments and systematically characterizing their behavior through controlled experiments—including vacuum inputs, positional perturbations, and graded quality responses—augmented by novel concepts such as “dark current” and “direction–stability decomposition.” Evaluations across three open-source models reveal significant differences: Llama-3.1-8B exhibits high dark current and conflicting positional preferences, whereas Qwen2.5-32B demonstrates superior performance. Furthermore, prompt engineering is found to modulate only the decision threshold, not the underlying discriminative capability.
本文提出了一种多级贝叶斯校准方法,通过融合异构数据来解决时间依赖的多组件系统的建模和测量不确定性问题。
Current evaluations of large language models as judges (LLM-as-a-Judge) predominantly rely on single scalar metrics, overlooking their systematic biases and psychometric properties. This work proposes the Judge Datasheet protocol, treating LLM judges as measurement instruments and systematically characterizing their behavior through controlled experiments—including vacuum inputs, positional perturbations, and graded quality responses—augmented by novel concepts such as “dark current” and “direction–stability decomposition.” Evaluations across three open-source models reveal significant differences: Llama-3.1-8B exhibits high dark current and conflicting positional preferences, whereas Qwen2.5-32B demonstrates superior performance. Furthermore, prompt engineering is found to modulate only the decision threshold, not the underlying discriminative capability.