Institution profile

Mitsubishi Heavy Industries Limited

Industry researchasia · jp
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

LLM Judges Have Dark Current: A Psychometric Datasheet for LLM-as-a-Judge Evaluation

Jun 14, 2026

Current evaluations of large language models as judges (LLM-as-a-Judge) predominantly rely on single scalar metrics, overlooking their systematic biases and psychometric properties. This work proposes the Judge Datasheet protocol, treating LLM judges as measurement instruments and systematically characterizing their behavior through controlled experiments—including vacuum inputs, positional perturbations, and graded quality responses—augmented by novel concepts such as “dark current” and “direction–stability decomposition.” Evaluations across three open-source models reveal significant differences: Llama-3.1-8B exhibits high dark current and conflicting positional preferences, whereas Qwen2.5-32B demonstrates superior performance. Furthermore, prompt engineering is found to modulate only the decision threshold, not the underlying discriminative capability.

0 citationsRead paper
Recent publications

Latest Papers

LLM Judges Have Dark Current: A Psychometric Datasheet for LLM-as-a-Judge Evaluation

Jun 14, 2026

Current evaluations of large language models as judges (LLM-as-a-Judge) predominantly rely on single scalar metrics, overlooking their systematic biases and psychometric properties. This work proposes the Judge Datasheet protocol, treating LLM judges as measurement instruments and systematically characterizing their behavior through controlled experiments—including vacuum inputs, positional perturbations, and graded quality responses—augmented by novel concepts such as “dark current” and “direction–stability decomposition.” Evaluations across three open-source models reveal significant differences: Llama-3.1-8B exhibits high dark current and conflicting positional preferences, whereas Qwen2.5-32B demonstrates superior performance. Furthermore, prompt engineering is found to modulate only the decision threshold, not the underlying discriminative capability.

0 citationsRead paper