Do Reasoning Representations Help Humans Evaluate LLM Outputs?

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了不同推理表示格式对人类评估大语言模型输出的帮助,通过控制实验发现简洁的思维链比复杂的规划分解更支持验证和理解。
📝 Abstract
Reasoning representations are increasingly used as explanations for large language model outputs. Yet they are typically evaluated with model-centric criteria, such as answer accuracy and faithfulness, leaving it unclear whether they help people evaluate model responses. In this work, we study reasoning representations as human-facing interfaces rather than proxies for model reasoning ability. We conduct a controlled human study of six reasoning formats across tasks of varying complexity, supported by a web-based framework that randomizes task domains, problem instances, and representation order. The study collects fine-grained judgments of structural understanding, error detection and localization, and trust calibration. Our study shows a mismatch between perceived preference and support for human evaluation. Participants prefer planning- and decomposition-based representations, but simpler chain-of-thought traces better support verification, trust, and interpretability. Preferred representations also introduce calibration risks, with more false alarms on correct traces and high trust despite low willingness to verify.
Problem

Research questions and friction points this paper is trying to address.

Reasoning Representations
Human Evaluation
Large Language Models
Structural Understanding
Trust Calibration
Innovation

Methods, ideas, or system contributions that make the work stand out.

reasoning representations
human evaluation
chain-of-thought
🔎 Similar Papers
No similar papers found.