When Text and Numbers Disagree: Evidence Arbitration in Large Language Models

📅 2026-08-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过引入合成基准探讨大型语言模型如何在文本与数值冲突时进行仲裁,发现模型偏好使用特定策略而非随机选择。
📝 Abstract
Large language models (LLMs) are increasingly used in settings where textual summaries, numerical observations, and external tool outputs may provide conflicting evidence. We study how LLMs arbitrate between such sources when they support opposing decisions. To do so, we introduce a controlled synthetic benchmark in which latent risk trajectories generate both numerical time series and natural language summaries, allowing us to construct conflicts where exactly one evidence source is aligned with the ground-truth label. This design lets us independently manipulate modality, temporal recency, source reliability, and evidence provenance. Across open-weight instruction-tuned models, we find that arbitration behaviour is systematic rather than random: models exhibit distinct text-versus-number preferences, follow temporal recency more consistently than explicit reliability cues, and can over-rely on external forecasts even when they conflict with direct contextual evidence. These results suggest that current LLMs often rely on heuristic arbitration strategies when integrating heterogeneous evidence, highlighting a failure mode for tool-augmented decision systems.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Evidence Arbitration
Conflicting Evidence
Innovation

Methods, ideas, or system contributions that make the work stand out.

controlled synthetic benchmark
evidence arbitration
heterogeneous evidence
temporal recency
🔎 Similar Papers
No similar papers found.
M
Mattia Carletti
Department of Engineering Science, University of Oxford, Oxford, UK
E
Edward Phillips
Department of Engineering Science, University of Oxford, Oxford, UK
F
Fredrik K. Gustafsson
Department of Engineering Science, University of Oxford, Oxford, UK
P
Patitapaban Palo
Department of Engineering Science, University of Oxford, Oxford, UK
Lei Clifton
Lei Clifton
Nuffield Department of Primary Care Health Sciences, University of Oxford
AI & Machine learningMedical statistics
Danielle Belgrave
Danielle Belgrave
GSK AI
StatisticsMachine LearningHealthcare
Xiao Gu
Xiao Gu
University of Oxford
AI for HealthcareBiomedical Signal ProcessingWearable/Ambient IntelligenceDeep Learning
David A. Clifton
David A. Clifton
Chair of Clinical Machine Learning, University of Oxford
Machine LearningClinical AIBiomedical Signal Processing