When Agents Disagree: Bayesian Backward Reasoning as a Label-Free Anchor for Multi-Agent Collective Decision-Making

๐Ÿ“… 2026-09-10
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆœฌๆ–‡้’ˆๅฏนๅคšๆ™บ่ƒฝไฝ“ๅ†ณ็ญ–ไธญๅ‡บ็Žฐๅ†ฒ็ช็ญ”ๆกˆ็š„้—ฎ้ข˜๏ผŒ้€š่ฟ‡่ดๅถๆ–ฏ้€†ๅ‘ๆŽจ็†ๆ–นๆณ•ๆไพ›ไบ†ไธ€ไธชๆ— ๆ ‡็ญพ็š„้”š็‚น๏ผŒไปฅๆ้ซ˜้›†ไฝ“ๅ†ณ็ญ–ๆ€ง่ƒฝใ€‚
๐Ÿ“ Abstract
When multiple LLM agents yield conflicting answers, the decision-making process dictates whether agent diversity improves performance or merely compounds shared errors. Existing collective decision-making methods, including voting, electoral rules, and LLM judges, rely on forward reasoning: they map evidence to labels in one direction. Although these methods can combine diverse forward traces, they still aggregate estimates that share this evidence-to-label factorization and can inherit correlated errors within the forward pool. We therefore construct a reverse posterior for each instance through Bayesian backward reasoning from an explicit likelihood. The forward and reverse posteriors provide differently factorized approximations of the underlying posterior. Because estimates from different factorizations may tend to share the same error less often, we use Jensen-Shannon divergence to rank agents by cross-path consistency. This cross-path consistency signal underlies three strategies: hard selection (MinJS), soft reweighting (FwdJS), and log-linear fusion (LogLin). Evaluated on DDXPlus across five LLM backbones, our proposed strategies show consistent improvements: MinJS outperforms random selection across all backbones, FwdJS generally improves over the strongest baseline, and LogLin achieves the best performance among the evaluated methods, with its largest gains on the subset where the agents disagree. Despite its weaker standalone accuracy, the reverse posterior serves as a more useful anchor than forward-only alternatives, providing complementary information for collective decision-making. When labeled data are available, a lightweight two-stage calibration can further refine the reverse anchor and improve aggregation performance.
Problem

Research questions and friction points this paper is trying to address.

conflicting answers
collective decision-making
forward reasoning
shared errors
Bayesian backward reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bayesian backward reasoning
Jensen-Shannon divergence
cross-path consistency
log-linear fusion
reverse posterior
๐Ÿ”Ž Similar Papers
No similar papers found.
K
Ken Chen
Department of Mechanical Engineering, The University of Melbourne, Melbourne, Australia
Wei Wang
Wei Wang
Department of Electrical and Electronic Engineering, University of Bristol
Sachith Seneviratne
Sachith Seneviratne
Research Fellow in Computer Vision, University Of Melbourne
Machine LearningComputer VisionNatural Language ProcessingUrban Informatics
H
Hansani Weeratunge
Department of Mechanical Engineering, Sri Lanka Institute of Information Technology, Sri Lanka
S
Saman Halgamuge
Department of Mechanical Engineering, The University of Melbourne, Melbourne, Australia