🤖 AI Summary
This study investigates whether structured mutual review and iterative prediction updating among heterogeneous large language models (LLMs)—specifically GPT-5, Claude Sonnet 4.5, and Gemini Pro 2.5—can improve binary event forecasting accuracy.
Method: Using 202 resolved questions from the Metaculus 2025 Q2 AI Forecasting Tournament, we evaluate performance via log loss and systematically compare four model-composition and information-sharing configurations.
Contribution/Results: A statistically significant log loss reduction of 0.020 (4% relative improvement, *p* = 0.017) is observed *only* under heterogeneous model composition with shared prediction information; homogeneous model deliberation and additional contextual inputs yield no gains. This constitutes the first empirical demonstration that cross-architectural LLM ensemble deliberation—conditional on shared predictions—yields statistically robust forecasting improvements. The findings identify prediction information sharing as a critical prerequisite, challenging prevailing assumptions about information aggregation in ensemble forecasting and establishing a novel paradigm for modeling LLM collective intelligence.
📝 Abstract
Structured deliberation has been found to improve the performance of human forecasters. This study investigates whether a similar intervention, i.e. allowing LLMs to review each other's forecasts before updating, can improve accuracy in large language models (GPT-5, Claude Sonnet 4.5, Gemini Pro 2.5). Using 202 resolved binary questions from the Metaculus Q2 2025 AI Forecasting Tournament, accuracy was assessed across four scenarios: (1) diverse models with distributed information, (2) diverse models with shared information, (3) homogeneous models with distributed information, and (4) homogeneous models with shared information. Results show that the intervention significantly improves accuracy in scenario (2), reducing Log Loss by 0.020 or about 4 percent in relative terms (p = 0.017). However, when homogeneous groups (three instances of the same model) engaged in the same process, no benefit was observed. Unexpectedly, providing LLMs with additional contextual information did not improve forecast accuracy, limiting our ability to study information pooling as a mechanism. Our findings suggest that deliberation may be a viable strategy for improving LLM forecasting.