Divergent-Convergent Reasoning: Scaling Test-Time Compute through Structured Solution Synthesis

πŸ“… 2026-08-15
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the opaque mechanisms and inefficient resource allocation in test-time computation by proposing the Diverge-Converge Reasoning (DCR) framework. This approach synthesizes structured solutions through recursive reconciliation of divergent candidate answers, leveraging a training-free dispersion metric to dynamically allocate computational resources. Furthermore, this work elucidates the scaling laws governing the relationship between solution divergence and test-time performance gains. Empirical evaluations demonstrate that DCR achieves accuracies of 93.3% and 92.0% on AIME 2024 and 2025, respectively, while reducing average computational overhead by 27%. These results confirm that the proposed framework effectively enhances both the reasoning capabilities and resource efficiency of large language models during inference.
πŸ“ Abstract
Test-time compute can substantially improve Large Language Model (LLM) reasoning performance, yet how and when additional compute helps remains poorly understood. We study Divergent-Convergent Reasoning (DCR), a simple two-phase primitive consisting of an exploration phase that generates multiple candidate solutions followed by a convergent reconciliation phase. We present three core results. First, we show that even a single reconciliation step can reliably amplify correct minority reports: across datasets, DCR often recovers the correct answer when correct exploration outputs are in the minority, a regime where majority voting fails. Second, we introduce recursive DCR, an autoregressive reconciliation system that iteratively analyzes disagreements and allocates additional test-time compute. Recursive DCR achieves higher accuracy than fixed-compute baselines-reaching 93.3% on AIME 2024 and 92.0% on AIME 2025-while using roughly 27% less compute on average, demonstrating that attentive resource allocation is superior to uniform scaling. Third, we analyze disagreement among exploration outputs via a simple, training-free dispersion metric. Dispersion reveals a structured relationship between disagreement and test-time gains: in regimes where DCR is effective, higher disagreement among exploration outputs is associated with larger accuracy improvements from reconciliation. Together, these results show that disagreement, often viewed as noise, can be systematically exploited to improve test-time reasoning and reveal emerging scaling laws for agentic LLM systems.
Problem

Research questions and friction points this paper is trying to address.

Test-time compute
LLM reasoning
Divergent-Convergent Reasoning
Disagreement exploitation
Scaling laws
Innovation

Methods, ideas, or system contributions that make the work stand out.

Divergent-Convergent Reasoning
Test-Time Compute
Recursive Reconciliation
Dispersion Metric
Scaling Laws
πŸ’Ό Related Jobs
No related jobs found.
B
Bo Wen
Enkira.ai, USA
Y
Yuhao Chen
School of Computing, Queen’s University, Kingston, Ontario, Canada
Erhan Bilal
Erhan Bilal
Enkira.ai, USA
C
Carla Agurto Rios
Enkira.ai, USA
C
Chen Wang
IBM T.J. Watson Research Center, USA
Junchen Jiang
Junchen Jiang
University of Chicago
Computer Networks