Steering Recurrent Reasoners at Inference Time with Readout Feedback

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过在推理时使用读出反馈(RoFB)方法引导循环模型的潜在动态,提高了模型在数独和迷宫任务上的性能,而无需重新训练。
📝 Abstract
Recurrent models, which repeatedly update latent states with shared computation blocks, have emerged as powerful architectures for solving complex reasoning tasks. Existing inference-time methods scale computation by running more steps or sampling more trajectories, but ignore information revealed within each trajectory. Here we show that recurrent models can be improved at inference time by using their own readout probabilities to steer latent dynamics without retraining. We introduce Readout Feedback (RoFB), a test-time intervention that converts intermediate predictions into token-wise pairwise coupling forces injected into the latent dynamics. Across three recurrent models (AKOrN, ItrSA++, TRM) on Sudoku and Maze, RoFB yields clear gains in four of six model-task pairs, achieving performance unattainable by merely running more steps or selecting from multiple trajectories, at comparable or lower computational cost. These results suggest that closed-loop steering of latent dynamics can serve as a complementary inference-time control mechanism for recurrent reasoning models.
Problem

Research questions and friction points this paper is trying to address.

Recurrent Models
Inference Time
Readout Probabilities
Latent Dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Readout Feedback
Recurrent Models
Latent Dynamics
Inference-time Control
🔎 Similar Papers