How to Speculate about Uncertainty in Agentic Coding? A Draft-Model Gate Method

📅 2026-09-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种名为推测不确定性(SU)的方法,通过分析代理输出的令牌来预测其失败的可能性,无需访问内部参数。该方法在软件工程中减少了执行错误率和成本。
📝 Abstract
LLM agents deployed for software engineering fail expensively: they act confidently wrong, and bad actions are recognized only after costly execution and retry. We present Speculative Uncertainty (SU), a method that recovers a predictive failure signal for a black-box agent from its output tokens alone, with no access to logits, weights, activations, or repeated sampling. Inverting speculative decoding, a small open-weight draft model scores the agent's already-generated trajectory in a single forward pass. From these speculative cross-likelihoods we extract phase-aware features by separating the reasoning and action spans, and calibrate them against a verifiable objective. SU produces a failure-likelihood score that any downstream policy, such as routing, human intervention, or extra test-time compute, can consume directly. To show the signal is actionable, we instantiate one such policy, a pre-execution veto gate, on software engineering agents Qwen3-Coder-480B and closed-source Claude 3.5 Sonnet, cutting execution error rate by 6-8 percentage points and token cost by 14-19% in deployment, transferring to out-of-distribution benchmarks without retraining, and generalizing across agent models.
Problem

Research questions and friction points this paper is trying to address.

Speculative Uncertainty
Software Engineering
Failure Prediction
Agent Output Tokens
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speculative Uncertainty
black-box agent
speculative decoding
phase-aware features
failure-likelihood score
🔎 Similar Papers
No similar papers found.