You Only Pass Once: Answering and Abstaining Together in a Single Forward Pass of a Frozen Language Model

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the underutilization of evidence and the coupling between guidance and abstention mechanisms during frozen language model inference by proposing YOPO. This system employs an unlabeled residual reconstruction network to simultaneously enhance reasoning and assess information sufficiency within a single forward pass, effectively decoupling multi-task interference without additional computational overhead. Experiments demonstrate that YOPO doubles three-class accuracy and comprehensively outperforms dual-channel baselines across ten backbone models. Furthermore, its gating mechanism achieves state-of-the-art in-domain performance and represents the only unlabeled method supporting cross-domain transferability. These findings establish the novel principle that abstention capabilities should not be explicitly trained, offering a robust solution for reliable inference in frozen models.
📝 Abstract
A frozen language model on reasoning tasks has two coupled weaknesses: it under-uses evidence its own residual stream already encodes, and it fails to detect when the input is insufficient to answer, so it confabulates. This paper consolidates two research lines that address these on the same residual stream: a conditional steering probe writes the stream at mid-stack layers and recovers reasoning accuracy from a frozen backbone, and a zero-shot sufficiency direction reads the stream and abstains when information is insufficient. Deployed in one forward pass they interfere: the steering write shifts the state the direction reads, costing up to 8 AUROC points of cross-domain transfer on small models; a separate clean pass doubles inference cost. We keep the direction fixed and train a small network to reconstruct the pre-steering residual from the steered one -- mean-squared error on (steered, clean) pairs, no sufficiency labels -- and read the direction on the reconstruction. The resulting system, YOPO (You Only Pass Once), answers, steers, and abstains in one forward pass of a frozen Qwen2.5 backbone (1.5B/3B/7B). End to end, three-way accuracy more than doubles the frozen baseline (0.375->0.798 on 1.5B alphaNLI) and one pass beats the two-pass reference at every scale (0.798/0.830/0.893 vs 0.753/0.790/0.863) and on ten backbones across six model families. We chart the capacity-transfer frontier quantifying the principle that abstention should not be trained in; a source-side audit catches our own alphaNLI construction leaking a surface artifact, so architectural claims are anchored on native-label replications (SQuAD2, RepLiQA, MuSiQue); and on the standard four-domain suite we contribute, to our knowledge, the first answer-or-abstain benchmark, where our gate tops every in-domain dataset and the label-free direction is the only gate family to survive domain transfer.
Problem

Research questions and friction points this paper is trying to address.

Frozen Language Model
Abstention
Hallucination
Single Forward Pass
Residual Stream Interference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Single Forward Pass
Residual Stream Reconstruction
Conditional Steering
Zero-shot Abstention
Frozen Language Model
💼 Related Jobs
No related jobs found.
Ziyang Luo
Ziyang Luo
Salesforce AI Research
AgentsLLMsMultimodal
Z
Zhongyao Chu
X
Xinjie He
Y
Youting Wang
X
Xukui Qin
R
Runxiong Wu
Y
Yan-Syuan Chen