Ask, Condition or Abstain: Reinforcement Learning for Missing-Premise Reasoning

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inability of reasoning models to effectively inquire, provide conditional answers, or refuse responses when premises are missing. To this end, we propose the ACA-RL framework alongside a Missing Premise Benchmark (MPB). By leveraging reasoning graph-guided data augmentation and structured reinforcement learning, our approach enables fine-grained modeling of five distinct response behaviors and optimizes multi-behavior rewards. Experimental results demonstrate that ACA-RL significantly enhances the performance of Qwen3 and Llama models on MPB while preserving reasoning capabilities on complete problems. Consequently, this work effectively strengthens model robustness and decision-making under uncertainty, providing a viable solution for handling incomplete information in complex reasoning tasks.
📝 Abstract
Answer-only reinforcement learning (RL) trains reasoning models to solve fully specified problems, but many realistic queries omit a premise needed for a unique answer. In this setting, the useful response is not always refusal: the model should ask for the missing premise, condition its answer on the unknown quantity, or abstain when no informative conditional response is available. We present \emph{Ask-Condition-Abstain Reinforcement Learning} (ACA-RL), a data-augmented RL framework for this setting. Its reasoning-graph-guided pipeline converts well-posed problems into missing-premise training instances with localized gap annotations; ACA-RL then trains on these instances with a structured reward over five observable response behaviors. We also introduce the \emph{Missing-Premise Benchmark} (MPB), a 274-instance human-verified benchmark spanning mathematical, logical, and real-world word problems. Across Qwen3 and Llama models, ACA-RL consistently improves on MPB while preserving competitive performance on well-posed reasoning tasks. Together with the released code, MPB, and training data, this work supports a new mission for NLP evaluation: measuring whether models can recognize when a task is underdetermined and handle uncertainty, not only whether they can answer fully specified questions.
Problem

Research questions and friction points this paper is trying to address.

Missing-Premise Reasoning
Underdetermined Tasks
Reinforcement Learning
Uncertainty Handling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Ask-Condition-Abstain RL
Missing-Premise Reasoning
Structured Reward
Reasoning Graph
Missing-Premise Benchmark
Yongqi Tong
Yongqi Tong
Alibaba Group, UCSD
Natural Language Processing
Z
Zhenyu Zhang
Ant International
Z
Zimi Liu
Ant Group
K
Kewei Fu
Ant International
M
Mingli Song
Zhejiang University
H
Haofei Zhang
Zhejiang University
J
Junshao Zhang
Dingtalk, Alibaba Group
H
Hong Zhu
Dingtalk, Alibaba Group
Jiang-Ming Yang
Jiang-Ming Yang
Ant Financial
PaymentsInformation RetrieverConsistency Maintenance
X
Xin Zhang
Ant International
J
Jianshe Li
Ant International