Answer Probing-Guided Search for Diverse Solution Exploration of LLMs

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决大型语言模型在生成多样化解决方案时的局限性,本文提出了一种基于答案探查指导的树搜索方法(APTS),通过探测潜在答案的隐藏状态和困惑度来提高解路径的多样性。
📝 Abstract
Generating multiple diverse and high-quality solutions is valuable for many applications, such as code-test generation and drug discovery. However, Large Language Models (LLMs) tend to converge on a single high-confidence solution during inference, limiting exploration of alternative valid solution paths. Existing test-time methods promote diversity through tree-like search and prune semantically similar branches using response-level semantic embeddings. However, we find that such embeddings are easily confounded by linguistic and stylistic similarities, making it difficult to distinguish genuinely distinct solution paths. To address this, we introduce Answer Probing, which probes the potential answer an LLM would reach from an intermediate reasoning path. We demonstrate that the hidden states of probed answers more effectively differentiate distinct solution paths than semantic embeddings, and the perplexity of probed answers serves as a practical proxy for reasoning correctness. Based on these findings, we propose Answer Probing-Guided Tree Search (APTS), which guides the tree search by the probed answers' hidden state similarity and perplexity. Experiments on three reasoning tasks across two LLMs show that APTS consistently enhances solution diversity, demonstrating its effectiveness and robustness.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Diverse Solutions
Solution Exploration
Semantic Embeddings
Reasoning Paths
Innovation

Methods, ideas, or system contributions that make the work stand out.

Answer Probing
hidden state similarity
perplexity
solution diversity
tree search
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yi Fang
University of Science and Technology of China
Q
Que Shen
Alibaba Group
Chengpeng Li
Chengpeng Li
USTC
LLMreasoningrecommendationreinforcement learning
Boyi Deng
Boyi Deng
University of Science and Technology of China
LLMsMechanistic Interpretability
W
Wei Shi
Shanghai Jiao Tong University
W
Wenjie Wang
University of Science and Technology of China
F
Fuli Feng
University of Science and Technology of China
Fengli Xu
Fengli Xu
Tsinghua University
LLM AgentData ScienceSocial ComputingScience of ScienceUrban Science
D
Dayiheng Liu
Alibaba Group