A Verifier-Guided Explainable Reasoning Framework with Gold-Anchored QLoRA, Task-Aware Mixture-of-Experts, and Group-Relative RLVR

📅 2026-09-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出一种验证器引导的可解释推理框架,通过黄金锚定QLoRA、任务感知混合专家系统和组相对RLVR方法,提高大语言模型在教育问答中的解释性和准确性。
📝 Abstract
Large language models (LLMs) show strong reasoning ability, but their explanations can remain inconsistent, weakly grounded, or difficult to verify. We propose a verifier-guided explainable reasoning framework for transparent educational question answering that combines gold-anchored QLoRA, task-aware symbolic routing, and group-relative RLVR. Qwen2.5-3B-Instruct is first adapted with field-weighted QLoRA supervision anchored to authoritative answers. A lightweight router then assigns logic problems to a FOL/Z3 verifier and physics problems to a formula- and unit aware symbolic solver. Verifier feedback is further used to support candidate evaluation, self-revision, and reward construction during RLVR. Candidate responses are evaluated along three complementary dimensions: P1 for answer correctness, P2 for evidence or unit consistency, and P3 for reasoning depth and explainability. At inference, gold-free self-consistency aggregates multiple candidate responses before an optional question-only physics verifier performs conservative system-level correction. On 438 held-out examples, RLVR increases P3 from 50.68% to 72.20%, while hybrid P1 remains approximately stable at 55.94%. Self-consistency improves model only P1 from 48.86% to 50.23%, with symbolic verification providing the remaining hybrid gain. These results indicate that RLVR primarily strengthens explicit reasoning structure, while symbolic verification complements the neural policy by improving answer reliability at the system level.
Problem

Research questions and friction points this paper is trying to address.

large language models
reasoning ability
explanations
inconsistent
weakly grounded
Innovation

Methods, ideas, or system contributions that make the work stand out.

Verifier-Guided
Explainable Reasoning
Gold-Anchored QLoRA
Task-Aware Mixture-of-Experts
Group-Relative RLVR
T
Thi Kim Trang Vo
University of Information Technology (UIT), Ho Chi Minh City, Vietnam; Vietnam National University, Ho Chi Minh City, Vietnam
N
Nam Tien Le
Ho Chi Minh City University of Technology (HCMUT), Vietnam; Vietnam National University, Ho Chi Minh City, Vietnam
T
Thi Kim Nguyet Vo
University of Economics Ho Chi Minh City (UEH), Vietnam; Viet Nam – The Netherlands Programme (VNP), Vietnam
M
Minh Khang Tran
University of Information Technology (UIT), Ho Chi Minh City, Vietnam; Vietnam National University, Ho Chi Minh City, Vietnam
D
Duy Phuong Tran
University of Information Technology (UIT), Ho Chi Minh City, Vietnam; Vietnam National University, Ho Chi Minh City, Vietnam