Beyond Solver Verdicts: Generative Reward Models for Autoformalization

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了神经符号系统中形式化翻译不保真问题,提出Generative Verification方法,通过生成参考等价分数提高验证准确性。
📝 Abstract
Neurosymbolic systems rely on mathematical solvers to guarantee reasoning correctness, yet solvers are fundamentally blind to whether a formal translation maintains strict reference-equivalence to a designated formalization. We formalize this vulnerability as Verdict-Preserving-Unfaithfulness (VPU): a failure mode where an incorrect encoding executes successfully and matches the expected verdict. We theoretically prove that structural, verdict-only verification heuristics are mathematically bounded to chance-level detection on these deceptively valid traces. To resolve this, we introduce Generative Verification (GenV), which distills an offline Z3-equivalence oracle into a reference-free, continuous reference-equivalence score by repurposing the language model's native vocabulary space. Mechanistic analysis via decision-projected logit lenses and sparse autoencoders shows this generative readout natively extracts precise spatial error coordinates without explicit localization training. Empirically, our oracle-mined verifier (GenV+HN) achieves 0.961 AUROC in reference-equivalence verification, generalizes zero-shot across unseen translators and divergent formal styles, and yields an 11.3-point downstream accuracy gain in agentic test-time compute allocation.
Problem

Research questions and friction points this paper is trying to address.

Neurosymbolic systems
Verdict-Preserving-Unfaithfulness (VPU)
reference-equivalence
mathematical solvers
formal translation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generative Verification
Verdict-Preserving-Unfaithfulness
reference-equivalence score
language model's native vocabulary space