FaithformBench: Benchmarking Faithfulness of Mathematical Chain-of-Thought Autoformalisation

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current automatic formalization systems lack evaluation methods that jointly assess faithfulness on both valid and invalid samples. This work proposes a low-cost, minimally assumptive benchmark that automatically generates perturbed reasoning steps to construct verifiable valid and invalid examples, introducing for the first time an evaluation mechanism based on invalid samples to expose a pervasive “sycophantic correction” behavior—where erroneous reasoning is silently transformed into provable formal statements. The framework employs dual metrics of validity-preserving and invalidity-preserving rates, making it applicable across diverse systems and datasets. Experiments on eight formalization systems and four mathematical datasets reveal that high validity-preserving performance often correlates with stronger sycophantic tendencies, uncovering a fundamental flaw in the faithfulness of current approaches.
📝 Abstract
Autoformalisation (AF) systems map natural language reasoning steps into formal statements in a proof assistant such as Lean. We consider how to assess the faithfulness of these systems. Existing approaches require expensive human-annotated ground truth, or rely on LLM judges or embedding models, which come with limited guarantees of accuracy. In addition, these methods typically only consider inputs that are known to be correct, and therefore do not assess whether the AF translates incorrect inputs faithfully. To address these limitations, we propose a new benchmark for AF faithfulness that is cheap to apply, sound under weak assumptions, and assesses both positive and negative examples. Our method is based on automatically generating perturbed reasoning steps that are designed to be invalid, and then measuring validity preservation on unperturbed steps and invalidity preservation on perturbed steps. We apply our method to eight AF systems across four mathematical datasets, and observe pervasive sycophancy: many AFs "silently correct" invalid inputs into provable statements. The most validity-preserving fine-tuned AFs are also the most sycophantic, suggesting a tension between validity and invalidity preservation in current AF systems.
Problem

Research questions and friction points this paper is trying to address.

Autoformalisation
faithfulness
mathematical reasoning
benchmarking
validity preservation
Innovation

Methods, ideas, or system contributions that make the work stand out.

faithfulness benchmarking
autoformalisation
validity preservation
invalidity preservation
sycophancy
🔎 Similar Papers
No similar papers found.