Institution profile

University of Richmond

Academic institutionnorthamerica · us
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Finding Answers in Thought Matters: Revisiting Evaluation on Large Language Models with Reasoning

Oct 16, 2025

Existing evaluations of large language models’ reasoning capabilities exhibit high sensitivity to answer extraction methods, resulting in unstable and inconsistent assessment outcomes. To address this, we propose Answer Regeneration (AR), an evaluation enhancement framework that decouples the reasoning process from answer extraction. AR introduces an additional reasoning step—prompting the model to regenerate its final answer based on its prior reasoning trace—thereby enabling robust answer extraction independent of heuristic rules. The framework is task-agnostic and applicable to diverse reasoning-intensive settings, including mathematical reasoning and open-domain question answering. Experiments across multiple benchmarks demonstrate that AR significantly improves evaluation robustness and accuracy, mitigating performance fluctuations induced by varying extraction strategies. Overall, AR provides a reliable, general-purpose solution for more stable and trustworthy assessment of LLM reasoning capabilities.

0 citationsRead paper
Recent publications

Latest Papers

Finding Answers in Thought Matters: Revisiting Evaluation on Large Language Models with Reasoning

Oct 16, 2025

Existing evaluations of large language models’ reasoning capabilities exhibit high sensitivity to answer extraction methods, resulting in unstable and inconsistent assessment outcomes. To address this, we propose Answer Regeneration (AR), an evaluation enhancement framework that decouples the reasoning process from answer extraction. AR introduces an additional reasoning step—prompting the model to regenerate its final answer based on its prior reasoning trace—thereby enabling robust answer extraction independent of heuristic rules. The framework is task-agnostic and applicable to diverse reasoning-intensive settings, including mathematical reasoning and open-domain question answering. Experiments across multiple benchmarks demonstrate that AR significantly improves evaluation robustness and accuracy, mitigating performance fluctuations induced by varying extraction strategies. Overall, AR provides a reliable, general-purpose solution for more stable and trustworthy assessment of LLM reasoning capabilities.

0 citationsRead paper