PTCG: Persona-guided Tree-based Counterargument Generation
为解决现有方法生成单一反论的问题,提出PTCG框架,通过结合树状思维启发的逐步生成与修剪及说话人角色选择,生成多样化且有说服力的反论。
为解决现有方法生成单一反论的问题,提出PTCG框架,通过结合树状思维启发的逐步生成与修剪及说话人角色选择,生成多样化且有说服力的反论。
Existing evaluations of large language models’ reasoning capabilities exhibit high sensitivity to answer extraction methods, resulting in unstable and inconsistent assessment outcomes. To address this, we propose Answer Regeneration (AR), an evaluation enhancement framework that decouples the reasoning process from answer extraction. AR introduces an additional reasoning step—prompting the model to regenerate its final answer based on its prior reasoning trace—thereby enabling robust answer extraction independent of heuristic rules. The framework is task-agnostic and applicable to diverse reasoning-intensive settings, including mathematical reasoning and open-domain question answering. Experiments across multiple benchmarks demonstrate that AR significantly improves evaluation robustness and accuracy, mitigating performance fluctuations induced by varying extraction strategies. Overall, AR provides a reliable, general-purpose solution for more stable and trustworthy assessment of LLM reasoning capabilities.
为解决现有方法生成单一反论的问题,提出PTCG框架,通过结合树状思维启发的逐步生成与修剪及说话人角色选择,生成多样化且有说服力的反论。
Existing evaluations of large language models’ reasoning capabilities exhibit high sensitivity to answer extraction methods, resulting in unstable and inconsistent assessment outcomes. To address this, we propose Answer Regeneration (AR), an evaluation enhancement framework that decouples the reasoning process from answer extraction. AR introduces an additional reasoning step—prompting the model to regenerate its final answer based on its prior reasoning trace—thereby enabling robust answer extraction independent of heuristic rules. The framework is task-agnostic and applicable to diverse reasoning-intensive settings, including mathematical reasoning and open-domain question answering. Experiments across multiple benchmarks demonstrate that AR significantly improves evaluation robustness and accuracy, mitigating performance fluctuations induced by varying extraction strategies. Overall, AR provides a reliable, general-purpose solution for more stable and trustworthy assessment of LLM reasoning capabilities.