Institution profile

Cardiff Metropolitan University

Academic institutioneurope · gb
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Large Language Models for Spreadsheets: Benchmarking Progress and Evaluating Performance with FLARE

Jun 18, 2025

Current large language models (LLMs) exhibit insufficient precise symbolic reasoning capabilities for spreadsheet tasks, particularly suffering from “hallucinatory” errors in multi-step logical formula generation and structured data manipulation. To address this gap, we propose FLARE—the first comprehensive, spreadsheet-oriented benchmark designed to evaluate rigorous logical reasoning in realistic office scenarios. FLARE comprises three task categories: formula generation, data manipulation, and logical auditing. It integrates both synthetically constructed and real-world spreadsheet cases into a hierarchical task suite and introduces a dual-verification mechanism to ensure correctness and robustness. Experimental results demonstrate that while state-of-the-art LLMs achieve strong performance on simple formula tasks, their accuracy degrades substantially on complex, multi-step reasoning tasks—revealing critical deficiencies in structured-data reasoning. This work establishes a novel evaluation paradigm and provides a foundational benchmark for advancing the reliability and trustworthiness of LLMs in spreadsheet and structured-data applications.

0 citationsRead paper
Recent publications

Latest Papers

Large Language Models for Spreadsheets: Benchmarking Progress and Evaluating Performance with FLARE

Jun 18, 2025

Current large language models (LLMs) exhibit insufficient precise symbolic reasoning capabilities for spreadsheet tasks, particularly suffering from “hallucinatory” errors in multi-step logical formula generation and structured data manipulation. To address this gap, we propose FLARE—the first comprehensive, spreadsheet-oriented benchmark designed to evaluate rigorous logical reasoning in realistic office scenarios. FLARE comprises three task categories: formula generation, data manipulation, and logical auditing. It integrates both synthetically constructed and real-world spreadsheet cases into a hierarchical task suite and introduces a dual-verification mechanism to ensure correctness and robustness. Experimental results demonstrate that while state-of-the-art LLMs achieve strong performance on simple formula tasks, their accuracy degrades substantially on complex, multi-step reasoning tasks—revealing critical deficiencies in structured-data reasoning. This work establishes a novel evaluation paradigm and provides a foundational benchmark for advancing the reliability and trustworthiness of LLMs in spreadsheet and structured-data applications.

0 citationsRead paper