Disentangling Topology and Diversity in Multi-Agent LLMs for Multilingual Low-Resource Emotion Detection
研究通过独立分析推理拓扑和代理间多样性来源,解决多语言低资源情感检测问题,使用不同配置的多代理LLM系统进行评估。
研究通过独立分析推理拓扑和代理间多样性来源,解决多语言低资源情感检测问题,使用不同配置的多代理LLM系统进行评估。
Current large language models (LLMs) exhibit insufficient precise symbolic reasoning capabilities for spreadsheet tasks, particularly suffering from “hallucinatory” errors in multi-step logical formula generation and structured data manipulation. To address this gap, we propose FLARE—the first comprehensive, spreadsheet-oriented benchmark designed to evaluate rigorous logical reasoning in realistic office scenarios. FLARE comprises three task categories: formula generation, data manipulation, and logical auditing. It integrates both synthetically constructed and real-world spreadsheet cases into a hierarchical task suite and introduces a dual-verification mechanism to ensure correctness and robustness. Experimental results demonstrate that while state-of-the-art LLMs achieve strong performance on simple formula tasks, their accuracy degrades substantially on complex, multi-step reasoning tasks—revealing critical deficiencies in structured-data reasoning. This work establishes a novel evaluation paradigm and provides a foundational benchmark for advancing the reliability and trustworthiness of LLMs in spreadsheet and structured-data applications.
研究通过独立分析推理拓扑和代理间多样性来源,解决多语言低资源情感检测问题,使用不同配置的多代理LLM系统进行评估。
Current large language models (LLMs) exhibit insufficient precise symbolic reasoning capabilities for spreadsheet tasks, particularly suffering from “hallucinatory” errors in multi-step logical formula generation and structured data manipulation. To address this gap, we propose FLARE—the first comprehensive, spreadsheet-oriented benchmark designed to evaluate rigorous logical reasoning in realistic office scenarios. FLARE comprises three task categories: formula generation, data manipulation, and logical auditing. It integrates both synthetically constructed and real-world spreadsheet cases into a hierarchical task suite and introduces a dual-verification mechanism to ensure correctness and robustness. Experimental results demonstrate that while state-of-the-art LLMs achieve strong performance on simple formula tasks, their accuracy degrades substantially on complex, multi-step reasoning tasks—revealing critical deficiencies in structured-data reasoning. This work establishes a novel evaluation paradigm and provides a foundational benchmark for advancing the reliability and trustworthiness of LLMs in spreadsheet and structured-data applications.