FlexEE: Self-Speculative and KV-Compatible Early Exiting for Offloading-Aware LLM Inference
为解决大语言模型推理中的计算和内存限制问题,FlexEE通过早期退出框架减少执行层数,利用层间监督、自推测解码和动态隐藏状态管理,实现高效推理。
为解决大语言模型推理中的计算和内存限制问题,FlexEE通过早期退出框架减少执行层数,利用层间监督、自推测解码和动态隐藏状态管理,实现高效推理。
研究提出了一种轻量级的多模态视觉-语言框架,使用TinyCLIP进行苹果幼果解剖结构分类,以支持果园精准操作。
Industrial table question answering (TQA) faces three key challenges: structural heterogeneity, difficulty in target data localization, and bottlenecks in complex reasoning. To address these, this paper proposes a large language model (LLM)-based programmable agent framework. The method replaces raw textual tables with structured schema representations and introduces a query-aware dynamic subtable scaling mechanism. It integrates Program-of-Thoughts (PoT) with the ReAct paradigm to construct an executable, iterative reasoning pipeline. Column selection and entity linking are incorporated to jointly enhance semantic understanding and code generation, thereby improving localization accuracy and reasoning controllability. Evaluated on DataBench and TableBench, the framework achieves absolute accuracy gains of 19.34% and 25%, respectively, demonstrating its effectiveness and strong scalability across multi-scale tabular data.
Amidst surging demand for mental health support, existing empathetic dialogue generation systems exhibit suboptimal performance. To address this, we propose a large language model (LLM) optimization framework integrating prompt engineering and efficient fine-tuning. Within a unified experimental setup, we systematically compare the impact of low-rank adaptation (LoRA) versus full-parameter fine-tuning on empathetic capability, and design a structured, empathy-oriented prompt template. Our method significantly enhances the model’s depth of emotional state understanding, empathetic response accuracy, and contextual coherence. Evaluated on the authoritative NLPCC 2025 Task 8 ESC benchmark, our best-performing model ranks second, demonstrating that our approach effectively balances generation quality, training efficiency, and empathetic efficacy. This work establishes a reproducible, scalable technical paradigm for trustworthy AI-powered psychological assistance.
Real-world table question answering (TQA) faces challenges including large-scale tables, incomplete column semantics, and entity ambiguity. Method: This paper proposes a novel framework integrating large language models (LLMs) with programmatic reasoning. It (1) introduces a multi-step schema linking strategy that dynamically generates structure-focused table representations to mitigate semantic ambiguity; (2) designs an iterative “Think–Reason–Reflect” architecture for joint structural and semantic modeling; and (3) incorporates LLM-driven programmable reasoning to generate interpretable, executable SQL-like queries. Contribution/Results: The framework achieves first place on both subtasks of SemEval-2025 Task 8, significantly improving reasoning accuracy and robustness on complex, large-scale, and low-quality tables. It establishes a new paradigm for semantic understanding of real-world tabular data characterized by structural sparsity and lexical ambiguity.
为解决大语言模型推理中的计算和内存限制问题,FlexEE通过早期退出框架减少执行层数,利用层间监督、自推测解码和动态隐藏状态管理,实现高效推理。
研究提出了一种轻量级的多模态视觉-语言框架,使用TinyCLIP进行苹果幼果解剖结构分类,以支持果园精准操作。
Industrial table question answering (TQA) faces three key challenges: structural heterogeneity, difficulty in target data localization, and bottlenecks in complex reasoning. To address these, this paper proposes a large language model (LLM)-based programmable agent framework. The method replaces raw textual tables with structured schema representations and introduces a query-aware dynamic subtable scaling mechanism. It integrates Program-of-Thoughts (PoT) with the ReAct paradigm to construct an executable, iterative reasoning pipeline. Column selection and entity linking are incorporated to jointly enhance semantic understanding and code generation, thereby improving localization accuracy and reasoning controllability. Evaluated on DataBench and TableBench, the framework achieves absolute accuracy gains of 19.34% and 25%, respectively, demonstrating its effectiveness and strong scalability across multi-scale tabular data.
Amidst surging demand for mental health support, existing empathetic dialogue generation systems exhibit suboptimal performance. To address this, we propose a large language model (LLM) optimization framework integrating prompt engineering and efficient fine-tuning. Within a unified experimental setup, we systematically compare the impact of low-rank adaptation (LoRA) versus full-parameter fine-tuning on empathetic capability, and design a structured, empathy-oriented prompt template. Our method significantly enhances the model’s depth of emotional state understanding, empathetic response accuracy, and contextual coherence. Evaluated on the authoritative NLPCC 2025 Task 8 ESC benchmark, our best-performing model ranks second, demonstrating that our approach effectively balances generation quality, training efficiency, and empathetic efficacy. This work establishes a reproducible, scalable technical paradigm for trustworthy AI-powered psychological assistance.
Real-world table question answering (TQA) faces challenges including large-scale tables, incomplete column semantics, and entity ambiguity. Method: This paper proposes a novel framework integrating large language models (LLMs) with programmatic reasoning. It (1) introduces a multi-step schema linking strategy that dynamically generates structure-focused table representations to mitigate semantic ambiguity; (2) designs an iterative “Think–Reason–Reflect” architecture for joint structural and semantic modeling; and (3) incorporates LLM-driven programmable reasoning to generate interpretable, executable SQL-like queries. Contribution/Results: The framework achieves first place on both subtasks of SemEval-2025 Task 8, significantly improving reasoning accuracy and robustness on complex, large-scale, and low-quality tables. It establishes a new paradigm for semantic understanding of real-world tabular data characterized by structural sparsity and lexical ambiguity.