Institution profile

Squirrel Ai Learning

Industry researchasia · cn
Official website
Research library92linked papers
Opportunities0open roles
Selected work

Representative Papers

MathAgent: Leveraging a Mixture-of-Math-Agent Framework for Real-World Multimodal Mathematical Error Detection

Mar 23, 2025

Accurately identifying errors in K–12 multimodal math assignments—comprising both handwritten or typeset text and diagrams—remains challenging, as current multimodal large language models (MLLMs) lack the capability to jointly reason over image and text modalities and precisely localize and attribute solution-step errors. Method: This paper proposes the first three-stage mathematical agent hybrid framework: (1) image-text consistency verification, (2) visual-semantic parsing, and (3) cross-modal error integration analysis—explicitly modeling multimodal associations between problem statements and solution steps. The framework integrates vision understanding, symbolic logical reasoning, and pedagogically grounded constraints within a specialized collaborative agent architecture. Contribution/Results: Evaluated on real-world educational data, the framework improves step-level error detection accuracy by 5% and error-type classification accuracy by 3%. It has been deployed at scale across a platform serving over one million students, achieving 90% user satisfaction and substantially reducing manual review overhead.

1 citationsRead paper

HalluTracer: Hallucination Detection via Depth-Averaging Truth Signals

Aug 17, 2026

This study addresses the limitation of existing white-box hallucination detection methods that overlook cross-layer truthfulness signals in large language models. We propose a Depth Averaging framework that exploits the geometric sparsity of layer-wise signals to reconstruct the detection paradigm from single-layer selection to depth aggregation. This approach effectively extracts linearly separable truthfulness features while suppressing noise. Extensive evaluations across six open-source models and five benchmarks demonstrate that our framework significantly outperforms current white-box baselines, achieving performance gains of 1 to 14 percentage points. These results validate the critical role of cross-layer signal aggregation in hallucination detection and offer a novel perspective for advancing white-box interpretability and reliability assessment in large language models.

0 citationsRead paper
Recent publications

Latest Papers

HalluTracer: Hallucination Detection via Depth-Averaging Truth Signals

Aug 17, 2026

This study addresses the limitation of existing white-box hallucination detection methods that overlook cross-layer truthfulness signals in large language models. We propose a Depth Averaging framework that exploits the geometric sparsity of layer-wise signals to reconstruct the detection paradigm from single-layer selection to depth aggregation. This approach effectively extracts linearly separable truthfulness features while suppressing noise. Extensive evaluations across six open-source models and five benchmarks demonstrate that our framework significantly outperforms current white-box baselines, achieving performance gains of 1 to 14 percentage points. These results validate the critical role of cross-layer signal aggregation in hallucination detection and offer a novel perspective for advancing white-box interpretability and reliability assessment in large language models.

0 citationsRead paper

TimeSage-EV: A Live Benchmark for Agentic Time Series Analysis in Evolving Environments

Aug 14, 2026

This study addresses the limitations of static temporal question-answering benchmarks by introducing the first dynamic, real-time benchmark spanning six domains and sixty scenarios, alongside a self-evolving agent framework equipped with a reusable skill library. Through multi-granularity data release and an LLM-based evaluation system, this work systematically assesses model capabilities in state recognition and prediction within evolving environments. The results reveal significant performance gaps in frontier models regarding temporal validity and context utilization. Furthermore, by providing monthly updated resources and comprehensive failure mode analyses, this research establishes a new paradigm for investigating dynamic temporal reasoning, overcoming the lack of timeliness assessment in existing methodologies.

0 citationsRead paper