Institution profile

Hangzhou Institute of Advance Studies, University of Chinese Academy of Science

Academic institutionasia · cn
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

Socratic agents for autonomous scientific discovery in high-dimensional physical systems

Jun 25, 2026

This work addresses the limited cognitive autonomy of traditional AI in scientific discovery by introducing AHOIS, a multi-agent AI scientist that incorporates a Socratic questioning mechanism into autonomous physical exploration. By leveraging causal interrogation, counterexample generation, and falsifiability-driven hypothesis refinement, AHOIS autonomously proposes, tests, and revises hypotheses without relying on prior models. The system integrates causal reasoning, constraint verification, sparse measurement optimization, and uncertainty calibration within a closed-loop experimental framework. Deployed on a multimode fiber platform, AHOIS autonomously discovered a stochastic interference encoding scheme, achieving classification accuracies of 76.97% on MNIST and 83.17% on Fashion-MNIST, while effectively diagnosing multiple failure modes and substantially enhancing the consistency and completeness of physical interpretability.

0 citationsRead paper

Beyond Endpoints: Path-Centric Reasoning for Vectorized Off-Road Network Extraction

Dec 11, 2025

To address topological errors (e.g., structural disconnections due to occlusion or ambiguous junctions) in vectorized off-road road extraction caused by domain shift, this paper introduces a path-centric inference paradigm, overcoming the robustness limitations of conventional node-centric approaches. Methodologically, we propose MaGRoad—a mask-aware geodesic road extraction framework integrating multi-scale visual evidence aggregation, geodesic path modeling, and lightweight vector decoding. Our contributions are threefold: (1) We release WildRoad, the first global off-road road dataset, accompanied by an interactive annotation tool; (2) We design MaGRoad to explicitly model continuous road centerlines via geodesic distance priors and mask-guided feature fusion; (3) MaGRoad achieves state-of-the-art performance on WildRoad, demonstrates strong cross-domain generalization to urban road datasets (e.g., DeepGlobe, Cowc), and operates 2.5× faster than prior methods.

0 citationsRead paper

Centering Emotion Hotspots: Multimodal Local-Global Fusion and Cross-Modal Alignment for Emotion Recognition in Conversations

Oct 07, 2025

Emotion Recognition in Conversations (ERC) faces challenges stemming from sparse, localized, and asynchronous multimodal evidence. To address these, we propose a multimodal ERC framework centered on “emotion hotspots”: (1) local emotion-critical segments are first identified within textual, acoustic, and visual modalities; (2) a hotspot-gated fusion mechanism adaptively weights local hotspots against global contextual representations; (3) a routing-based hybrid aligner enables fine-grained cross-modal alignment; and (4) a dialogue-structure graph models inter-utterance dependencies. The method integrates local-global feature modeling, graph neural networks, dynamic attention, and gating mechanisms. Evaluated on standard benchmarks, our approach significantly outperforms strong baselines. Ablation studies confirm the effectiveness of each component. Overall, the framework achieves robust, interpretable, fine-grained emotion recognition by jointly leveraging modality-specific saliency and structured conversational context.

0 citationsRead paper

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition

Jul 16, 2025

Dynamic Facial Expression Recognition (DFER) faces two key challenges: insufficient exploitation of fine-grained affective cues from generated textual descriptions, and difficulty suppressing facial motions irrelevant to emotion. To address these, we propose GRACE—a framework that achieves token-level cross-modal alignment between linguistic cues and visually salient regions via coarse-to-fine emotional text enhancement and motion-difference-weighted attention. GRACE further incorporates dynamic motion modeling, semantic text refinement, and entropy-regularized optimal transport for precise spatiotemporal localization of emotion-relevant features. Evaluated on three benchmark datasets, GRACE achieves state-of-the-art performance, particularly improving recognition accuracy for ambiguous classes (e.g., “surprise” vs. “fear”) and long-tailed categories. It attains superior Unweighted Average Recall (UAR) and Weighted Average Recall (WAR) compared to existing methods.

0 citationsRead paper

A Comprehensive Benchmark for Electrocardiogram Time-Series

Jul 14, 2025

Existing ECG analysis studies often overlook the electrophysiological characteristics of ECG signals and clinical application requirements, leading to inadequate evaluation frameworks. To address this, we propose ECG-Bench—the first comprehensive, multi-task benchmark for ECG time-series analysis—covering four clinically relevant downstream tasks: rhythm classification, anomaly detection, lesion localization, and risk prediction. We introduce novel evaluation metrics tailored to ECG spectral properties and clinical interpretability. Furthermore, we design ECGFormer, a lightweight, physiology-aware temporal modeling architecture. Leveraging a large-scale pretraining benchmarking framework, systematic evaluations demonstrate that our metrics improve assessment accuracy by +12.3% on average, while ECGFormer achieves a mean F1-score of 94.7% across six mainstream datasets (e.g., MIT-BIH), outperforming state-of-the-art models by 3.1%. This work establishes a standardized evaluation paradigm and delivers a high-performance foundational model for intelligent ECG analysis.

0 citationsRead paper
Recent publications

Latest Papers

Socratic agents for autonomous scientific discovery in high-dimensional physical systems

Jun 25, 2026

This work addresses the limited cognitive autonomy of traditional AI in scientific discovery by introducing AHOIS, a multi-agent AI scientist that incorporates a Socratic questioning mechanism into autonomous physical exploration. By leveraging causal interrogation, counterexample generation, and falsifiability-driven hypothesis refinement, AHOIS autonomously proposes, tests, and revises hypotheses without relying on prior models. The system integrates causal reasoning, constraint verification, sparse measurement optimization, and uncertainty calibration within a closed-loop experimental framework. Deployed on a multimode fiber platform, AHOIS autonomously discovered a stochastic interference encoding scheme, achieving classification accuracies of 76.97% on MNIST and 83.17% on Fashion-MNIST, while effectively diagnosing multiple failure modes and substantially enhancing the consistency and completeness of physical interpretability.

0 citationsRead paper

Beyond Endpoints: Path-Centric Reasoning for Vectorized Off-Road Network Extraction

Dec 11, 2025

To address topological errors (e.g., structural disconnections due to occlusion or ambiguous junctions) in vectorized off-road road extraction caused by domain shift, this paper introduces a path-centric inference paradigm, overcoming the robustness limitations of conventional node-centric approaches. Methodologically, we propose MaGRoad—a mask-aware geodesic road extraction framework integrating multi-scale visual evidence aggregation, geodesic path modeling, and lightweight vector decoding. Our contributions are threefold: (1) We release WildRoad, the first global off-road road dataset, accompanied by an interactive annotation tool; (2) We design MaGRoad to explicitly model continuous road centerlines via geodesic distance priors and mask-guided feature fusion; (3) MaGRoad achieves state-of-the-art performance on WildRoad, demonstrates strong cross-domain generalization to urban road datasets (e.g., DeepGlobe, Cowc), and operates 2.5× faster than prior methods.

0 citationsRead paper

Centering Emotion Hotspots: Multimodal Local-Global Fusion and Cross-Modal Alignment for Emotion Recognition in Conversations

Oct 07, 2025

Emotion Recognition in Conversations (ERC) faces challenges stemming from sparse, localized, and asynchronous multimodal evidence. To address these, we propose a multimodal ERC framework centered on “emotion hotspots”: (1) local emotion-critical segments are first identified within textual, acoustic, and visual modalities; (2) a hotspot-gated fusion mechanism adaptively weights local hotspots against global contextual representations; (3) a routing-based hybrid aligner enables fine-grained cross-modal alignment; and (4) a dialogue-structure graph models inter-utterance dependencies. The method integrates local-global feature modeling, graph neural networks, dynamic attention, and gating mechanisms. Evaluated on standard benchmarks, our approach significantly outperforms strong baselines. Ablation studies confirm the effectiveness of each component. Overall, the framework achieves robust, interpretable, fine-grained emotion recognition by jointly leveraging modality-specific saliency and structured conversational context.

0 citationsRead paper

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition

Jul 16, 2025

Dynamic Facial Expression Recognition (DFER) faces two key challenges: insufficient exploitation of fine-grained affective cues from generated textual descriptions, and difficulty suppressing facial motions irrelevant to emotion. To address these, we propose GRACE—a framework that achieves token-level cross-modal alignment between linguistic cues and visually salient regions via coarse-to-fine emotional text enhancement and motion-difference-weighted attention. GRACE further incorporates dynamic motion modeling, semantic text refinement, and entropy-regularized optimal transport for precise spatiotemporal localization of emotion-relevant features. Evaluated on three benchmark datasets, GRACE achieves state-of-the-art performance, particularly improving recognition accuracy for ambiguous classes (e.g., “surprise” vs. “fear”) and long-tailed categories. It attains superior Unweighted Average Recall (UAR) and Weighted Average Recall (WAR) compared to existing methods.

0 citationsRead paper

A Comprehensive Benchmark for Electrocardiogram Time-Series

Jul 14, 2025

Existing ECG analysis studies often overlook the electrophysiological characteristics of ECG signals and clinical application requirements, leading to inadequate evaluation frameworks. To address this, we propose ECG-Bench—the first comprehensive, multi-task benchmark for ECG time-series analysis—covering four clinically relevant downstream tasks: rhythm classification, anomaly detection, lesion localization, and risk prediction. We introduce novel evaluation metrics tailored to ECG spectral properties and clinical interpretability. Furthermore, we design ECGFormer, a lightweight, physiology-aware temporal modeling architecture. Leveraging a large-scale pretraining benchmarking framework, systematic evaluations demonstrate that our metrics improve assessment accuracy by +12.3% on average, while ECGFormer achieves a mean F1-score of 94.7% across six mainstream datasets (e.g., MIT-BIH), outperforming state-of-the-art models by 3.1%. This work establishes a standardized evaluation paradigm and delivers a high-performance foundational model for intelligent ECG analysis.

0 citationsRead paper