Institution profile

State Key Laboratory of General Artificial Intelligence

Academic institutionasia · cn
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Reconstructing 12-Lead ECG from 3-Lead ECG using Variational Autoencoder to Improve Cardiac Disease Detection of Wearable ECG Devices

Oct 13, 2025

Portable 3-lead wearable electrocardiograms (ECGs) lack the diagnostic fidelity of clinical 12-lead ECGs, limiting their utility in cardiac screening. Method: We propose WearECG—a novel variational autoencoder (VAE) explicitly modeling spatiotemporal dependencies in ECG signals—to synthesize high-fidelity 12-lead ECGs from 3-lead inputs. The model employs a multi-objective loss combining mean squared error (MSE), mean absolute error (MAE), and Fréchet Inception Distance (FID), and is fine-tuned on ECGFounder for multi-label disease classification. Results: Evaluated on MIMIC-IV, reconstructed ECGs demonstrate physiological plausibility and clinical diagnostic validity, as confirmed by expert blind assessment. Downstream detection of myocardial infarction and other conditions achieves performance comparable to ground-truth 12-lead ECGs. This work pioneers the tight integration of generative modeling with rigorous clinical validation, establishing a scalable, low-cost paradigm for population-level cardiac screening.

0 citationsRead paper

Beyond Isolated Facts: Synthesizing Narrative and Grounded Supervision for VideoQA

Sep 29, 2025

Existing VideoQA models rely on shallow supervision signals from isolated question-answer pairs, limiting their ability to model the narrative logic and causal structure of video events. To address this, we propose a question-driven narrativized supervision paradigm: leveraging Question-Based Paraphrasing (QBP) and Question-Based Captioning (QBC), we reconstruct discrete QA pairs into coherent narrative paragraphs grounded in fine-grained visual evidence. The resulting narratives are trained end-to-end within a unified next-token prediction framework. This approach elevates video understanding supervision from a “collection of facts” to a “structured narrative” for the first time, substantially enhancing models’ capacity to capture deep event semantics. Our method achieves new state-of-the-art results on STAR and NExT-QA: a 3B-parameter model improves accuracy on STAR by 4.9 points to 72.5%, while a 7B model attains 80.8% on NExT-QA. It also demonstrates improved cross-dataset generalization and faster training convergence.

0 citationsRead paper

AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical Reasoning

Aug 11, 2025

Visual robotic manipulation (VRM) suffers from scarce robot interaction data and high costs of multimodal annotation. Existing vision-language pretraining approaches either rely on non-task-specific web data or employ implicit modeling (e.g., frame prediction), resulting in poor generalization under few-shot settings. To address this, we propose an analogy-based cross-modal action transfer framework. Our method explicitly extracts action knowledge from human hand keypoints—establishing a structured analogical mapping between human motion and robot actuator dynamics for the first time. It integrates keypoint-driven vision-language pretraining, human action video retrieval, historical observation alignment, and an analogy reasoning network. Evaluated on the CALVIN benchmark and real-robot experiments, our approach significantly outperforms state-of-the-art methods in few-shot scenarios, demonstrating that human motion priors effectively enhance robotic generalization across tasks and environments.

0 citationsRead paper
Recent publications

Latest Papers

Reconstructing 12-Lead ECG from 3-Lead ECG using Variational Autoencoder to Improve Cardiac Disease Detection of Wearable ECG Devices

Oct 13, 2025

Portable 3-lead wearable electrocardiograms (ECGs) lack the diagnostic fidelity of clinical 12-lead ECGs, limiting their utility in cardiac screening. Method: We propose WearECG—a novel variational autoencoder (VAE) explicitly modeling spatiotemporal dependencies in ECG signals—to synthesize high-fidelity 12-lead ECGs from 3-lead inputs. The model employs a multi-objective loss combining mean squared error (MSE), mean absolute error (MAE), and Fréchet Inception Distance (FID), and is fine-tuned on ECGFounder for multi-label disease classification. Results: Evaluated on MIMIC-IV, reconstructed ECGs demonstrate physiological plausibility and clinical diagnostic validity, as confirmed by expert blind assessment. Downstream detection of myocardial infarction and other conditions achieves performance comparable to ground-truth 12-lead ECGs. This work pioneers the tight integration of generative modeling with rigorous clinical validation, establishing a scalable, low-cost paradigm for population-level cardiac screening.

0 citationsRead paper

Beyond Isolated Facts: Synthesizing Narrative and Grounded Supervision for VideoQA

Sep 29, 2025

Existing VideoQA models rely on shallow supervision signals from isolated question-answer pairs, limiting their ability to model the narrative logic and causal structure of video events. To address this, we propose a question-driven narrativized supervision paradigm: leveraging Question-Based Paraphrasing (QBP) and Question-Based Captioning (QBC), we reconstruct discrete QA pairs into coherent narrative paragraphs grounded in fine-grained visual evidence. The resulting narratives are trained end-to-end within a unified next-token prediction framework. This approach elevates video understanding supervision from a “collection of facts” to a “structured narrative” for the first time, substantially enhancing models’ capacity to capture deep event semantics. Our method achieves new state-of-the-art results on STAR and NExT-QA: a 3B-parameter model improves accuracy on STAR by 4.9 points to 72.5%, while a 7B model attains 80.8% on NExT-QA. It also demonstrates improved cross-dataset generalization and faster training convergence.

0 citationsRead paper

AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical Reasoning

Aug 11, 2025

Visual robotic manipulation (VRM) suffers from scarce robot interaction data and high costs of multimodal annotation. Existing vision-language pretraining approaches either rely on non-task-specific web data or employ implicit modeling (e.g., frame prediction), resulting in poor generalization under few-shot settings. To address this, we propose an analogy-based cross-modal action transfer framework. Our method explicitly extracts action knowledge from human hand keypoints—establishing a structured analogical mapping between human motion and robot actuator dynamics for the first time. It integrates keypoint-driven vision-language pretraining, human action video retrieval, historical observation alignment, and an analogy reasoning network. Evaluated on the CALVIN benchmark and real-robot experiments, our approach significantly outperforms state-of-the-art methods in few-shot scenarios, demonstrating that human motion priors effectively enhance robotic generalization across tasks and environments.

0 citationsRead paper