Institution profile

State Key Laboratory of Multimedia Information Processing

Academic institutionasia · cn
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Teaching Audio Models to Reason: A Unified Framework for Source- and Layer-wise Distillation

Sep 22, 2025

Large audio-language models exhibit limited performance on complex reasoning tasks, primarily due to the audio–text modality gap and the absence of structured intermediate supervision. To address this, we propose a unified knowledge distillation framework that transfers symbolic reasoning capabilities from a large text-based teacher model to an audio-based student model while preserving its acoustic understanding. Our approach introduces dual-dimensional distillation—across source modalities (text and audio teachers) and across hierarchical model layers—enabling fine-grained, layer-aligned knowledge transfer. Crucially, we incorporate structured intermediate supervision signals to bridge semantic discrepancies between acoustic representations and symbolic reasoning. Experiments demonstrate substantial improvements in multi-step reasoning performance for audio models, achieving state-of-the-art results across multiple benchmarks and effectively narrowing the semantic gap between speech representation and symbolic reasoning.

0 citationsRead paper

RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot

Jul 05, 2025

Significant visual and kinematic discrepancies between human hand demonstrations and robotic manipulation, coupled with reliance on specialized teleoperation hardware, hinder scalable and cost-effective demonstration data collection. Method: We propose an end-to-end gesture-to-gripper motion generation framework. Using a wrist-mounted GoPro fisheye camera, we capture first-person gesture videos and construct a paired human-hand–robot SE(3) action dataset. A spatiotemporally aligned generative model directly maps hand keypoint sequences to robot gripper trajectories—without requiring physical robots during demonstration recording. Contribution/Results: This work introduces the first calibration-free, teleoperation-free cross-modal motion generation method based solely on monocular fisheye video. Experiments show that policies trained on generated demonstrations achieve performance comparable to those trained on ground-truth demonstrations across diverse dexterous manipulation tasks. Data acquisition efficiency improves by over 5×, substantially enhancing the practicality and scalability of robotic imitation learning.

0 citationsRead paper
Recent publications

Latest Papers

Teaching Audio Models to Reason: A Unified Framework for Source- and Layer-wise Distillation

Sep 22, 2025

Large audio-language models exhibit limited performance on complex reasoning tasks, primarily due to the audio–text modality gap and the absence of structured intermediate supervision. To address this, we propose a unified knowledge distillation framework that transfers symbolic reasoning capabilities from a large text-based teacher model to an audio-based student model while preserving its acoustic understanding. Our approach introduces dual-dimensional distillation—across source modalities (text and audio teachers) and across hierarchical model layers—enabling fine-grained, layer-aligned knowledge transfer. Crucially, we incorporate structured intermediate supervision signals to bridge semantic discrepancies between acoustic representations and symbolic reasoning. Experiments demonstrate substantial improvements in multi-step reasoning performance for audio models, achieving state-of-the-art results across multiple benchmarks and effectively narrowing the semantic gap between speech representation and symbolic reasoning.

0 citationsRead paper

RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot

Jul 05, 2025

Significant visual and kinematic discrepancies between human hand demonstrations and robotic manipulation, coupled with reliance on specialized teleoperation hardware, hinder scalable and cost-effective demonstration data collection. Method: We propose an end-to-end gesture-to-gripper motion generation framework. Using a wrist-mounted GoPro fisheye camera, we capture first-person gesture videos and construct a paired human-hand–robot SE(3) action dataset. A spatiotemporally aligned generative model directly maps hand keypoint sequences to robot gripper trajectories—without requiring physical robots during demonstration recording. Contribution/Results: This work introduces the first calibration-free, teleoperation-free cross-modal motion generation method based solely on monocular fisheye video. Experiments show that policies trained on generated demonstrations achieve performance comparable to those trained on ground-truth demonstrations across diverse dexterous manipulation tasks. Data acquisition efficiency improves by over 5×, substantially enhancing the practicality and scalability of robotic imitation learning.

0 citationsRead paper