Institution profile

Central South University

Academic institutionasia · cn
Official website
Research library583linked papers
Opportunities0open roles
Selected work

Representative Papers

FocusTrack: One-Stage Focus-and-Suppress Framework for 3D Point Cloud Object Tracking

Oct 27, 2025ACM Multimedia

This work addresses the limitations of existing two-stage 3D point cloud object tracking methods, which rely on explicit foreground segmentation and consequently suffer from error accumulation and computational bottlenecks. To overcome these issues, we propose the first end-to-end single-stage tracking framework that jointly models motion and semantics without explicit segmentation, enabling both efficiency and accuracy. The core innovation lies in a focus-suppression attention mechanism, integrated with a temporal difference Siamese encoder to model inter-frame motion dynamics, thereby adaptively enhancing foreground features while suppressing background noise. Extensive experiments demonstrate that our method achieves state-of-the-art performance on major benchmarks—including KITTI, nuScenes, and Waymo—while running at an impressive inference speed of 105 FPS.

6 citationsRead paper

A Convoy of Magnetic Millirobots Transports Endoscopic Instruments for Minimally-Invasive Surgery.

Jul 01, 2024Advancement of science

In minimally invasive surgery, microrobots suffer from insufficient traction on slippery, soft-tissue surfaces, hindering reliable transport of elongated instruments (e.g., endoscopes, catheters). To address this, we present TrainBot—a magnetically actuated millirobotic convoy system—where multiple millirobots cooperatively form a “train-like” configuration to enable stable, heavy-load instrument transport within narrow anatomical lumens (e.g., bile ducts, intestines). Key contributions include: (i) the first demonstration of millirobotic swarm-based cargo transport, achieving a twofold increase in output force; (ii) bioinspired, biocompatible microstructured feet that enhance individual propulsion force by threefold; and (iii) the world’s first millirobot-assisted electrodilatation procedure for biliary stricture relief. Integrated with wireless permanent-magnet actuation and multi-robot closed-loop control, TrainBot successfully validated biliary obstruction clearance, drainage tunnel creation, and targeted drug delivery in human-scale organ phantoms—significantly advancing precision instrument delivery in minimally invasive interventions.

2 citationsRead paper

Large Language Model(LLM) assisted End-to-End Network Health Management based on Multi-Scale Semanticization

Jun 12, 2024arXiv.org

To address the challenge of real-time, accurate, and interpretable health-state diagnosis for devices in Dynamic Heterogeneous Networks (DHNs), this paper proposes an end-to-end intelligent operations and maintenance framework. The frontend introduces a Multi-Scale Semantic Anomaly Detection Model (MSADM), integrating semantic rule trees with attention mechanisms to enable fine-grained perception across heterogeneous network entities. The backend employs a Chain-of-Thought (CoT) large language model to autonomously generate fault root-cause analyses and optimization strategies, thereby closing the detection–analysis–decision loop. This work establishes the first multi-scale semantic anomaly detection paradigm for DHNs and pioneers the deep integration of CoT-based LLMs into network fault diagnosis pipelines. Experimental results demonstrate that MSADM achieves 91.31% accuracy in heterogeneous network anomaly detection—significantly outperforming existing distributed approaches—while supporting real-time, adaptive, and interpretable network health management.

2 citationsRead paper

MIND: Benchmarking Memory Consistency and Action Control in World Models

Feb 08, 2026

Existing world models lack a unified open-domain closed-loop benchmark, making it difficult to systematically evaluate their memory consistency and action control capabilities. To address this gap, this work proposes MIND, a benchmark comprising 250 high-resolution (1080p/24 FPS) multi-view synchronized video sequences that span a diverse action space—including variations in movement speed and camera rotation—and introduces a closed-loop interactive evaluation framework. Additionally, we present MIND-World, the first Video-to-World baseline method designed for open-domain scenarios. Experimental results demonstrate that current models still face significant challenges in long-term memory stability and generalization across actions, highlighting MIND as a reliable platform for future research in world modeling.

1 citationsRead paper

ConLA: Contrastive Latent Action Learning from Human Videos for Robotic Manipulation

Jan 31, 2026

This work addresses the challenge of learning transferable robotic manipulation policies from human demonstration videos without explicit action labels, while avoiding shortcut learning and representation entanglement caused by reconstructing visual appearance. The authors propose an unsupervised pretraining framework that leverages a contrastive disentanglement mechanism, integrating action category priors with temporal dynamics to effectively separate motion semantics from visual content. This approach yields clean, semantically consistent latent action representations. Notably, it is the first method to surpass the performance of models pretrained on real robot trajectories when using only human videos for pretraining. Extensive experiments demonstrate its strong generalization and practical utility across multiple robotic manipulation benchmarks, significantly enhancing both the disentanglement and transferability of learned action representations.

1 citationsRead paper
Recent publications

Latest Papers