SMASH: Mastering Scalable Whole-Body Skills for Humanoid Ping-Pong with Egocentric Vision

📅 2026-04-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work proposes a modular system for humanoid table tennis that operates solely on onboard first-person vision, eliminating reliance on external perception. By integrating low-latency visual perception, a generative action prior, and a scalable whole-body motion control framework, the approach achieves continuous and precise ball returns without external cameras for the first time. The system enables diverse striking maneuvers—including powerful smashes and low crouching shots—and demonstrates robust upper-lower body coordination during high-speed rallies in real-world settings. This significantly enhances the autonomy and dynamic motor capabilities of humanoids in fast-paced, interactive tasks.
📝 Abstract
Existing humanoid table tennis systems remain limited by their reliance on external sensing and their inability to achieve agile whole-body coordination for precise task execution. These limitations stem from two core challenges: achieving low-latency and robust onboard egocentric perception under fast robot motion, and obtaining sufficiently diverse task-aligned strike motions for learning precise yet natural whole-body behaviors. In this work, we present \methodname, a modular system for agile humanoid table tennis that unifies scalable whole-body skill learning with onboard egocentric perception, eliminating the need for external cameras during deployment. Our work advances prior humanoid table-tennis systems in three key aspects. First, we achieve agile and precise ball interaction with tightly coordinated whole-body control, rather than relying on decoupled upper- and lower-body behaviors. This enables the system to exhibit diverse strike motions, including explosive whole-body smashes and low crouching shots. Second, by augmenting and diversifying strike motions with a generative model, our framework benefits from scalable motion priors and produces natural, robust striking behaviors across a wide workspace. Third, to the best of our knowledge, we demonstrate the first humanoid table-tennis system capable of consecutive strikes using onboard sensing alone, despite the challenges of low-latency perception, ego-motion-induced instability, and limited field of view. Extensive real-world experiments demonstrate stable and precise ball exchanges under high-speed conditions, validating scalable, perception-driven whole-body skill learning for dynamic humanoid interaction tasks.
Problem

Research questions and friction points this paper is trying to address.

humanoid robotics
egocentric vision
whole-body coordination
table tennis
onboard perception
Innovation

Methods, ideas, or system contributions that make the work stand out.

egocentric vision
whole-body control
generative motion prior
onboard perception
humanoid table tennis
🔎 Similar Papers
No similar papers found.
J
Junli Ren
The University of Hong Kong Kinetix AI
Y
Yinghui Li
The University of Hong Kong Kinetix AI
K
Kai Zhang
The University of Hong Kong Kinetix AI
P
Penglin Fu
The University of Hong Kong Kinetix AI
H
Haoran Jiang
The University of Hong Kong Kinetix AI
Y
Yixuan Pan
The University of Hong Kong Kinetix AI
G
Guangjun Zeng
The University of Hong Kong Kinetix AI
Tao Huang
Tao Huang
Assistant Professor of Law, City University of Hong Kong
Constitutional LawLaw and Technology
W
Weizhong Guo
The University of Hong Kong Kinetix AI
P
Peng Lu
The University of Hong Kong Kinetix AI
T
Tianyu Li
The University of Hong Kong Kinetix AI
J
Jingbo Wang
The University of Hong Kong Kinetix AI
Li Chen
Li Chen
Professor, Hong Kong Baptist University
Conversational AIExplainable AIRecommender SystemsHuman-Computer Interaction
Hongyang Li
Hongyang Li
Assistant Professor, University of Hong Kong
Computer VisionAutonomous DrivingRobotics
Ping Luo
Ping Luo
National University of Defense Technology
distributed_computing