ForceU-VLA: A Force-Aware Vision-Language-Action Model for Embodied Ultrasound Scanning

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of loose force-image coupling and lack of phase awareness in embodied ultrasound scanning by proposing a force-aware vision-language-action model. The method integrates force-ultrasound synergistic fusion with a phase-adaptive modulation mechanism to enable autonomous, precise scanning under multimodal feedback, supported by a newly constructed real-world synchronized dataset comprising 100,000 frames. Experimental results demonstrate that the proposed model significantly improves contact stability and probe pressure regulation while effectively enhancing dynamic interaction capture accuracy, task execution quality, and overall system reliability. These advancements establish a novel paradigm for embodied medical manipulation, bridging critical gaps in sensorimotor integration for robotic ultrasound procedures.
📝 Abstract
Embodied intelligent ultrasound scanning enables the automation and standardization of the ultrasound examination process by integrating perception, decision-making, and execution capabilities. However, existing methods suffer from loosely coupled modeling between force and ultrasound modalities and lack awareness of scanning stages, which limits their ability to capture dynamic probe-tissue interactions. To address these issues, we propose ForceU-VLA, a force-aware Vision-Language-Action model for autonomous embodied ultrasound scanning, which leverages force signals and ultrasound image feedback throughout the scanning process to enable accurate and high-quality ultrasound acquisition. Firstly, we propose a Force-Ultrasound Synergistic Fusion Module (FUSFM) that synergistically fuses ultrasound visual and force-feedback information to provide stable, reliable guidance for probe motion. Secondly, a Stage-Adaptive Modulation Mechanism (SAMM) is proposed to accommodate the task requirements across different scanning stages by adaptively modulating multimodal features to enhance their representation quality. Additionally, we introduce ForceU-VLA-Data, a real-world, force-aware embodied ultrasound dataset that integrates visual, force, and action signals, including data from two organs across five representative clinical scanning views, and comprising 450 expert-collected trajectories with approximately 100,000 synchronized multimodal frames. Extensive experimental results demonstrate that ForceU-VLA significantly improves contact stability and probe pressure regulation in embodied ultrasound scanning, thereby effectively enhancing task execution quality and overall system reliability. The source code is available at https://github.com/VMVLab/ForceU-VLA.
Problem

Research questions and friction points this paper is trying to address.

Embodied Ultrasound Scanning
Force-Ultrasound Coupling
Scanning Stage Awareness
Probe-Tissue Interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Force-Ultrasound Synergistic Fusion
Stage-Adaptive Modulation
Vision-Language-Action Model
Embodied Ultrasound Scanning
Multimodal Dataset
🔎 Similar Papers
No similar papers found.
X
Xingzheng Wu
Faculty of Computer Science and Technology, Ocean University of China
C
Cheng Zhang
Faculty of Computer Science and Technology, Ocean University of China
G
Guihao Yan
Faculty of Computer Science and Technology, Ocean University of China
X
Xifeng Hu
School of Information Science and Engineering, Shandong University
Zhi Liu
Zhi Liu
Johns Hopkins University School of Medicine
Postdoctoral Fellow
Q
Qing Cai
Innovation School of Artificial Intelligence, Hefei University of Technology