🤖 AI Summary
This study addresses the absence of tactile feedback and data acquisition challenges in first-person video by proposing EgoTac, a generalizable model for precise visual-to-tactile inference. Leveraging joint training and zero-shot transfer learning on 5.7 million image-tactile pairs, this approach effectively overcomes cross-domain generalization bottlenecks. Experimental results demonstrate that EgoTac achieves an in-domain force estimation error below 0.06 N while outperforming state-of-the-art methods in out-of-domain contact prediction. By validating the feasibility of inferring rich tactile signals from purely visual inputs, this work establishes a novel paradigm for embodied intelligence perception.
📝 Abstract
Touch is fundamental to dexterous manipulation, yet most egocentric human data increasingly used for robot learning lacks tactile information. Directly collecting large-scale tactile data is challenging due to sensor limitations, while human video data is abundant, contact-rich, and easily scalable. This motivates a natural question: can tactile signals be inferred purely from vision? To address this, we introduce EgoTac, a generalizable model that predicts rich tactile information directly from egocentric human videos. EgoTac is trained on a unified corpus of over 5.7M image-tactile pairs, covering both continuous force measurements and binary contacts. By learning from this diverse dataset, EgoTac captures nuanced touch dynamics across varied interactions. Experiments demonstrate strong performance: in-domain prediction achieves an average force error below 0.06N. On out-of-domain contact prediction benchmarks, EgoTac consistently outperforms the state-of-the-art contact estimator. It also captures the rise and fall patterns of real tactile data and enables zero-shot predictions on unconstrained real-world videos. Scaling analyses further reveal that both data diversity and volume improve performance steadily. Overall, EgoTac provides a scalable pathway to extract tactile priors from egocentric human videos, enabling broadly applicable tactile-aware robot learning.