FoldNet++: a Large-Scale Synthetic Dataset for Robotic T-Shirt Folding and Unfolding
为解决机器人T恤折叠和展开难题,本文构建了一个大规模合成数据集FoldNet++,并通过基于规则的框架生成操作演示来训练视觉运动策略。
为解决机器人T恤折叠和展开难题,本文构建了一个大规模合成数据集FoldNet++,并通过基于规则的框架生成操作演示来训练视觉运动策略。
研究通过UniPart模型解决了零样本语言驱动的3D部分分割问题,利用CLIP文本嵌入和大规模标注数据集LangPart-1M进行训练。
本文提出RoboGesture框架,通过设计数据、建模和控制来解决人形机器人同步生成语义手势的问题,采用层次语义-声学对齐器和流式条件运动生成器等方法。
Existing evaluations of humanoid motion tracking rely primarily on kinematic errors, which fail to capture physically implausible distortions perceptible to humans—such as foot sliding or incorrect contact—and suffer from small-scale, low-diversity test sets. To address these limitations, this work introduces HumanTracker, a large-scale and diverse benchmark comprising 153 hours of optical motion capture data from professional actors, covering four action categories with accompanying textual annotations. Furthermore, the authors propose HumanScore, a human-aligned evaluation metric derived from a preference model trained on 12K motion pairs (24K individual motions). HumanScore enables fine-grained diagnosis of critical physical properties like contact fidelity and support stability, significantly outperforming conventional metrics across multiple state-of-the-art trackers, accurately predicting human preferences, and uncovering previously overlooked physical inconsistencies.
Humanoid robot control faces significant challenges in whole-body coordination, real-time responsiveness, and cross-scenario generalization, with existing approaches limited in scalability and universality. This work proposes a scalable behavioral foundation model that achieves breakthrough performance through three core innovations: a global-frame-based motion tracking learning paradigm, a co-designed strategy involving the number of policy rollouts and diversity of reference motions, and a novel Humanoid Transformer architecture. The study systematically uncovers a scalable pathway for behavioral foundation models, enabling structured behavioral representations to emerge naturally. Evaluated in both simulation and real-world deployment, the approach substantially improves performance, reducing MPKPE by over 10% on local motion patterns and by 82% on global motion patterns in the test set.
为解决机器人T恤折叠和展开难题,本文构建了一个大规模合成数据集FoldNet++,并通过基于规则的框架生成操作演示来训练视觉运动策略。
研究通过UniPart模型解决了零样本语言驱动的3D部分分割问题,利用CLIP文本嵌入和大规模标注数据集LangPart-1M进行训练。
本文提出RoboGesture框架,通过设计数据、建模和控制来解决人形机器人同步生成语义手势的问题,采用层次语义-声学对齐器和流式条件运动生成器等方法。
Existing evaluations of humanoid motion tracking rely primarily on kinematic errors, which fail to capture physically implausible distortions perceptible to humans—such as foot sliding or incorrect contact—and suffer from small-scale, low-diversity test sets. To address these limitations, this work introduces HumanTracker, a large-scale and diverse benchmark comprising 153 hours of optical motion capture data from professional actors, covering four action categories with accompanying textual annotations. Furthermore, the authors propose HumanScore, a human-aligned evaluation metric derived from a preference model trained on 12K motion pairs (24K individual motions). HumanScore enables fine-grained diagnosis of critical physical properties like contact fidelity and support stability, significantly outperforming conventional metrics across multiple state-of-the-art trackers, accurately predicting human preferences, and uncovering previously overlooked physical inconsistencies.
Humanoid robot control faces significant challenges in whole-body coordination, real-time responsiveness, and cross-scenario generalization, with existing approaches limited in scalability and universality. This work proposes a scalable behavioral foundation model that achieves breakthrough performance through three core innovations: a global-frame-based motion tracking learning paradigm, a co-designed strategy involving the number of policy rollouts and diversity of reference motions, and a novel Humanoid Transformer architecture. The study systematically uncovers a scalable pathway for behavioral foundation models, enabling structured behavioral representations to emerge naturally. Evaluated in both simulation and real-world deployment, the approach substantially improves performance, reducing MPKPE by over 10% on local motion patterns and by 82% on global motion patterns in the test set.