Institution profile

Skild AI

Industry researchnorthamerica · us
Official website
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

FlowDPG: Deterministic Policy Gradient on Flow Matching Policies for Real-World Manipulation

Jun 20, 2026

This work addresses the computational burden and numerical instability of flow-matching–based reinforcement learning in real-world robotic manipulation, which stems from its reliance on backpropagation through time (BPTT). To overcome this limitation, the authors propose FlowDPG, a BPTT-free DDPG-style algorithm that distills critic gradients into a velocity field, thereby integrating demonstration-driven motion with critic-guided corrections during policy optimization. Theoretical analysis demonstrates that FlowDPG’s update direction aligns with the classical deterministic policy gradient. Evaluated on a multi-stage dual-arm AirPods assembly task, FlowDPG achieves a 92% end-to-end success rate, substantially outperforming existing approaches based on value conditioning, auxiliary module adaptation, and adjoint gradients.

0 citationsRead paper

HyReach: Vision-Guided Hybrid Manipulator Reaching in Unseen Cluttered Environments

Mar 22, 2026

This work addresses the challenge of high-precision robotic grasping in unstructured, previously unseen cluttered environments, where conventional rigid manipulators struggle due to limited compliance and adaptability. The authors propose a real-time hybrid rigid-soft continuum manipulator system that integrates vision-guided 3D scene reconstruction, shape-aware motion planning, and a learning-based hybrid controller. This approach enables, for the first time, generalization to entirely novel scenes without requiring environment-specific retraining. By synergistically combining the compliance of soft robotics with the precision of rigid mechanisms, the system achieves an average end-effector positioning error of less than 2 cm across diverse real-world cluttered settings, substantially improving task success rates and robustness in open-world manipulation scenarios.

0 citationsRead paper

ViPRA: Video Prediction for Robot Actions

Nov 11, 2025

This paper addresses robot continuous control from unlabeled video demonstrations. We propose a motion-aware latent action representation coupled with a chunked flow-matching decoder. Our method explicitly models *what changes* and *how it changes* in visual dynamics, jointly optimizing a perception loss and optical flow consistency constraint within a pretrain-fine-tune framework to enable cross-morphology generalization. With only ~100 demonstration videos, it achieves smooth, high-frequency (22 Hz) continuous control. On the SIMPLER benchmark, our approach improves performance by 16%; on real-world manipulation tasks, it yields a 13% gain; and it supports cross-platform transfer. Key contributions are: (1) a motion-centric latent action representation that captures spatiotemporal dynamics directly from visual change; and (2) a lightweight, efficient flow-matching decoder explicitly regularized by optical flow constraints to ensure physically plausible action decoding.

0 citationsRead paper

LocoFormer: Generalist Locomotion via Long-context Adaptation

Sep 28, 2025

Existing motion controllers rely heavily on manual parameter tuning and lack cross-morphology generalization. Method: We propose a universal locomotion control framework based on procedural robot architecture and strong domain randomization for large-scale reinforcement learning training, augmented by a long-context Transformer to enable cross-episode dynamics modeling and test-time adaptation—without requiring precise kinematic priors. Contribution/Results: The resulting policy achieves robust walking on unseen legged and wheeled robots. Experiments demonstrate seamless deployment across heterogeneous hardware platforms; the controller maintains stability under severe disturbances—including sudden payload changes, single-motor failure, and post-fall recovery—exhibiting strong generalization and emergent robustness. Notably, no morphology-specific fine-tuning or explicit dynamics modeling is required.

0 citationsRead paper

Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations

Jul 01, 2025

This work addresses the challenge of teaching robots complex manipulation skills—such as pouring, wiping, and stirring—using only AI-generated videos, without physical demonstrations or robot-side training. Methodologically, it leverages text-to-video diffusion models to synthesize action videos, employs vision-language models to automatically filter semantically consistent video segments, extracts object motion trajectories via 6D pose estimation, and maps these trajectories to robot execution in an embodiment-agnostic manner. Its key contribution is the first demonstration of using purely synthetic video as supervision for end-to-end closed-loop transfer from generated visual data to real-world robot control. Experiments show that the approach achieves performance on par with human demonstrations, improves consistently with higher-generation video fidelity, and significantly outperforms baseline methods—including keypoint prediction and dense feature tracking—in both accuracy and generalization.

0 citationsRead paper
Recent publications

Latest Papers

FlowDPG: Deterministic Policy Gradient on Flow Matching Policies for Real-World Manipulation

Jun 20, 2026

This work addresses the computational burden and numerical instability of flow-matching–based reinforcement learning in real-world robotic manipulation, which stems from its reliance on backpropagation through time (BPTT). To overcome this limitation, the authors propose FlowDPG, a BPTT-free DDPG-style algorithm that distills critic gradients into a velocity field, thereby integrating demonstration-driven motion with critic-guided corrections during policy optimization. Theoretical analysis demonstrates that FlowDPG’s update direction aligns with the classical deterministic policy gradient. Evaluated on a multi-stage dual-arm AirPods assembly task, FlowDPG achieves a 92% end-to-end success rate, substantially outperforming existing approaches based on value conditioning, auxiliary module adaptation, and adjoint gradients.

0 citationsRead paper

HyReach: Vision-Guided Hybrid Manipulator Reaching in Unseen Cluttered Environments

Mar 22, 2026

This work addresses the challenge of high-precision robotic grasping in unstructured, previously unseen cluttered environments, where conventional rigid manipulators struggle due to limited compliance and adaptability. The authors propose a real-time hybrid rigid-soft continuum manipulator system that integrates vision-guided 3D scene reconstruction, shape-aware motion planning, and a learning-based hybrid controller. This approach enables, for the first time, generalization to entirely novel scenes without requiring environment-specific retraining. By synergistically combining the compliance of soft robotics with the precision of rigid mechanisms, the system achieves an average end-effector positioning error of less than 2 cm across diverse real-world cluttered settings, substantially improving task success rates and robustness in open-world manipulation scenarios.

0 citationsRead paper

ViPRA: Video Prediction for Robot Actions

Nov 11, 2025

This paper addresses robot continuous control from unlabeled video demonstrations. We propose a motion-aware latent action representation coupled with a chunked flow-matching decoder. Our method explicitly models *what changes* and *how it changes* in visual dynamics, jointly optimizing a perception loss and optical flow consistency constraint within a pretrain-fine-tune framework to enable cross-morphology generalization. With only ~100 demonstration videos, it achieves smooth, high-frequency (22 Hz) continuous control. On the SIMPLER benchmark, our approach improves performance by 16%; on real-world manipulation tasks, it yields a 13% gain; and it supports cross-platform transfer. Key contributions are: (1) a motion-centric latent action representation that captures spatiotemporal dynamics directly from visual change; and (2) a lightweight, efficient flow-matching decoder explicitly regularized by optical flow constraints to ensure physically plausible action decoding.

0 citationsRead paper

LocoFormer: Generalist Locomotion via Long-context Adaptation

Sep 28, 2025

Existing motion controllers rely heavily on manual parameter tuning and lack cross-morphology generalization. Method: We propose a universal locomotion control framework based on procedural robot architecture and strong domain randomization for large-scale reinforcement learning training, augmented by a long-context Transformer to enable cross-episode dynamics modeling and test-time adaptation—without requiring precise kinematic priors. Contribution/Results: The resulting policy achieves robust walking on unseen legged and wheeled robots. Experiments demonstrate seamless deployment across heterogeneous hardware platforms; the controller maintains stability under severe disturbances—including sudden payload changes, single-motor failure, and post-fall recovery—exhibiting strong generalization and emergent robustness. Notably, no morphology-specific fine-tuning or explicit dynamics modeling is required.

0 citationsRead paper

Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations

Jul 01, 2025

This work addresses the challenge of teaching robots complex manipulation skills—such as pouring, wiping, and stirring—using only AI-generated videos, without physical demonstrations or robot-side training. Methodologically, it leverages text-to-video diffusion models to synthesize action videos, employs vision-language models to automatically filter semantically consistent video segments, extracts object motion trajectories via 6D pose estimation, and maps these trajectories to robot execution in an embodiment-agnostic manner. Its key contribution is the first demonstration of using purely synthetic video as supervision for end-to-end closed-loop transfer from generated visual data to real-world robot control. Experiments show that the approach achieves performance on par with human demonstrations, improves consistently with higher-generation video fidelity, and significantly outperforms baseline methods—including keypoint prediction and dense feature tracking—in both accuracy and generalization.

0 citationsRead paper