Institution profile

Agile Robots AG

Industry researcheurope · ch
Official website
Research library13linked papers
Opportunities0open roles
Selected work

Representative Papers

TacUMI: A Multi-Modal Universal Manipulation Interface for Contact-Rich Tasks

Jan 21, 2026

This work addresses the challenge of accurately identifying semantic event boundaries in long-horizon manipulation tasks rich in physical contact, where reliance solely on visual and proprioceptive cues proves insufficient for effective task segmentation. To overcome this limitation, the authors propose TacUMI—a compact, multimodal data acquisition system that integrates ViTac visuo-tactile sensing with force-torque and pose perception—and, for the first time, embed it within a general-purpose manipulation interface to enable highly synchronized multimodal recording. Building upon this hardware foundation, they further introduce a temporal modeling–based multimodal fusion framework to automatically extract event boundaries from human demonstrations. Evaluated on a cable assembly task, the method achieves over 90% segmentation accuracy, significantly outperforming unimodal baselines and demonstrating the critical role of multimodal perception in enhancing task decomposition performance.

1 citationsRead paper

From Flow to One Step: Real-Time Multi-Modal Trajectory Policies via Implicit Maximum Likelihood Estimation-based Distribution Distillation

Mar 10, 2026

This work addresses the high latency of existing diffusion- or flow-matching-based generative strategies, which rely on iterative ODE solvers and are thus ill-suited for high-frequency closed-loop control, while single-step acceleration methods often suffer from distribution collapse and loss of multimodal behavior. To overcome these limitations, we propose a distribution distillation framework based on Implicit Maximum Likelihood Estimation (IMLE), which distills a conditional flow matching (CFM) teacher model into a single-step student model. By introducing a bidirectional Chamfer distance to optimize set-level objectives, our approach ensures both coverage and fidelity of multimodal action distributions in a single forward pass. Integrated with a geometric-aware encoder that fuses multimodal perception (RGB, depth, point clouds, and proprioception), the method enables real-time replanning at high control frequencies and demonstrates enhanced robustness under dynamic perturbations, effectively mitigating distribution collapse in single-step generation.

0 citationsRead paper

OmniDexVLG: Learning Dexterous Grasp Generation from Vision Language Model-Guided Grasp Semantics, Taxonomy and Functional Affordance

Dec 03, 2025

Existing dexterous grasp generation methods struggle to jointly model grasp categorization, contact semantics, and functional affordances, resulting in weak semantic controllability and poor human interpretability. This paper proposes a multimodal semantic-aware framework that, for the first time, jointly embeds these three semantic dimensions into a vision-language model, enabling fine-grained, natural-language-instruction-driven grasp control. We introduce multi-agent collaborative reasoning, retrieval-augmented generation, and chain-of-thought prompting, integrated with physics-based optimization and category-aware differential force-closure sampling to ensure pose feasibility and diversity. Evaluated in both simulation and real-robot settings, our method significantly outperforms state-of-the-art approaches, improving semantic consistency (+28.6%), contact structural richness (+34.1%), and functional affordance coverage (+41.3%).

0 citationsRead paper

ResponsibleRobotBench: Benchmarking Responsible Robot Manipulation using Multi-modal Large Language Models

Dec 03, 2025

This work addresses the lack of accountability and reliability evaluation for multimodal large language models (MLLMs) in high-risk robotic manipulation. We introduce the first benchmark dedicated to responsible robotic manipulation, comprising 23 multi-stage tasks spanning electrical, chemical, and personal safety-critical scenarios, with validated sim-to-real transferability. Methodologically, we integrate MLLMs with visual perception, in-context learning, hazard detection, and multi-representation action execution within a reproducible framework, supported by a novel multimodal dataset. Our approach uniquely unifies risk-aware reasoning, moral inference, and physics-grounded planning, and introduces new evaluation metrics—including “safety success rate”—to quantify responsible behavior. Experimental results demonstrate that the benchmark effectively discriminates among agents across safety compliance, robustness to environmental perturbations, and generalization to unseen tasks and hazards.

0 citationsRead paper
Recent publications

Latest Papers

From Flow to One Step: Real-Time Multi-Modal Trajectory Policies via Implicit Maximum Likelihood Estimation-based Distribution Distillation

Mar 10, 2026

This work addresses the high latency of existing diffusion- or flow-matching-based generative strategies, which rely on iterative ODE solvers and are thus ill-suited for high-frequency closed-loop control, while single-step acceleration methods often suffer from distribution collapse and loss of multimodal behavior. To overcome these limitations, we propose a distribution distillation framework based on Implicit Maximum Likelihood Estimation (IMLE), which distills a conditional flow matching (CFM) teacher model into a single-step student model. By introducing a bidirectional Chamfer distance to optimize set-level objectives, our approach ensures both coverage and fidelity of multimodal action distributions in a single forward pass. Integrated with a geometric-aware encoder that fuses multimodal perception (RGB, depth, point clouds, and proprioception), the method enables real-time replanning at high control frequencies and demonstrates enhanced robustness under dynamic perturbations, effectively mitigating distribution collapse in single-step generation.

0 citationsRead paper

TacUMI: A Multi-Modal Universal Manipulation Interface for Contact-Rich Tasks

Jan 21, 2026

This work addresses the challenge of accurately identifying semantic event boundaries in long-horizon manipulation tasks rich in physical contact, where reliance solely on visual and proprioceptive cues proves insufficient for effective task segmentation. To overcome this limitation, the authors propose TacUMI—a compact, multimodal data acquisition system that integrates ViTac visuo-tactile sensing with force-torque and pose perception—and, for the first time, embed it within a general-purpose manipulation interface to enable highly synchronized multimodal recording. Building upon this hardware foundation, they further introduce a temporal modeling–based multimodal fusion framework to automatically extract event boundaries from human demonstrations. Evaluated on a cable assembly task, the method achieves over 90% segmentation accuracy, significantly outperforming unimodal baselines and demonstrating the critical role of multimodal perception in enhancing task decomposition performance.

1 citationsRead paper

OmniDexVLG: Learning Dexterous Grasp Generation from Vision Language Model-Guided Grasp Semantics, Taxonomy and Functional Affordance

Dec 03, 2025

Existing dexterous grasp generation methods struggle to jointly model grasp categorization, contact semantics, and functional affordances, resulting in weak semantic controllability and poor human interpretability. This paper proposes a multimodal semantic-aware framework that, for the first time, jointly embeds these three semantic dimensions into a vision-language model, enabling fine-grained, natural-language-instruction-driven grasp control. We introduce multi-agent collaborative reasoning, retrieval-augmented generation, and chain-of-thought prompting, integrated with physics-based optimization and category-aware differential force-closure sampling to ensure pose feasibility and diversity. Evaluated in both simulation and real-robot settings, our method significantly outperforms state-of-the-art approaches, improving semantic consistency (+28.6%), contact structural richness (+34.1%), and functional affordance coverage (+41.3%).

0 citationsRead paper

ResponsibleRobotBench: Benchmarking Responsible Robot Manipulation using Multi-modal Large Language Models

Dec 03, 2025

This work addresses the lack of accountability and reliability evaluation for multimodal large language models (MLLMs) in high-risk robotic manipulation. We introduce the first benchmark dedicated to responsible robotic manipulation, comprising 23 multi-stage tasks spanning electrical, chemical, and personal safety-critical scenarios, with validated sim-to-real transferability. Methodologically, we integrate MLLMs with visual perception, in-context learning, hazard detection, and multi-representation action execution within a reproducible framework, supported by a novel multimodal dataset. Our approach uniquely unifies risk-aware reasoning, moral inference, and physics-grounded planning, and introduces new evaluation metrics—including “safety success rate”—to quantify responsible behavior. Experimental results demonstrate that the benchmark effectively discriminates among agents across safety compliance, robustness to environmental perturbations, and generalization to unseen tasks and hazards.

0 citationsRead paper