DeepInsight II: One Trace from Benchmark to Robot

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the fragmentation in embodied intelligence evaluation and the lack of empirical continuity in sim-to-real deployment by proposing a native sim-to-real tracking reduction framework alongside a five-level handover diagnostic system. By integrating MotionBench unified metrics, cross-domain identity binding, and System2-1-0 combinatorial attribution techniques, we establish a shared tracking mechanism between simulation and physical platforms. The research reproduces multi-benchmark checkpoints and achieves trajectory alignment, effectively validating fault attribution and remediation capabilities under hardware constraints. Consequently, this work establishes a complete empirical closed loop from benchmark execution to physical verification, significantly enhancing both the interpretability and deployment reliability of embodied intelligence systems.
📝 Abstract
Across a Physical AI stack, evaluation maturity is inversely aligned with deployment risk: foundation models enjoy mature, standardized harnesses, while the embodied layers on which deployment actually turns remain fragmented across benchmark-specific simulators, embodiments, and interfaces. The first DeepInsight report (v1) unified evaluation across this stack behind three abstractions---task, resource, and result---but its quantitative evidence centered on the foundation-model layer; navigation and manipulation (System 1) and whole-body control (System 0) remained simulation case studies, and physical execution was outside its empirical scope. DeepInsight II keeps that substrate fixed and quantifies the embodied half. First, it reproduces released-checkpoint references across two navigation and four manipulation benchmarks under their native protocols. Second, MotionBench places four released whole-body controllers under one workload and metric contract, then carries a qualified within-family cohort from parallel simulation to matched real-robot trials in which simulated and physical rollouts share a parent trace identity while retaining execution-domain-specific records, making the sim-to-real gap a native reduction rather than a reconciliation across toolchains. Third, a composed System 2--1--0 study extends trace localization into five evidence-grounded handoff labels, each mapped to a concrete repair action, with a measured repairability criterion and physical episodes testing the same attribution under hardware-observable state. The contribution is therefore not a new evaluation architecture, but empirical continuity from benchmark execution to matched robot evidence and repair-oriented diagnosis.
Problem

Research questions and friction points this paper is trying to address.

Physical AI evaluation
Embodied AI
Sim-to-real gap
Benchmark fragmentation
Robot deployment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sim-to-Real Trace Identity
MotionBench
Repair-Oriented Diagnosis
System 2-1-0 Handoff Labels
Empirical Continuity
🔎 Similar Papers
No similar papers found.
S
Siyi Li
XPENG Robotics
Y
Yuchen Kang
XPENG Robotics
W
Wuliang Wang
XPENG Robotics
Z
Zhengjie Zhang
XPENG Robotics
J
Jiangpin Liu
XPENG Robotics
J
Jianhao Yao
XPENG Robotics
J
Jie Chen
XPENG Robotics