🤖 AI Summary
This study addresses the confounding effect of scenario discrepancies in cross-domain evaluation of autonomous driving Vision-Language-Action (VLA) models by proposing Syn2Sim2Phy, an event-matching framework. By constructing consistent safety-critical interaction event chains and introducing a novel event-specification anchoring mechanism, this approach eliminates domain assumption biases to enable reproducible benchmarking across synthetic, simulation, and physical domains. Integrating semantic slot mapping with multidimensional scoring, we quantify performance variations in Cut-in and VRU scenarios, revealing that optimal evaluation domains shift dynamically with scene context. Notably, Alpamayo-R1 achieves the highest score of 0.405, establishing an evidence-compliant assessment baseline for VLA behaviors.
📝 Abstract
Vision-language-action (VLA) models for autonomous driving jointly produce scene interpretation, language-based reasoning, and driving trajectories. Existing evaluations often use independently selected synthetic, simulated, and physical data, so measured performance gaps can be confounded by changes in scenario content rather than genuine domain sensitivity. We propose SSP (Synthetic-Simulation-Physical), an event-matched Syn2Sim2Phy evaluation framework that anchors cross-domain comparison to the same safety-critical interaction. Starting from a synthetic long-tail video, SSP builds a validated event specification that preserves road topology, participant roles, relative motion, conflict evolution, passing order, response constraints, and event phases. Platform-specific realizations are then constructed in CARLA and on a closed proving ground and are evaluated only after transfer audits confirm preservation of mandatory event properties. SSP maps heterogeneous outputs from OpenEMMA, LLaViDA, and Alpamayo-R1 into common semantic slots and a 1 s trajectory window to assess output validity, semantic accuracy, critical-interaction recognition, trajectory quality, and risk response. Across Cut-in and vulnerable-road-user crossing cases, the macro-averaged Integrated VLA Capability Scores are 0.259, 0.291, and 0.325 in the Synthetic, Simulation, and Physical domains, respectively, while the best domain varies by scenario. Alpamayo-R1, OpenEMMA, and LLaViDA obtain scores of 0.405, 0.338, and 0.131. SSP provides a reproducible scene-transfer chain and an evidence-qualified evaluation of VLA behavior without assuming that the Physical domain is universally superior.