Auditing Instruction-Trajectory Mismatches in Multimodal Robot Demonstrations
This work addresses the implicit misalignment between language instructions and behavioral trajectories in multimodal robot demonstration data by proposing MMPF, a training-free auditing framework. Treating each modality as an expert, MMPF estimates task label distributions through a combination of local neighborhood consistency and global prototype similarity, then fuses multimodal information via prediction entropy weighting to achieve high-accuracy mismatch detection and label correction. As the first method specifically designed to handle data errors where actions are correct but instructions are erroneous, MMPF achieves state-of-the-art performance in instruction-trajectory matching (ITM) detection and correction on both the LIBERO benchmark and real-world robot datasets. Experimental results demonstrate that MMPF significantly enhances downstream policy learning, with physical robot trials validating an effective trade-off between filtering and relabeling strategies.