FabriMAE I Trust Myself? Self-Evaluating VLA Action Generation with Markov Attention Entropy
This study addresses the challenge of self-assessing action reliability in Vision-Language-Action (VLA) models lacking external supervision. We propose the Markov Attention Entropy (MAE) framework, which converts internal attention entropy into reliability scores and integrates a multi-sampling strategy to enable unsupervised self-evaluation and test-time action selection. Furthermore, we introduce the LIBERO-Reflect benchmark to elucidate cross-architectural abstractions in action generation. Experimental results demonstrate that MAE outperforms state-of-the-art baselines across multiple metrics, significantly enhancing the robustness of PI-series models with minimal runtime overhead. Consequently, this work provides an efficient solution for the trustworthy deployment of VLA systems by facilitating reliable autonomous decision-making without requiring additional labeled data or external feedback mechanisms during inference.