Moving Beyond More Views: Redundancy-Aware Ego-Exo Fusion for Proficiency Estimation

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对EgoExo动作质量评估中多视角冗余和过拟合问题,提出AdaMVS和VIB-GB模块以选择并融合最有效视角,提升性能。
📝 Abstract
EgoExo proficiency estimation aims to assess action quality by integrating fine-grained motion cues from egocentric (1st-person) views with spatial context from multiple exocentric (3rd-person) views. Simply adding more exocentric views degrades EgoExo performance, as redundant or noisy perspectives dilute useful motion cues. Our analysis identifies two key causes: (1) Multiview redundancy - From the data perspective, certain views provide limited or noisy information, diluting discriminative cues; (2) Overfitting - From the feature perspective, conventional fusion increases representational complexity, causing the model to memorise view-specific patterns rather than learn generalisable representations. To address these issues, we propose two complementary modules: AdaMVS, which adaptively identifies and fuses the most informative view tokens under weak supervision from the data perspective, and VIB-GB, which combines Gradient Blending and Variational Information Bottleneck regularisation from the feature perspective to compress redundant signals and suppress overfitting during training. Experiments on EgoExo-4D and EgoExo-Fitness demonstrate that our method learns both which view to look at and how to fuse them, achieving new state-of-the-art results. Our source code is available at https://github.com/dx199771/AdaMVS
Problem

Research questions and friction points this paper is trying to address.

EgoExo proficiency estimation
multiview redundancy
overfitting
Innovation

Methods, ideas, or system contributions that make the work stand out.

AdaMVS
VIB-GB
redundancy-aware
Ego-Exo fusion
proficiency estimation