MAMS: Model-Agnostic Module Selection Framework for Video Captioning
To address the critical issue of information loss or redundancy caused by fixed-frame sampling in video captioning, this paper proposes the first model-agnostic dynamic module selection framework. Our method jointly adapts the number of visual tokens and the scale of generation modules via adaptive visual token subset construction and a learnable attention masking mechanism. Key contributions include: (1) decoupling module selection from the backbone model to enable plug-and-play integration; (2) introducing a lightweight gating network for token importance estimation and subset selection; and (3) designing an adaptive attention mask to enhance modeling of salient spatiotemporal regions. Evaluated on MSVD, MSR-VTT, and VATEX, our framework consistently improves BLEU-4 and CIDEr scores across three representative video captioning architectures—achieving average gains of +2.1 and +3.7, respectively—while reducing computational overhead by 18%–25%.