Institution profile

Hippo T&C Co., Ltd.

Industry researchasia · jp
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

MAMS: Model-Agnostic Module Selection Framework for Video Captioning

Jan 30, 2025

To address the critical issue of information loss or redundancy caused by fixed-frame sampling in video captioning, this paper proposes the first model-agnostic dynamic module selection framework. Our method jointly adapts the number of visual tokens and the scale of generation modules via adaptive visual token subset construction and a learnable attention masking mechanism. Key contributions include: (1) decoupling module selection from the backbone model to enable plug-and-play integration; (2) introducing a lightweight gating network for token importance estimation and subset selection; and (3) designing an adaptive attention mask to enhance modeling of salient spatiotemporal regions. Evaluated on MSVD, MSR-VTT, and VATEX, our framework consistently improves BLEU-4 and CIDEr scores across three representative video captioning architectures—achieving average gains of +2.1 and +3.7, respectively—while reducing computational overhead by 18%–25%.

0 citationsRead paper
Recent publications

Latest Papers

MAMS: Model-Agnostic Module Selection Framework for Video Captioning

Jan 30, 2025

To address the critical issue of information loss or redundancy caused by fixed-frame sampling in video captioning, this paper proposes the first model-agnostic dynamic module selection framework. Our method jointly adapts the number of visual tokens and the scale of generation modules via adaptive visual token subset construction and a learnable attention masking mechanism. Key contributions include: (1) decoupling module selection from the backbone model to enable plug-and-play integration; (2) introducing a lightweight gating network for token importance estimation and subset selection; and (3) designing an adaptive attention mask to enhance modeling of salient spatiotemporal regions. Evaluated on MSVD, MSR-VTT, and VATEX, our framework consistently improves BLEU-4 and CIDEr scores across three representative video captioning architectures—achieving average gains of +2.1 and +3.7, respectively—while reducing computational overhead by 18%–25%.

0 citationsRead paper