Institution profile

Oosto

Industry researcheurope · il
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Cap2Sum: Learning to Summarize Videos by Generating Captions

Aug 23, 2024arXiv.org

To address the high cost of manual annotation and limited data scale in video summarization—which severely hinder model generalization—this paper proposes a large-scale training paradigm leveraging dense video captions as weak supervision. Methodologically, it pioneers the use of dense captions instead of human-generated summaries as supervisory signals; incorporates CLIP’s vision-language priors to explicitly recover salient objects missing from captions; and designs a Transformer-based cross-modal generation architecture that integrates weakly supervised learning with zero-shot transfer and cross-dataset fine-tuning strategies. Evaluated on two newly constructed benchmarks—TVSum-Caption and SumMe-Caption—the approach substantially outperforms prior methods, achieving significant improvements in summary quality and cross-domain generalization. This work establishes a viable pathway toward low-cost, large-scale video summarization.

1 citationsRead paper

Cross-Modal Distillation For Widely Differing Modalities

Jul 22, 2025

To address overfitting in knowledge distillation caused by modality heterogeneity in multimodal learning, this paper proposes a teacher-student framework tailored for discriminative cross-modal knowledge transfer. Methodologically, it integrates cross-modal knowledge distillation, joint feature-classifier alignment, and dynamic sample weighting—without requiring strong inter-modal alignment assumptions. Its key contributions are: (1) a two-level soft-constraint distillation strategy that jointly aligns heterogeneous modalities in both feature space and classifier output space; and (2) a data-quality-aware adaptive sample weighting mechanism to enhance model robustness. Evaluated on speaker identification and image classification tasks, the method significantly improves cross-modal knowledge transfer efficiency and generalization across vision, language, and speech modalities. Notably, it demonstrates superior robustness on low-quality samples, validating its effectiveness under realistic, noisy conditions.

0 citationsRead paper
Recent publications

Latest Papers

Cross-Modal Distillation For Widely Differing Modalities

Jul 22, 2025

To address overfitting in knowledge distillation caused by modality heterogeneity in multimodal learning, this paper proposes a teacher-student framework tailored for discriminative cross-modal knowledge transfer. Methodologically, it integrates cross-modal knowledge distillation, joint feature-classifier alignment, and dynamic sample weighting—without requiring strong inter-modal alignment assumptions. Its key contributions are: (1) a two-level soft-constraint distillation strategy that jointly aligns heterogeneous modalities in both feature space and classifier output space; and (2) a data-quality-aware adaptive sample weighting mechanism to enhance model robustness. Evaluated on speaker identification and image classification tasks, the method significantly improves cross-modal knowledge transfer efficiency and generalization across vision, language, and speech modalities. Notably, it demonstrates superior robustness on low-quality samples, validating its effectiveness under realistic, noisy conditions.

0 citationsRead paper

Cap2Sum: Learning to Summarize Videos by Generating Captions

Aug 23, 2024arXiv.org

To address the high cost of manual annotation and limited data scale in video summarization—which severely hinder model generalization—this paper proposes a large-scale training paradigm leveraging dense video captions as weak supervision. Methodologically, it pioneers the use of dense captions instead of human-generated summaries as supervisory signals; incorporates CLIP’s vision-language priors to explicitly recover salient objects missing from captions; and designs a Transformer-based cross-modal generation architecture that integrates weakly supervised learning with zero-shot transfer and cross-dataset fine-tuning strategies. Evaluated on two newly constructed benchmarks—TVSum-Caption and SumMe-Caption—the approach substantially outperforms prior methods, achieving significant improvements in summary quality and cross-domain generalization. This work establishes a viable pathway toward low-cost, large-scale video summarization.

1 citationsRead paper