QCaption: Video Captioning and Q&A through Fusion of Large Multimodal Models
This work proposes a fully self-contained multimodal fusion architecture to address the limitations of insufficient semantic understanding and reliance on external services in video captioning and question-answering tasks. By integrating keyframe extraction, a large vision-language model (LVM), and a large language model (LLM), the framework enables end-to-end, efficient video understanding without dependence on external APIs, thereby supporting fully local deployment. Experimental results demonstrate significant performance gains, with up to 44.2% improvement in video captioning and 48.9% enhancement in video-based question answering, substantially advancing the system’s accuracy, practicality, and deployability in real-world scenarios.