From Understanding to Engagement: Personalized pharmacy Video Clips via Vision Language Models (VLMs)
This work addresses the inefficiency and inconsistency of manual annotation in traditional pharmaceutical video processing, which struggles to leverage multimodal information—particularly in large-scale, long-form videos such as clinical trial interviews. The authors propose an end-to-end framework for automatically generating highlight clips by integrating vision-language models (VLMs) and audio-language models (ALMs), enhanced with role-based prompting to enable personalized editing tailored to marketing, training, and regulatory scenarios. Key innovations include a reproducible Cut & Merge algorithm ensuring audiovisual synchronization and smooth transitions, a role-prompt-driven personalization mechanism, and a highly efficient, low-cost pipeline. Evaluated on the Video MME benchmark and a dataset of 16,159 pharmaceutical videos, the system achieves a 3–4× speedup and 4× cost reduction compared to baselines, while outperforming advanced models like Gemini 2.5 Pro in both coherence (0.348) and informativeness (0.721).