From Understanding to Engagement: Personalized pharmacy Video Clips via Vision Language Models (VLMs)

📅 2026-01-08
🏛️ arXiv.org
📈 Citations: 1
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the inefficiency and inconsistency of manual annotation in traditional pharmaceutical video processing, which struggles to leverage multimodal information—particularly in large-scale, long-form videos such as clinical trial interviews. The authors propose an end-to-end framework for automatically generating highlight clips by integrating vision-language models (VLMs) and audio-language models (ALMs), enhanced with role-based prompting to enable personalized editing tailored to marketing, training, and regulatory scenarios. Key innovations include a reproducible Cut & Merge algorithm ensuring audiovisual synchronization and smooth transitions, a role-prompt-driven personalization mechanism, and a highly efficient, low-cost pipeline. Evaluated on the Video MME benchmark and a dataset of 16,159 pharmaceutical videos, the system achieves a 3–4× speedup and 4× cost reduction compared to baselines, while outperforming advanced models like Gemini 2.5 Pro in both coherence (0.348) and informativeness (0.721).

Technology Category

Application Category

📝 Abstract
Vision Language Models (VLMs) are poised to revolutionize the digital transformation of pharmacyceutical industry by enabling intelligent, scalable, and automated multi-modality content processing. Traditional manual annotation of heterogeneous data modalities (text, images, video, audio, and web links), is prone to inconsistencies, quality degradation, and inefficiencies in content utilization. The sheer volume of long video and audio data further exacerbates these challenges, (e.g. long clinical trial interviews and educational seminars). Here, we introduce a domain adapted Video to Video Clip Generation framework that integrates Audio Language Models (ALMs) and Vision Language Models (VLMs) to produce highlight clips. Our contributions are threefold: (i) a reproducible Cut&Merge algorithm with fade in/out and timestamp normalization, ensuring smooth transitions and audio/visual alignment; (ii) a personalization mechanism based on role definition and prompt injection for tailored outputs (marketing, training, regulatory); (iii) a cost efficient e2e pipeline strategy balancing ALM/VLM enhanced processing. Evaluations on Video MME benchmark (900) and our proprietary dataset of 16,159 pharmacy videos across 14 disease areas demonstrate 3 to 4 times speedup, 4 times cost reduction, and competitive clip quality. Beyond efficiency gains, we also report our methods improved clip coherence scores (0.348) and informativeness scores (0.721) over state of the art VLM baselines (e.g., Gemini 2.5 Pro), highlighting the potential of transparent, custom extractive, and compliance supporting video summarization for life sciences.
Problem

Research questions and friction points this paper is trying to address.

Vision Language Models
multi-modality content processing
video summarization
pharmaceutical industry
heterogeneous data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision Language Models
Video Summarization
Personalized Content Generation
Audio Language Models
Cut and Merge Algorithm
S
Suyash Mishra
Roche
Q
Qiang Li
Accenture
S
Srikanth Patil
Involead
Anubhav Girdhar
Anubhav Girdhar
Data Engineer
Generative AI