Multimedia Asset Personalization via Multimodal Embeddings at Netflix

📅 2026-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过使用多模态嵌入方法改进了Netflix的内容推荐系统,解决了传统模型无法处理新内容的问题,并提高了个性化推广素材的效果。
📝 Abstract
Personalized promotional assets, namely artwork images and video preview clips, are critical to content discovery on Netflix. Traditional models for asset selection rely on ID-based interaction history, leaving them blind to asset content and unable to serve newly launched titles and assets. We describe how multimodal embeddings reshaped production systems at Netflix and report transferable lessons for practitioners adopting foundation-model embeddings into recommender systems. First, pretrained image embeddings unlock cross-title, cross-canvas knowledge transfer. Augmenting a two-tower model with CLIP image embeddings lets a single model serve all five Netflix artwork canvas types, replacing five separately trained per-canvas models and substantially improving cold-start performance. A lightweight extension reuses CLIP's joint text-image space to make artwork personalization query-aware in search. Second, multimodality decisively beats any single modality for video preview personalization. We describe MediaFM, our in-house tri-modal foundation model trained on a large-scale corpus of shots from the Netflix show catalog, fusing visual (SeqCLIP), audio (wav2vec 2.0), and timed-text signals; adopted for video preview personalization, it outperforms strong visual-only baselines both offline and in online A/B tests. Third, a simple offline proxy task whose performance correlates with online outcomes can accelerate the experimentation and productization cycle. Predicting the popularity-based winner from embeddings alone ranks embedding models and versions, pruning the choice space before any end-to-end integration or A/B test; it now gates every new MediaFM checkpoint. We also share the production engineering decisions (shared embedding infrastructure, low-latency serving, cheap screening) that made these deployments viable, along with the design tradeoffs and failure modes we encountered.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Embeddings
Personalization
Content Discovery
Cold-Start
Innovation

Methods, ideas, or system contributions that make the work stand out.

multimodal embeddings
cross-canvas knowledge transfer
tri-modal foundation model
offline proxy task
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
E
Emma Yanyang Kong
Netflix, Los Gatos, California, USA
Aditya Deshpande
Aditya Deshpande
Research Scientist, Apple
Generative AIComputer VisionMachine LearningGraphicsGPU Computing
Bowei Yan
Bowei Yan
A
Asad Abbasi
Netflix, Los Gatos, California, USA
Santiago Castro
Santiago Castro
Netflix, Los Gatos, California, USA
A
Avneesh Saluja
Cohere, San Francisco, California, USA
D
David Fagnan
Netflix, Los Gatos, California, USA
A
Ashish Rastogi
Netflix, Los Gatos, California, USA