A Dual-Transformer for Multi-Camera View Recommendation

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种新的Dual-Transformer架构,通过交叉注意力机制解决多摄像机视角推荐问题,显著提高了TVMCE数据集上的性能。
📝 Abstract
Multi-camera systems are foundational to modern media production, and multi-camera editing is a critical task. This involves the proper selection of the appropriate camera view at each moment. In this paper, we propose a novel Dual-Transformer architecture with Cross-Attention that heavily outperformed the current SOTA models over the TVMCE dataset (TV Shows Multicamera Editing dataset). Our model decouples these tasks: (1) a dedicated temporal encoder first processes the sequence of past frames to build a rich memory of the recent history, and (2) the candidate camera views then act as queries to this memory via a cross-attention module, allowing each candidate to independently interrogate the historical context and find the most relevant information for its own evaluation. Our approach achieved 56.60% Precision@0.5, representing a substantial improvement over the prior best result of 37.16%. We further conducted an ablation study exploring the use of lightweight backbone architectures, where the SwinV2 backbone yielded the best performance, achieving 69.65% Precision@0.5. Using this best-performing configuration, we then investigated the feasibility of adapting the model to replicate the editing style of a specific human editor. To this end, we fine-tuned the model using varying proportions of the initial segment of a target video. Our results demonstrate that even with only 20% of the video used for fine-tuning, the model exhibited measurable improvements in Precision@0.5, indicating strong potential for data-efficient personalization of editing style adapted to each individual TV show or producer.
Problem

Research questions and friction points this paper is trying to address.

Multi-Camera Systems
View Selection
Media Production
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dual-Transformer
Cross-Attention
SwinV2 backbone
data-efficient personalization
🔎 Similar Papers
No similar papers found.
J
Josep Cabacas-Maso
eHealth Center, Faculty of Computer Science, Multimedia and Telecommunications, Universitat Oberta de Catalunya, 08018 Barcelona, Spain
Carles Ventura
Carles Ventura
Universitat Oberta de Catalunya (UOC)
Computer visionImage and video segmentation
I
Ismael Benito-Altamirano
eHealth Center, Faculty of Computer Science, Multimedia and Telecommunications, Universitat Oberta de Catalunya, 08018 Barcelona, Spain