MLLM-Assisted Audio VOS: A 3rd Place Report for the MeViS-Audio Track, 8th LSVOS Challenge

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种无需训练的框架,通过结合多模态大语言模型与基于SAM的分割模型,解决了音频引导的视频对象分割问题。
📝 Abstract
In this technical report, we present a training-free framework for audio-guided video object segmentation, which integrates Multimodal Large Language Models (MLLMs) with SAM-based segmentation models. We decompose the task into several stages and identify suitable foundation models for each stage. Without introducing additional model training or task-specific fine-tuning, our approach leverages the strong multimodal reasoning capabilities of MLLMs to model text-visual correspondence and employs SAM-based models for accurate object mask generation. The proposed framework demonstrates the effectiveness of leveraging foundation models for audio-guided video segmentation and achieves competitive performance in the MeViS-Audio Track of the 8th LSVOS Challenge.
Problem

Research questions and friction points this paper is trying to address.

audio-guided video object segmentation
Multimodal Large Language Models
SAM-based models
foundation models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Large Language Models
SAM-based segmentation models
audio-guided video object segmentation
🔎 Similar Papers
No similar papers found.
L
Liangtao Shi
Hefei University of Technology
J
Jinxia Xie
Nanjing University of Science and Technology
Xiantao Hu
Xiantao Hu
Nanjing University of Science & Technology
Computer VIsion
T
Ting Liu
Hunan Police College