MMDS-Bench: Benchmarking Multimodal Large Language Models on Dynamic Stance in Social Media Interactions

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过构建MMDS-Bench数据集,评估了12种多模态大语言模型在社交媒体中动态立场分类问题上的表现,特别是跨模态理解与推理能力。
📝 Abstract
Dynamic stance classification models how a reply responds to its direct parent message, rather than how a post relates to a fixed topic. Existing work has mainly studied this problem in text-only settings, while social media interactions increasingly rely on images, screenshots, memes, reaction images, and cross-modal references. We introduce MMDS-Bench, a diagnostic benchmark for multimodal dynamic stance classification in social media parent-reply interactions. MMDS-Bench contains 3,482 multimodal instances annotated with a seven-label dynamic stance taxonomy, together with an 800-instance diagnostic subset that requires structured reasoning over parent understanding, reply understanding, and stance-relation inference. We further annotate each instance with five challenge factors covering multimodal fusion, parent framing, non-literal expression, interaction reasoning, and label-boundary ambiguity. We evaluate 12 closed-source and open-source multimodal large language models and propose a reference-grounded LLM-judge protocol for assessing reasoning quality. Results show that current MLLMs still struggle with multimodal dynamic stance understanding, especially in cases that require relational inference beyond separate parent and reply comprehension.
Problem

Research questions and friction points this paper is trying to address.

Dynamic Stance
Multimodal
Social Media Interactions
Parent-Reply
Innovation

Methods, ideas, or system contributions that make the work stand out.

multimodal dynamic stance
social media interactions
structured reasoning
challenge factors
reference-grounded LLM-judge protocol
🔎 Similar Papers
No similar papers found.