When Irrelevant Text Matters: Affine Margin Shifts in Multimodal Large Language Models

📅 2026-06-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了无关文本对多模态大语言模型视觉任务的影响,通过控制干预和决策边界分析,揭示了无关文本导致的模型偏好可估计畸变。
📝 Abstract
Multimodal large language models (MLLMs) are frequently exposed to auxiliary textual context, the impact of which on visually grounded tasks remains underexplored. In this paper, we investigate the influence of task-irrelevant context by formulating it as a controlled intervention within a binary visual judgment framework. By maintaining an invariant prompt structure while varying auxiliary inputs, we observe that irrelevant text consistently biases model predictions across diverse benchmarks. To move beyond performance metrics, we characterize this sensitivity through a decision margin defined by the log-probability difference between binary candidates. Our analysis reveals a robust geometric regularity: contextconditioned margins follow a consistent affine transformation of their context-free counterparts. This finding demonstrates that irrelevant context does not manifest as unstructured stochastic noise but as a estimable distortion of model preference. We further interpret the fitted affine parameters as metrics for visual commitment preservation and directional answer bias. These findings provide a margin-level diagnostic view of irrelevant-context effects in MLLMs and offer a basis for future studies on noisy-context robustness
Problem

Research questions and friction points this paper is trying to address.

multimodal large language models
task-irrelevant context
visual judgment
decision margin
affine transformation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Affine Transformation
Decision Margin
Context Sensitivity
Multimodal Large Language Models
🔎 Similar Papers
No similar papers found.