Deeply Interleaved Text-Image Contexts for Multimodal LLMs Assessment

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决多模态模型在交错文本-图像场景中的评估与训练不足问题,通过构建TIC-Bench基准测试,涵盖逻辑、时间、空间关联三大领域,以提升模型整合文本-图像信息的能力。
📝 Abstract
Current evaluations and training of multimodal models predominantly focus on multi-image tasks, largely overlooking interleaved text-image scenarios. In such multi-image tasks, text typically serves merely as task instructions, lacking deep semantic interaction with the visual content. In contrast, realworld applications like text-image co-creation, character tracking, and spatial reconstruction require constant interaction between text and images. Consequently, models must possess a deep understanding of these interleaved contexts. To bridge this gap, we introduce a novel benchmark, TIC-Bench (deeply interleaved Text-Image Contexts), designed to evaluate the capability of models to integrate text-image clues and recover the ground truth facts within deeply interleaved contexts. This benchmark encompasses three core domains: Logical, Temporal, and Spatial Association, which are further categorized into eight specific types, comprising a total of 2,280 questions. We evaluated 10 state-of-the-art MLLMs and observed a substantial performance gap compared to human experts, together with persistent difficulties in integrating evidence distributed across interleaved visual and textual inputs. Ultimately, this benchmark provides a valuable analytical tool for assessing and advancing the ability of multimodal models to effectively integrate text and image information in deeply interleaved contexts. TIC-Bench is publicly available at https://huggingface.co/datasets/pino10010/TIC-Bench
Problem

Research questions and friction points this paper is trying to address.

multimodal models
interleaved text-image scenarios
semantic interaction
real-world applications
deep understanding
Innovation

Methods, ideas, or system contributions that make the work stand out.

TIC-Bench
interleaved text-image contexts
multimodal models
semantic interaction
Zihao Wang
Zihao Wang
Harbin Institute of Technology
Generative model
X
Xi Xiang
Harbin Institute of Technology
Y
Yuwen Sun
Harbin Institute of Technology
Y
Yingyu Li
Harbin Institute of Technology
Y
Yabo Zhang
Harbin Institute of Technology
Y
Yihan Zeng
Huawei Noah’s Ark Lab
F
Fan Li
Huawei Noah’s Ark Lab
Wangmeng Zuo
Wangmeng Zuo
School of Computer Science and Technology, Harbin Institute of Technology
Computer VisionImage ProcessingGenerative AIDeep LearningBiometrics