Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study reveals a critical blind spot in current large vision-language models (LVLMs): their inability to reliably assess the temporal and logical coherence of image sequences, particularly in detecting semantically misordered or contradictory narratives. Through systematic evaluation on pairwise temporal discrimination tasks, we uncover a structural failure in LVLMs driven by primacy and recency effects, causing them to over-rely on frame position rather than genuine narrative logic. This bias stems from architectural choices such as causal masking and rotary embeddings. Diagnostic probes, positional perturbations, and temporal discrimination experiments demonstrate that while LVLMs perform adequately on frame-level scoring, their performance sharply degrades in long-range temporal reasoning. These findings challenge the prevailing snapshot-centric evaluation paradigm and question the suitability of LVLMs as trustworthy evaluators of visual narrative coherence.
📝 Abstract
As generative multimedia evolves from static image synthesis to complex, interleaved visual narratives, a foundational bottleneck has emerged: the judgment crisis. While human perception naturally synthesizes the temporal and logical flow of a story, automated evaluation systems remain largely "blind" to sequential continuity, often failing to distinguish between a coherent narrative and a semantically shuffled or contradictory sequence. This work identifies a critical structural gap in current multimodal evaluation paradigms, arguing that the reliance on Large Vision-Language Models (LVLMs) as judges is fundamentally limited by architectural biases. Our analysis reveals a profound performance dichotomy: while models may appear competent in isolated pointwise scoring, they suffer a catastrophic collapse when required to perform pairwise discrimination of temporal order. We demonstrate that this is not merely a data-scarcity issue but a structural one. Through a series of diagnostic probes, we uncover systematic positional asymmetries, specifically primacy and recency effects, where a model's judgment of a story is significantly influenced by the placement of a frame, often more than by its semantic consistency. These biases, potentially rooted in causal masking and rotary embeddings, suggest that current transformer-based judges are inherently ill-equipped for long-form visual reasoning. By exposing these blind spots, we challenge the multimedia community to move beyond snapshot-centric metrics and instead pioneer Temporally-Aware Evaluation paradigms that treat visual sequences as unified logical structures rather than unordered collections of frames.
Problem

Research questions and friction points this paper is trying to address.

temporal reasoning
image sequences
Large Vision-Language Models
sequential continuity
multimodal evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

temporal reasoning
Large Vision-Language Models
sequence evaluation
positional bias
multimodal narrative
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Martina Ianaro
University of Bologna, Bologna, Italy
Guilherme Fernandes
Guilherme Fernandes
NOVA School of Science and Technology, NOVA Laboratory for Computer Science and Informatics, Lisbon, Portugal
Maurizio Gabbrielli
Maurizio Gabbrielli
Professor of Computer Science, University of Bologna
Programming languagesartificial intelligenceSOC
J
João Magalhães
NOVA School of Science and Technology, NOVA Laboratory for Computer Science and Informatics, Lisbon, Portugal