Right Answer, Wrong Heat: Explanation-Aware Evaluation and Thermal-Grounded Feedback for MLLMs on Infrared Images

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the tendency of multimodal large language models to produce correct answers on infrared image tasks while generating explanations that lack grounding in thermal characteristics. To tackle this issue, the paper introduces the first explanation-aware framework that explicitly treats “thermal groundedness” as an independent evaluation dimension. Through a dual-model consensus evaluator, human-anchored calibration, and controlled experiments with visible-light renderings, the study reveals that high-capability models often rely on visible-spectrum cues when infrared input is absent. Furthermore, the authors propose a training-free Thermal Groundedness Feedback (TGF) mechanism that significantly enhances the thermal basis of generated explanations without altering the predicted answers, demonstrating that answer accuracy alone is insufficient for evaluating true infrared understanding.
📝 Abstract
General-purpose multimodal large language models (MLLMs) are increasingly applied to infrared images, where they are commonly scored by answer accuracy alone. However, a correct answer does not ensure that the model's explanation is grounded in infrared thermal evidence. We introduce an explanation-aware evaluation framework that separates answer correctness, output-level explanation groundedness, and thermal grounding for infrared visual questions. Using a Dual-LLM Consensus Judge with a preliminary human-anchor calibration check, we find that correct answers can still rely on weak or visible-light evidence; withholding the original infrared image and showing only a visible-like rendering erodes thermal grounding with little accuracy change; and this erosion is observed most strongly for more capable models but disappears when infrared remains available. We further propose Thermal-Grounded Feedback (TGF), a training-free feedback loop that diagnoses explanation-side failures and revises the explanation while preserving the selected answer. On local paired-input validation, TGF improves explanation-side grounding without changing answers. These findings suggest that future trustworthy MLLMs for infrared scene understanding should be evaluated and developed to produce thermally grounded explanations rather than merely accurate answers.
Problem

Research questions and friction points this paper is trying to address.

infrared images
multimodal large language models
explanation groundedness
thermal grounding
answer accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

explanation-aware evaluation
thermal grounding
multimodal LLMs
infrared images
Thermal-Grounded Feedback
🔎 Similar Papers
2024-03-22IEEE transactions on circuits and systems for video technology (Print)Citations: 2