GRACE: Boosting Video MLLMs with Grounded Action-Centric Evidence for Viewer Sentiment Prediction
This work addresses the challenge of predicting viewer sentiment in video advertisements, where fine-grained emotion-relevant behaviors and visual cues are difficult to capture from full-frame inputs. The authors propose an action-centric, structured evidence enhancement framework that extracts temporal subject-predicate-object triplets, crops visual patches of participating entities, and integrates visible text to construct explicit, spatially localizable multimodal reasoning cues. This approach uniquely combines action triplets with entity-specific visual crops to guide interpretable sentiment reasoning in multimodal large language models (Qwen2.5-VL/Qwen3-VL). Evaluated on the Pitts dataset, the method significantly outperforms baseline approaches, and transfer experiments on AdsQA and TVQA subsets demonstrate its strong generalization capability.