AffectOmni: RL-Verifiable People-Centric Grounded Affective Reasoning for Social and Art-Related Scenes

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出AffectOmni框架,通过引入People Focus和Temporal Order奖励以及采用组内比较评分方法,解决多模态大语言模型在情感推理中忽视以人为中心线索的问题。
📝 Abstract
Multimodal large language models (MLLMs) achieve strong performance on VQA and scene understanding, yet affective reasoning remains vulnerable to shortcut behavior. Models may predict correct answers while neglecting people-centric cues such as micro expressions and body language, which weakens traceability and external verification. Prior reinforcement learning approaches mainly reward context or logical coherence without explicitly enforcing attention to human evidence. In addition, LLM as a Judge scoring often suffers from score clustering, which reduces reward discriminability. We propose AffectOmni, a GRPO trained framework for verifiable affective reasoning. AffectOmni introduces People Focus and Temporal Order rewards to encourage people-centric evidence selection and temporally structured reasoning, and it adopts within-group comparative scoring to produce more stable and discriminative reward signals. For verification, a Thinking Summarizer converts free form rationales into executable evidence instructions, which are grounded into pixel level evidence regions via SAM3 to provide an externally auditable interface outside the training loop. Experiments on IntentBench, Daily Omni, and WorldSense show consistent improvements over open source 7B scale baselines, including gains of 4.66% on emotion recognition and +14.29% on temporally sensitive tasks. Code is available at https://github.com/eliot127825-rgb/AffectOmni_nobody.
Problem

Research questions and friction points this paper is trying to address.

affective reasoning
shortcut behavior
people-centric cues
reinforcement learning
reward discriminability
Innovation

Methods, ideas, or system contributions that make the work stand out.

AffectOmni
People Focus
Temporal Order
within-group comparative scoring
Thinking Summarizer
Y
Yibo Wang
Lanzhou University, Lanzhou, China
R
Rui Yang
Lanzhou University, Lanzhou, China
J
Jisheng Dang
Lanzhou University, Lanzhou, China
B
Bimei Wang
Lanzhou University, Lanzhou, China
Y
Yitao Wu
Hainan University, Haikou, China
P
Pengfei Cao
Lanzhou University, Lanzhou, China
W
Wencan Zhang
School of Computing, National University of Singapore, Singapore
Hong Peng
Hong Peng
Vice Professor of Physice, Lanzhou University
EEGAffective ComputingDepressionAnxiety neurosis
B
Bin Hu
Lanzhou University, Lanzhou, China
T
Tat-Seng Chua
School of Computing, National University of Singapore, Singapore