TY - JOUR
T1 - AffectOmni
T2 - RL-Verifiable People-Centric Grounded Affective Reasoning for Social and Art-Related Scenes
AU - Wang, Yibo
AU - Yang, Rui
AU - Dang, Jisheng
AU - Wang, Bimei
AU - Wu, Yitao
AU - Cao, Pengfei
AU - Zhang, Wencan
AU - Peng, Hong
AU - Hu, Bin
AU - Chua, Tat Seng
N1 - Publisher Copyright:
© 2010-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Multimodal large language models (MLLMs) achieve strong performance on VQA and scene understanding, yet affective reasoning remains vulnerable to shortcut behavior. Models may predict correct answers while neglecting people-centric cues such as micro expressions and body language, which weakens traceability and external verification. Prior reinforcement learning approaches mainly reward context or logical coherence without explicitly enforcing attention to human evidence. In addition, LLM as a Judge scoring often suffers from score clustering, which reduces reward discriminability. We propose AffectOmni, a GRPO trained framework for verifiable affective reasoning. AffectOmni introduces People Focus and Temporal Order rewards to encourage people-centric evidence selection and temporally structured reasoning, and it adopts within-group comparative scoring to produce more stable and discriminative reward signals. For verification, a Thinking Summarizer converts free form rationales into executable evidence instructions, which are grounded into pixel level evidence regions via SAM3 to provide an externally auditable interface outside the training loop. Experiments on IntentBench, Daily Omni, and WorldSense show consistent improvements over open source 7B scale baselines, including gains of 4.66% on emotion recognition and +14.29% on temporally sensitive tasks.
AB - Multimodal large language models (MLLMs) achieve strong performance on VQA and scene understanding, yet affective reasoning remains vulnerable to shortcut behavior. Models may predict correct answers while neglecting people-centric cues such as micro expressions and body language, which weakens traceability and external verification. Prior reinforcement learning approaches mainly reward context or logical coherence without explicitly enforcing attention to human evidence. In addition, LLM as a Judge scoring often suffers from score clustering, which reduces reward discriminability. We propose AffectOmni, a GRPO trained framework for verifiable affective reasoning. AffectOmni introduces People Focus and Temporal Order rewards to encourage people-centric evidence selection and temporally structured reasoning, and it adopts within-group comparative scoring to produce more stable and discriminative reward signals. For verification, a Thinking Summarizer converts free form rationales into executable evidence instructions, which are grounded into pixel level evidence regions via SAM3 to provide an externally auditable interface outside the training loop. Experiments on IntentBench, Daily Omni, and WorldSense show consistent improvements over open source 7B scale baselines, including gains of 4.66% on emotion recognition and +14.29% on temporally sensitive tasks.
KW - Affective Computing
KW - Emotion Recognition
KW - Multimodal Large Language Models
KW - Reinforcement Learning
KW - Social Scene Understanding
KW - Visual Grounding
UR - https://www.scopus.com/pages/publications/105043711295
U2 - 10.1109/TAFFC.2026.3707634
DO - 10.1109/TAFFC.2026.3707634
M3 - Article
AN - SCOPUS:105043711295
SN - 1949-3045
JO - IEEE Transactions on Affective Computing
JF - IEEE Transactions on Affective Computing
ER -