跳到主要导航 跳到搜索 跳到主要内容

AffectOmni: RL-Verifiable People-Centric Grounded Affective Reasoning for Social and Art-Related Scenes

  • Yibo Wang
  • , Rui Yang
  • , Jisheng Dang
  • , Bimei Wang
  • , Yitao Wu
  • , Pengfei Cao
  • , Wencan Zhang
  • , Hong Peng
  • , Bin Hu*
  • , Tat Seng Chua
  • *此作品的通讯作者
  • Lanzhou University
  • Hainan University
  • National University of Singapore

科研成果: 期刊稿件文章同行评审

摘要

Multimodal large language models (MLLMs) achieve strong performance on VQA and scene understanding, yet affective reasoning remains vulnerable to shortcut behavior. Models may predict correct answers while neglecting people-centric cues such as micro expressions and body language, which weakens traceability and external verification. Prior reinforcement learning approaches mainly reward context or logical coherence without explicitly enforcing attention to human evidence. In addition, LLM as a Judge scoring often suffers from score clustering, which reduces reward discriminability. We propose AffectOmni, a GRPO trained framework for verifiable affective reasoning. AffectOmni introduces People Focus and Temporal Order rewards to encourage people-centric evidence selection and temporally structured reasoning, and it adopts within-group comparative scoring to produce more stable and discriminative reward signals. For verification, a Thinking Summarizer converts free form rationales into executable evidence instructions, which are grounded into pixel level evidence regions via SAM3 to provide an externally auditable interface outside the training loop. Experiments on IntentBench, Daily Omni, and WorldSense show consistent improvements over open source 7B scale baselines, including gains of 4.66% on emotion recognition and +14.29% on temporally sensitive tasks.

源语言英语
期刊IEEE Transactions on Affective Computing
DOI
出版状态已接受/待刊 - 2026
已对外发布

学术指纹

探究 'AffectOmni: RL-Verifiable People-Centric Grounded Affective Reasoning for Social and Art-Related Scenes' 的科研主题。它们共同构成独一无二的学术指纹。

引用此