TY - GEN
T1 - Multi-Agent Reinforcement Learning with Relational Trust for Adversarial Cloud Settings
AU - Wang, Hanlin
AU - Gai, Keke
AU - Yu, Jing
AU - Xu, Lei
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Cloud-based multi-agent systems that involve adversarial participants face a critical challenge. Agents must infer the trustworthiness of their partners from ongoing interactions before they can coordinate effectively. Existing multi-agent reinforcement learning methods such as Multi-Agent Proximal Policy Optimization (MAPPO) rely solely on instantaneous observations and discard the relational signals that emerge from repeated agent interactions. This paper proposes SEM-MAPPO, which augments MAPPO with a Social-Emotional Memory (SEM) module that implements a three-stage relational knowledge lifecycle. The module acquires pairwise interaction signals through an appraisal-inspired affective encoder, maintains evolving relational states in an embedding matrix, and distills the stored knowledge into trust-weighted social features for decision augmentation. We evaluate SEM-MAPPO on five Multi-Agent Particle Environment (MPE) scenarios covering cooperative, communication, and mixed settings. On the mixed Simple adversary task, SEM-MAPPO improves the final return from -2.03 to 0.21 under the same MAPPO training protocol, while preserving stable performance on cooperative tasks. Three-way ablation confirms that each lifecycle stage contributes, and the computational overhead is only 3 to 7 percent. These results suggest that dynamic relational knowledge is most beneficial when agents must distinguish cooperators from adversaries online, a setting that motivates trust-aware coordination in cloud security monitoring.
AB - Cloud-based multi-agent systems that involve adversarial participants face a critical challenge. Agents must infer the trustworthiness of their partners from ongoing interactions before they can coordinate effectively. Existing multi-agent reinforcement learning methods such as Multi-Agent Proximal Policy Optimization (MAPPO) rely solely on instantaneous observations and discard the relational signals that emerge from repeated agent interactions. This paper proposes SEM-MAPPO, which augments MAPPO with a Social-Emotional Memory (SEM) module that implements a three-stage relational knowledge lifecycle. The module acquires pairwise interaction signals through an appraisal-inspired affective encoder, maintains evolving relational states in an embedding matrix, and distills the stored knowledge into trust-weighted social features for decision augmentation. We evaluate SEM-MAPPO on five Multi-Agent Particle Environment (MPE) scenarios covering cooperative, communication, and mixed settings. On the mixed Simple adversary task, SEM-MAPPO improves the final return from -2.03 to 0.21 under the same MAPPO training protocol, while preserving stable performance on cooperative tasks. Three-way ablation confirms that each lifecycle stage contributes, and the computational overhead is only 3 to 7 percent. These results suggest that dynamic relational knowledge is most beneficial when agents must distinguish cooperators from adversaries online, a setting that motivates trust-aware coordination in cloud security monitoring.
KW - cloud security
KW - dynamic relational knowledge
KW - Multi-agent reinforcement learning
KW - social-emotional memory
KW - trust-based coordination
UR - https://www.scopus.com/pages/publications/105047535717
U2 - 10.1109/SmartCloud69481.2026.00020
DO - 10.1109/SmartCloud69481.2026.00020
M3 - Conference contribution
AN - SCOPUS:105047535717
T3 - Proceedings - 2026 IEEE 11th International Conference on Smart Cloud, SmartCloud 2026
SP - 66
EP - 71
BT - Proceedings - 2026 IEEE 11th International Conference on Smart Cloud, SmartCloud 2026
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 11th IEEE International Conference on Smart Cloud, SmartCloud 2026
Y2 - 8 May 2026 through 10 May 2026
ER -