TY - GEN
T1 - Temporal Multimodal Encoding for Reinforcement Learning-Driven Quadruped Robotic Control
AU - Chen, Chongming
AU - Ren, Xuemei
AU - Zheng, Dongdong
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2026.
PY - 2026
Y1 - 2026
N2 - In recent years, reinforcement learning (RL) has made significant progress in the field of legged motion control. Unlike traditional methods that rely on precise parameters, vision-free strategies based on RL extract features from a robot’s proprioceptive information through an encoder, enabling stable motion without visual input. However, multilayer perceptron (MLP) is predominantly utilized in existing encoders, and the static feature extraction mechanisms of MLP face difficulties in capturing dynamic characteristics in time-series data, consequently limiting the robot’s adaptability in complex environments. To tackle this problem, a temporal multimodal encoding (TME) method is proposed. In this method, the advantages of gated recurrent unit (GRU) and MLP are combined, facilitating the deep fusion of temporal and spatial features. Additionally, a contrastive learning mechanism is introduced, where external modalities are contrastively learned alongside features extracted from the robot’s historical proprioceptive information, thus improving the model’s representation capability and generalization performance. The effectiveness of the proposed method has been verified through simulation experiments on a quadruped robot.
AB - In recent years, reinforcement learning (RL) has made significant progress in the field of legged motion control. Unlike traditional methods that rely on precise parameters, vision-free strategies based on RL extract features from a robot’s proprioceptive information through an encoder, enabling stable motion without visual input. However, multilayer perceptron (MLP) is predominantly utilized in existing encoders, and the static feature extraction mechanisms of MLP face difficulties in capturing dynamic characteristics in time-series data, consequently limiting the robot’s adaptability in complex environments. To tackle this problem, a temporal multimodal encoding (TME) method is proposed. In this method, the advantages of gated recurrent unit (GRU) and MLP are combined, facilitating the deep fusion of temporal and spatial features. Additionally, a contrastive learning mechanism is introduced, where external modalities are contrastively learned alongside features extracted from the robot’s historical proprioceptive information, thus improving the model’s representation capability and generalization performance. The effectiveness of the proposed method has been verified through simulation experiments on a quadruped robot.
KW - Contrastive learning
KW - Reinforcement learning
UR - https://www.scopus.com/pages/publications/105040368111
U2 - 10.1007/978-981-95-6557-3_28
DO - 10.1007/978-981-95-6557-3_28
M3 - Conference contribution
AN - SCOPUS:105040368111
SN - 9789819565566
T3 - Lecture Notes in Electrical Engineering
SP - 286
EP - 296
BT - Proceedings of 2025 Chinese Intelligent Systems Conference - Volume 2
A2 - Jia, Yingmin
A2 - Zhang, Weicun
A2 - Fu, Yongling
A2 - Liu, Yang
PB - Springer Science and Business Media Deutschland GmbH
T2 - 21st Chinese Intelligent Systems Conference, CISC 2025
Y2 - 25 October 2025 through 26 October 2025
ER -