TY - JOUR
T1 - Human-Guided Online Reward Adaptation for Real-Robot Arm Manipulation
AU - Zhou, Tianxing
AU - Ao, Haojia
AU - Lu, Haoyang
AU - Chen, Guangyan
AU - Zhou, Zichen
AU - Cui, Te
AU - Yu, Chao
AU - Yue, Yufeng
N1 - Publisher Copyright:
© 2016 IEEE.
PY - 2026
Y1 - 2026
N2 - Real-world reinforcement learning enables robots to adapt directly to complex physical dynamics, avoiding the reality gap inherent in simulation and rendering previously infeasible manipulation skills a reality. However, the high cost of real-world exploration strictly limits sample availability. Such tight interaction budgets make sparse rewards insufficient for sample-efficient learning in real-robot arm manipulation, creating the need for dense reward supervision. Although pretrained reward models provide useful semantic priors, their generic predictions are often misaligned with embodiment-specific and online dynamics during physical deployment. To address this issue, we propose adaptive reward via human interaction (ARHI), a parameter-efficient framework for online adaptation of VLM-based rewards in real-robot reinforcement learning. Rather than treating the pretrained reward as a fixed module, ARHI continuously recalibrates it using lightweight online updates, while preserving the semantic priors of the frozen backbone. To support this online reward adaptation under sparse human feedback, we further integrate the method into an asynchronous dual-loop system that converts sparse interventions into dense supervisory signals for reward and policy learning. Extensive real-world experiments across six manipulation tasks demonstrate that ARHI outperforms strong baselines, reducing human effort by 28.0%, while achieving a 47.6% reduction in performance drop under challenging deployment shifts.
AB - Real-world reinforcement learning enables robots to adapt directly to complex physical dynamics, avoiding the reality gap inherent in simulation and rendering previously infeasible manipulation skills a reality. However, the high cost of real-world exploration strictly limits sample availability. Such tight interaction budgets make sparse rewards insufficient for sample-efficient learning in real-robot arm manipulation, creating the need for dense reward supervision. Although pretrained reward models provide useful semantic priors, their generic predictions are often misaligned with embodiment-specific and online dynamics during physical deployment. To address this issue, we propose adaptive reward via human interaction (ARHI), a parameter-efficient framework for online adaptation of VLM-based rewards in real-robot reinforcement learning. Rather than treating the pretrained reward as a fixed module, ARHI continuously recalibrates it using lightweight online updates, while preserving the semantic priors of the frozen backbone. To support this online reward adaptation under sparse human feedback, we further integrate the method into an asynchronous dual-loop system that converts sparse interventions into dense supervisory signals for reward and policy learning. Extensive real-world experiments across six manipulation tasks demonstrate that ARHI outperforms strong baselines, reducing human effort by 28.0%, while achieving a 47.6% reduction in performance drop under challenging deployment shifts.
KW - Reinforcement learning
KW - human-in-the-loop learning
KW - reward adaptation
KW - robot learning
KW - visual-language models
UR - https://www.scopus.com/pages/publications/105041000721
U2 - 10.1109/LRA.2026.3699150
DO - 10.1109/LRA.2026.3699150
M3 - Article
AN - SCOPUS:105041000721
SN - 2377-3766
VL - 11
SP - 9072
EP - 9079
JO - IEEE Robotics and Automation Letters
JF - IEEE Robotics and Automation Letters
IS - 8
ER -