Skip to main navigation Skip to search Skip to main content

Human-Guided Online Reward Adaptation for Real-Robot Arm Manipulation

  • Tianxing Zhou
  • , Haojia Ao
  • , Haoyang Lu
  • , Guangyan Chen
  • , Zichen Zhou
  • , Te Cui
  • , Chao Yu*
  • , Yufeng Yue*
  • *Corresponding author for this work
  • Beijing Institute of Technology
  • Zhongguancun Academy
  • Tsinghua University

Research output: Contribution to journalArticlepeer-review

Abstract

Real-world reinforcement learning enables robots to adapt directly to complex physical dynamics, avoiding the reality gap inherent in simulation and rendering previously infeasible manipulation skills a reality. However, the high cost of real-world exploration strictly limits sample availability. Such tight interaction budgets make sparse rewards insufficient for sample-efficient learning in real-robot arm manipulation, creating the need for dense reward supervision. Although pretrained reward models provide useful semantic priors, their generic predictions are often misaligned with embodiment-specific and online dynamics during physical deployment. To address this issue, we propose adaptive reward via human interaction (ARHI), a parameter-efficient framework for online adaptation of VLM-based rewards in real-robot reinforcement learning. Rather than treating the pretrained reward as a fixed module, ARHI continuously recalibrates it using lightweight online updates, while preserving the semantic priors of the frozen backbone. To support this online reward adaptation under sparse human feedback, we further integrate the method into an asynchronous dual-loop system that converts sparse interventions into dense supervisory signals for reward and policy learning. Extensive real-world experiments across six manipulation tasks demonstrate that ARHI outperforms strong baselines, reducing human effort by 28.0%, while achieving a 47.6% reduction in performance drop under challenging deployment shifts.

Original languageEnglish
Pages (from-to)9072-9079
Number of pages8
JournalIEEE Robotics and Automation Letters
Volume11
Issue number8
DOIs
Publication statusAccepted/In press - 2026
Externally publishedYes

Keywords

  • Reinforcement learning
  • human-in-the-loop learning
  • reward adaptation
  • robot learning
  • visual-language models

Fingerprint

Dive into the research topics of 'Human-Guided Online Reward Adaptation for Real-Robot Arm Manipulation'. Together they form a unique fingerprint.

Cite this