Skip to main navigation Skip to search Skip to main content

Constructing LTL-based Reward Machines for Reinforcement Learning via Bayesian Approach

  • Feiyu Yu
  • , Qizhen Wu
  • , Lei Chen*
  • *Corresponding author for this work
  • Beijing Institute of Technology
  • Beihang University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Reward Machines (RMs) can tackle sparse reward challenges in Reinforcement Learning, thus leading to more efficient training. However, their practical application is hindered by the requirement of extensive expert knowledge for manual design. Here, we introduce a novel Bayesian method to construct RMs without hand-crafting. Embed with linear temporal logic, our method discovers temporal specifications from complex environment tasks, and the specifications are subsequently translated into RMs. The constructed RMs enable agents to perform more goal-directed actions and clearly understand task progression when facing sparse rewards. Experimental results demonstrate that our method quickly generates highly interpretable and robust RMs compared to existing approaches, even from noisy data in partially observable environments.

Original languageEnglish
Title of host publicationProceedings - 2025 China Automation Congress, CAC 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages6634-6639
Number of pages6
ISBN (Electronic)9798331589677
DOIs
Publication statusPublished - 2025
Externally publishedYes
Event2025 China Automation Congress, CAC 2025 - Harbin, China
Duration: 26 Sept 202528 Sept 2025

Publication series

NameProceedings - 2025 China Automation Congress, CAC 2025

Conference

Conference2025 China Automation Congress, CAC 2025
Country/TerritoryChina
CityHarbin
Period26/09/2528/09/25

Keywords

  • Bayesian framework
  • Linear Temporal Logic
  • Reinforcement Learning
  • Reward Machines

Fingerprint

Dive into the research topics of 'Constructing LTL-based Reward Machines for Reinforcement Learning via Bayesian Approach'. Together they form a unique fingerprint.

Cite this