Abstract
The International Lunar Research Station will be established near the moon's south pole through advanced unmanned rovers at the beginning. Given the region's limited daylight and the inherent uncertainties in mobility and resource collection, traditional remote control methods prove inadequate. Moreover, assessing each rover's time and resource consumption during operations is inappropriate due to the inherent uncertainties associated with mobility and material collection. This is compounded by the complexities involved in optimizing these activities. We have selected to solve the problem with reinforcement learning (RL) because it can tackle uncertainty and optimization. To ensure the safety of time and resources on the rovers, we propose a constrained reinforcement learning approach to generate safe actions while adhering to critical constraints to time and resources and applying optimizations on time consumption. To solve the task planning problem with time and resource constraints and uncertainty, we demonstrate a new Markov Decision Process (MDP) to illustrate the states, constraints, uncertainties, and the decision amount on actions. The novelty of our method lies in a three-level progressive constraint and uncertainty handling mechanism: the action mask for hard feasibility constraints, the Lagrangian method for redundancy, and rewards with penalties providing experience against uncertainty. These guarantee safety and peak reduction in consumption on rovers. Experiments show that the reinforcement learning planner can complete task planning, significantly outperforming baseline methods, even with an exceeded target without violating any constraints. Moreover, our planner can plan within seconds. In contrast, the Expressive Numeric Heuristic Search Planner cannot solve the original problem without uncertainty in 1h. Still, our approach can solve the simplified problem with 99.82% less time consumption than the planner above, whose plan also consumes twice the time and resources. This research enhances the efficiency and safety of rover task planning and contributes to developing autonomous lunar base construction and further exploration tasks.
| Original language | English |
|---|---|
| Article number | 115750 |
| Journal | Applied Soft Computing |
| Volume | 202 |
| DOIs | |
| Publication status | Published - Oct 2026 |
Keywords
- Lunar rover
- Reinforcement learning
- Task planning
- Uncertainty
Fingerprint
Dive into the research topics of 'A constrained reinforcement learning approach for strict constraint task planning on lunar rovers with uncertain time and resource consumption'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver