TY - JOUR
T1 - TopoAgent
T2 - A Constraint-Structured Reinforcement Learning Agent for Heterogeneous Satellite Mission Scheduling
AU - Ren, Yi
AU - Liu, Shuyi
AU - Chen, Xiao
AU - Gao, Yuan
AU - Zhang, Zeyu
AU - Li, Ruide
N1 - Publisher Copyright:
© 2026 by the authors.
PY - 2026/6
Y1 - 2026/6
N2 - With more satellites, richer payload resources, and more diverse service functions, satellite systems are increasingly operated as large space–ground networks. These networks must schedule arriving missions under changing topology, gateway access, beam availability, weather-affected links, spectrum compatibility, and mission time windows. Offline optimization can compute high-quality schedules when the mission set, satellite visibility windows, and resource states are known before execution, but repeated replanning is costly for asynchronous arrivals. Online heuristics make faster decisions from local route rules, but they do not evaluate how an accepted service path changes the capacity left for later requests. Reinforcement-learning schedulers can adapt from delayed scheduling outcomes. However, many generic policies rely on fixed-step state updates or flat compound-action scores, whereas online satellite scheduling makes decisions at irregular arrivals over continuously evolving topology and capacity-coupled service paths. We propose TopoAgent, an online reinforcement-learning agent for heterogeneous satellite mission scheduling. TopoAgent models each request as a service-path decision, propagates compound feasibility through the satellite–gateway–beam hierarchy, and uses a capacity-aware policy to choose among feasible paths. A deterministic constraint manager places the selected path in time, while SRV guides the policy toward assignments that preserve reusable beam capacity. In a high-fidelity simulator, TopoAgent achieves a 74.7% mission completion rate and a 75.5% high-priority completion ratio over five seeds.
AB - With more satellites, richer payload resources, and more diverse service functions, satellite systems are increasingly operated as large space–ground networks. These networks must schedule arriving missions under changing topology, gateway access, beam availability, weather-affected links, spectrum compatibility, and mission time windows. Offline optimization can compute high-quality schedules when the mission set, satellite visibility windows, and resource states are known before execution, but repeated replanning is costly for asynchronous arrivals. Online heuristics make faster decisions from local route rules, but they do not evaluate how an accepted service path changes the capacity left for later requests. Reinforcement-learning schedulers can adapt from delayed scheduling outcomes. However, many generic policies rely on fixed-step state updates or flat compound-action scores, whereas online satellite scheduling makes decisions at irregular arrivals over continuously evolving topology and capacity-coupled service paths. We propose TopoAgent, an online reinforcement-learning agent for heterogeneous satellite mission scheduling. TopoAgent models each request as a service-path decision, propagates compound feasibility through the satellite–gateway–beam hierarchy, and uses a capacity-aware policy to choose among feasible paths. A deterministic constraint manager places the selected path in time, while SRV guides the policy toward assignments that preserve reusable beam capacity. In a high-fidelity simulator, TopoAgent achieves a 74.7% mission completion rate and a 75.5% high-priority completion ratio over five seeds.
KW - action masking
KW - constrained decision making
KW - event-driven scheduling
KW - reinforcement learning
KW - resource allocation
KW - satellite scheduling
UR - https://www.scopus.com/pages/publications/105041460351
U2 - 10.3390/electronics15112456
DO - 10.3390/electronics15112456
M3 - Article
AN - SCOPUS:105041460351
SN - 2079-9292
VL - 15
JO - Electronics (Switzerland)
JF - Electronics (Switzerland)
IS - 11
M1 - 2456
ER -