← Latest papers
💻 computer science

ER-Curriculum-DDQN for Critical Task Offloading in UAV-MEC via Transparent LEO Relays

This paper proposes ER-Curriculum-DDQN, a deep reinforcement learning framework that integrates feasibility masking, critical task pressure features, and reservation-guided curriculum learning to optimize online task offloading in UAV-MEC systems via transparent LEO relays, thereby maximizing the timely completion ratio of critical tasks under strict service window constraints.

Original authors: Ziang Zhang, Zhenjiang Zhang, Jiaxing Du, Zhaoyuan Liu, Yujie Wang

Published 2026-09-23
📖 5 min read🧠 Deep dive

Original authors: Ziang Zhang, Zhenjiang Zhang, Jiaxing Du, Zhaoyuan Liu, Yujie Wang

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the vast, silent stretches of the sky where the ground's infrastructure fades away, a new kind of network is taking shape. Imagine a fleet of drones, or unmanned aerial vehicles, hovering over a disaster zone or a remote mountain range, carrying the weight of urgent data from sensors and cameras. These drones act as flying computers, but they have limits. To solve complex problems, they need to send their heavy data to powerful servers on the ground or in the clouds. When the ground is too far or the roads are blocked, these drones look up to the stars for help. They connect to low-orbit satellites that act as transparent mirrors, simply bouncing signals between the drone and a distant gateway without stopping to process the information themselves. This creates a fleeting bridge, a narrow window of time where the drone, the satellite, and the ground station are all aligned. The challenge is that this window is short and unforgiving. If a drone tries to send a massive, non-urgent file through this bridge, it might lock up the entire connection for the duration of the transaction. If a critical emergency message arrives while that bridge is occupied, it cannot be sent until the first task is finished, potentially missing a life-saving deadline. The core question for scientists is how to manage this scarce, time-limited connection so that the most important jobs always get through, even when the system is crowded.

Researchers at Beijing Jiaotong University and the Army Arms University of the PLA have tackled this problem by designing a smart decision-making system for these flying networks. They focused on a scenario where twenty drones are working together, sharing two ground-based computing centers and two satellite paths. The system must decide, the moment a new task arrives, whether to process it locally on the drone, send it to a nearby ground server, or forward it through the satellite. The catch is that the satellite path operates on a strict "first-come, first-served" rule that cannot be interrupted once started. If a routine task is admitted to the satellite path, it reserves the entire connection for its journey, blocking any other task, no matter how urgent, until the round trip is complete. The researchers needed a way to predict when to say "no" to a routine task to save the satellite path for a future emergency, all without knowing what tasks would arrive next.

To solve this, the team developed a new method called ER-Curriculum-DDQN. This is a type of artificial intelligence that learns by trial and error, but with a specific set of rules to keep it grounded in reality. The system uses a "feasibility mask," which acts like a filter that instantly removes any impossible choices. For example, if the satellite window is closed or the local drone memory is full, the system simply cannot consider those options. This prevents the AI from wasting time thinking about actions that would fail due to physical constraints. Beyond just knowing what is possible, the system tracks a "pressure" feature. This is a measure of how much demand there has been for the satellite path from critical tasks in the recent past. If the system senses that critical tasks are arriving frequently, it becomes more cautious about letting routine tasks use the satellite, effectively saving the bridge for the emergencies that might come next.

The learning process itself was guided by a "curriculum," a teaching strategy that starts simple and gets harder. In the early stages of training, the AI was explicitly told to feel a penalty whenever it let a routine task use the satellite path if that action increased the pressure on the system. This taught the AI the value of holding back resources. As the training progressed, this explicit penalty was gradually faded away, forcing the AI to internalize the lesson and make the right decisions based on the patterns it had learned, rather than following a simple rule. The goal was not just to get tasks done, but to ensure that the most critical ones were finished on time.

The researchers tested their system through extensive computer simulations, pitting it against other known methods and simple rule-based approaches. They varied the conditions, changing how many tasks arrived per second, how long the satellite windows lasted, and how many of the tasks were critical emergencies. In every scenario, especially when the system was under heavy load or the satellite windows were very short, the new method outperformed the others. It achieved the highest rate of on-time completion for critical tasks. For instance, when the number of incoming tasks was high and the satellite windows were tight, the new system managed to complete critical tasks about 69% of the time, while other methods struggled significantly lower. The results showed that by combining a strict filter for impossible moves with a learned sense of future demand, the system could protect the most important data without needing to know the future.

The study confirms that in these high-stakes, time-sensitive environments, the best strategy is not just to react to what is happening right now, but to understand the scarcity of the connection itself. The researchers found that simply prioritizing critical tasks by giving them higher weights was not enough; the system had to understand the cost of occupying the satellite path with a routine task. By using a learning approach that respects the physical limits of the network and learns from the history of demand, the drones can make smarter choices. This work does not claim to have solved every problem in space communications, nor does it promise to work in every possible future scenario, but it demonstrates a clear and effective way to manage the delicate balance between routine work and urgent needs in a network where time is the most valuable resource.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →