Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents
This paper introduces MissionBench, a benchmark for evaluating the zero-shot mission-level capabilities of multimodal large language models in aerial 3D environments, revealing that while scaling improves performance, current models still significantly lag behind humans in complex, long-horizon embodied tasks requiring multi-step planning and adaptive reasoning.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you've just handed a brand-new, super-smart robot a map and a mission: "Go find that red car and tell me its license plate." You expect the robot to zoom off, scan the scene, and come back with the answer. But here's the catch: this robot has never flown a drone before. It hasn't been taught how to hover, how to dodge a tree, or how to read a sign from a bird's-eye view. It only knows how to talk and look at pictures because it was trained on the entire internet. This is the world of Multimodal Large Language Models (MLLMs)—AI brains that can see images and read text, now being asked to control physical machines. The big question scientists are asking is: Can these general-purpose "smart brains" figure out how to fly a drone and solve a complex mission just by reading a single sentence, without any special training? It's like asking a person who has read every book in a library to suddenly pilot a spaceship to a specific asteroid just because they read a story about it.
In this paper, the researchers decided to put these AI brains to the ultimate test. They built a video game world called MissionBench, which is basically a giant, high-tech playground for drones. Inside this playground, there are 120 different missions, ranging from "Patrol this neighborhood" to "Inspect a burning ship" or "Drop a package on a tent." The twist? The AI agents (the drone pilots) are given only a text instruction and a camera view from the drone's perspective. They have to figure out where to go, how to turn, how fast to fly, and what to report back, all on their own. It's a closed-loop game: the AI moves, the world changes, the AI sees the new view, and it has to decide what to do next.
The results were a mix of "wow" and "whoops." The researchers tested 22 different AI models, from open-source ones to the most powerful commercial giants. The best-performing AI managed to succeed on fewer than 35% of the missions. To put that in perspective, a human pilot playing the same game with the same limited controls succeeded 70% of the time, and with full keyboard control, humans hit 84.4%. So, while the AI isn't terrible, it's still struggling to keep up with a human. The paper suggests that simply making the AI "bigger" (adding more data and parameters) helps; the largest models did significantly better than the smaller ones, suggesting that raw size gives them a bit more "embodied" intuition. However, even the smartest AI still gets confused. It often crashes into things, flies in circles, or gives up too early, thinking it's finished when it's actually still miles away from the target.
The study also looked at why the AI fails. It turns out that being good at spotting an object in a single picture (like finding a car in a photo) doesn't guarantee the AI can fly the drone to that car. The AI needs to do a lot more than just "see"; it needs to plan a path, adjust its height, keep the target in view, and adapt when things go wrong. The researchers found that the AI often suffers from "drift and oscillation"—it gets stuck going back and forth like a confused moth—or "premature termination," where it declares victory after just a few seconds because it got tired or confused.
Ultimately, this paper suggests that while general-purpose AI is getting better at controlling drones without special training, it's not ready for the real world yet. The gap between what the AI can do and what a human can do is still huge. The authors warn that as these models get bigger and smarter, they might become more capable, but they also carry risks if deployed too soon. For now, MissionBench serves as a reality check: having a brain that knows everything about the world doesn't automatically mean you know how to fly a drone through it. The journey from "reading about flying" to "actually flying" is still a long one.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.