Pre-Flight: A Benchmark for Evaluating Large Language Models on Aviation Operational Knowledge
This paper introduces "Pre-Flight," an open-source benchmark of 300 expert-authored multiple-choice questions designed to evaluate large language models on aviation operational knowledge, revealing that even the most advanced models currently fall significantly short of expert-level reliability and highlighting the necessity of domain-specific evaluation for responsible AI deployment in aviation.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the aviation industry as a massive, incredibly complex library. Inside, there are millions of books: some are shiny new digital manuals, while others are dusty, handwritten logs from the 1960s. This library contains the rules for how planes move on the ground, how to talk to air traffic control, and how to keep passengers safe.
For a long time, we've been trying to teach computers (specifically, Large Language Models or "LLMs") to read this library and act as helpful assistants. But here's the problem: we've been testing these computers with general trivia questions, like "Who was the first president?" or "What is the capital of France?"
The Problem: The "General Knowledge" Trap
The authors of this paper argue that passing a general trivia test doesn't mean a computer is ready to work in an airport. It's like testing a pilot's ability to fly a plane by asking them to recite the lyrics to a pop song. They might know the words perfectly, but that doesn't mean they can handle a storm or a mechanical failure. In aviation, a wrong answer isn't just a bad grade; it can cost money, reputation, or worse.
The Solution: "Pre-Flight"
To fix this, the authors created a new test called Pre-Flight. Think of this as a specialized "driving test" specifically for airport ground operations.
- The Test: It consists of 300 multiple-choice questions.
- The Source: The questions come from real, official rulebooks (like the ICAO international standards and US FAA regulations) and safety manuals used by real pilots and ground crews.
- The Graders: The questions were written and checked by real aviation experts—people who have actually managed air traffic, flown planes, and worked on the tarmac.
The Results: The "Expert Gap"
The authors took the smartest AI models available (up to mid-2026) and gave them this test. Here is what they found:
- The Human Benchmark: When a small group of real aviation professionals took a similar quiz, they scored about 95%. This is the "gold standard" of reliability.
- The AI Performance: The very best AI model managed to score 82.7%.
- The Gap: There is a significant gap between the AI and the human experts. Even though AI scores are getting better, they are improving very slowly. It's like a student who has been studying for two years and finally went from a 75% to an 83%, but still hasn't reached the 95% needed to graduate with honors.
Why is this happening?
The paper suggests a few reasons for the struggle:
- US Regulations are Tricky: The AI models struggled the most with specific US rules (FAA). This might be because those specific details are less common in the general internet data the AI was trained on compared to international rules.
- Reasoning vs. Memorization: The AI is good at remembering facts (like "What is the speed limit?"), but it gets confused when it has to do multi-step logic (like "If it's snowing and the runway is short, which gate should we use?").
- Overconfidence: Sometimes the AI is very confident in its wrong answers, which is dangerous in a high-stakes field like aviation.
The "Outlier" and the "Hard Mode"
- One model (ALLaM 2 7B) scored only 25.3%, which is basically guessing at random. This shows that a model can be great at other things but completely fail when faced with aviation-specific rules.
- The authors also mention they are building a "Hard Mode" version of the test that is even harder. They are keeping this private for now so that as AI gets smarter, they can keep testing it without the AI having "cheated" by memorizing the answers from the internet.
The Bottom Line
The paper concludes that before we let AI handle important, non-safety-critical jobs in aviation (like scheduling or customer service), we must prove it can pass this specific "Pre-Flight" test. Currently, the AI is not quite there yet. It's a promising student, but it's not ready to be the captain of the ship.
What the paper does NOT say:
- It does not say AI should be used for safety-critical tasks (like actually flying the plane or making emergency decisions).
- It does not claim that AI will replace human pilots or ground crews soon.
- It does not suggest that the current AI scores are good enough for immediate, unmonitored deployment in real-world operations.
The main message is simple: We need a specialized test for aviation, and right now, the AI is failing to reach the level of a human expert.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.