Falcon: Functional Assembly and Language for Compositional Reasoning in X-ray
This paper introduces Falcon, a multimodal framework that addresses the limitations of object-centric vision-language models in X-ray baggage screening by formalizing threat detection as compositional reasoning over spatially dispersed components, supported by the new Falcon-X benchmark for evaluating structured safety inference.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a security guard looking at an X-ray of a suitcase. Most computer programs today act like a very strict librarian: they look at the image and say, "I see a battery," "I see a knife," or "I see a bottle." They are great at spotting individual items.
But here is the problem: A single battery is harmless. A single detonator is harmless. A pile of explosives is just a pile of chemicals. The real danger only appears when these specific items are together and connected in a way that makes them work as a bomb.
Current AI struggles with this. It sees the parts but misses the "story" of how they fit together. It's like a child who can name every piece of a Lego set but doesn't understand that if you snap the red block to the blue block, you've built a car.
Enter "Falcon."
Falcon is a new AI system designed specifically to solve this "puzzle" problem in X-ray baggage screening. Instead of just listing items, Falcon acts like a detective who understands how things work together.
Here is how it works, broken down into simple steps:
1. The "Safety State" (The Detective's Notebook)
Most AI models try to guess the answer directly from the picture. Falcon takes a different approach. Before it answers a question, it fills out a structured "notebook" (which the paper calls a Structured Safety State).
In this notebook, it writes down three things:
- Who is there? (Is there a battery? A detonator? Explosives?)
- Do they fit? (If the battery is there, does it look like it could connect to the detonator?)
- How dangerous is this? (Based on who is there and how they fit, what is the risk score?)
Think of this like a chef checking a recipe. Before saying "This cake is ready," the chef checks: "Do I have flour? Yes. Eggs? Yes. Are they mixed? Yes. Okay, now I can say it's a cake." Falcon does this safety check before it speaks.
2. The "Functional Grounding" (Finding the Team, Not Just the Players)
If you ask a normal AI, "Where is the battery?" it points to the battery.
If you ask Falcon, "Which items in this bag could form a bomb?" it doesn't just point to one thing. It draws a circle around the battery, the detonator, and the explosives, and says, "These three are the team."
It understands that the danger isn't in the individual objects, but in the relationship between them. It can even tell you what is missing. If it sees a battery and a detonator but no explosives, it says, "I see two parts of a bomb, but the main explosive part is missing. The risk is high, but not 100% yet."
3. The "Falcon-X" Benchmark (The Training Ground)
To teach Falcon this skill, the researchers created a new dataset called Falcon-X.
- Old Datasets: Were like a photo album of "Bad Guys" (guns, knives, scissors).
- Falcon-X: Is like a collection of "Disassembled Puzzles." It contains thousands of X-rays where bomb parts are scattered, hidden under other clothes, or mixed with clutter. It forces the AI to learn that a battery hidden under a shirt is still a battery, and if it's near a detonator, it's a problem.
4. The Results: Why It Matters
The paper tested Falcon against other smart AI models.
- The Competition: Other models were good at saying, "That looks like a battery." But when asked, "Is this a bomb?" they often got confused or hallucinated (made things up).
- Falcon: Because it forces itself to check the "relationships" first, it is much better at spotting the functional threat. It can look at a messy, cluttered bag and say, "These three specific items, even though they are far apart, could work together to make a bomb."
The Bottom Line
The paper claims that by forcing the AI to stop and think about how parts connect (compositional reasoning) rather than just what parts exist (object detection), we can build much safer security systems. Falcon doesn't just see the pieces; it understands the machine they could build.
In short: Falcon is an AI that doesn't just count the bricks; it knows when the bricks are stacked in a way that could bring the whole building down.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.