Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions
This survey systematically reviews 74 studies to establish an 11-category taxonomy for agentic and generative AI in OSINT, identifies a critical gap between the recognized risk of hallucinations and their empirical measurement, maps current research strengths and weaknesses across the intelligence lifecycle, and proposes a ten-point agenda advocating for a human-in-the-loop co-pilot model as the most viable near-term deployment strategy.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a massive mystery. The clues are scattered everywhere: millions of social media posts, news articles, hidden forums, and encrypted messages. In the past, you had to read every single one of these clues by hand. It was slow, exhausting, and you often missed things.
Now, imagine you have a super-smart robot assistant (an AI) that can read all those clues in seconds, connect the dots, and write a report for you. This paper is a massive review of 74 different studies trying to figure out if this robot assistant is ready for the job.
Here is the simple breakdown of what the researchers found, using some everyday analogies:
1. The Robot is Fast, But We Don't Know If It's Honest
The paper says the robot (called "Agentic AI" or "Generative AI") is incredibly good at gathering information and organizing it. It's like a super-fast librarian who can find books you didn't even know existed.
However, there is a huge problem: We don't really know how often the robot makes things up.
- The "Hallucination" Problem: In the world of AI, "hallucination" means the robot confidently tells you a lie that sounds true.
- The Finding: The researchers looked at 74 studies. More than 20 of them admitted, "Hey, this robot might lie." But only one study actually tried to measure how often it lies in a real OSINT (Open Source Intelligence) situation. That one study found a 4% lie rate, but it was done under perfect, easy conditions.
- The Analogy: It's like having a student who claims they can pass a math test. 20 people say, "I think they might cheat," but only one person actually gave them a test. The rest just assumed they were honest without checking.
2. The Robot is Great at Starting, But Bad at Finishing
The researchers mapped out the detective's workflow:
- Gathering clues (Collection)
- Reading and sorting them (Analysis)
- Checking if they are true (Verification)
- Writing the final report (Reporting)
- Making the final decision (Decision Support)
The Finding: The robot is excellent at steps 1 and 2. It can gather and sort clues better than any human. But it is almost completely ignored for steps 3, 4, and 5.
- The Analogy: It's like hiring a robot to do the heavy lifting of moving furniture (gathering and sorting), but then expecting a human to do the delicate work of arranging the art on the walls and deciding where the sofa goes. The paper says we are over-focusing on the heavy lifting and ignoring the delicate, critical parts where mistakes matter most.
3. The "Test Score" vs. The "Real Job"
There is a big disagreement in how we test these robots.
- The Multiple-Choice Test: Some studies test the robot with simple multiple-choice questions (like a trivia quiz). The robot gets a 91% score, beating human experts.
- The Real Job Test: Other studies test the robot by giving it a messy, real-world case where it has to read hundreds of documents, find hidden connections, and write a report. In this test, the robot fails. It can't handle the complexity or the uncertainty.
- The Analogy: It's like a driver who gets a perfect score on a driving simulator (the multiple-choice test) but crashes the moment they get behind the wheel in real traffic with rain and pedestrians (the real job). The paper warns us not to be fooled by the high test scores.
4. The "Poisoned Well" Risk
The paper highlights a scary risk: What if the clues the robot finds are actually fake?
- The Finding: Researchers showed that bad actors can use similar AI tools to write fake news or fake threat reports. When the robot reads these fake reports, it believes them and adds them to its "knowledge base."
- The Analogy: Imagine your robot assistant is gathering water from a river to drink. But someone has secretly poisoned the river upstream. The robot doesn't know the water is poisoned; it just drinks it and tells you, "This is great water!" The paper says the robot has no built-in filter to tell the difference between real water and poisoned water.
5. The Solution: The "Co-Pilot" Model
So, should we fire the robot? No. The paper says we should keep it, but change how we use it.
- The Recommendation: We should use a "Human-in-the-Loop" model. Think of it like a Co-Pilot in an airplane.
- The Robot (Co-Pilot) does the boring, heavy work: gathering data, sorting files, and writing the first draft.
- The Human (Captain) sits in the seat, checks the robot's work, verifies the facts, and makes the final decision.
- Why? Because the robot is fast but prone to lying and being tricked. The human is slower but has the judgment to spot the lies. Until we can prove the robot is 100% reliable, the human must always be the one holding the controls.
Summary
The paper concludes that while AI is a powerful new tool for intelligence work, it is currently too risky to use on its own. It's like a powerful engine that hasn't been fully safety-tested yet. We need to build better safety checks (tests for lying, tests for fake news, and legal rules) before we let the robot drive the car alone. For now, the best approach is a partnership: the robot does the heavy lifting, and the human makes sure it's telling the truth.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.