The Phish, The Spam, and The Valid: Generating Feature-Rich Emails for Benchmarking LLMs
This paper introduces PhishFuzzer, an open-source framework that generates a large-scale, metadata-rich dataset of 23,100 diverse emails with strict three-class labels to rigorously benchmark the reliability and robustness of state-of-the-art LLMs in detecting phishing and spam under varying structural conditions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are the bouncer at a very exclusive, high-tech nightclub. Your job is to stop three types of people from getting in:
- The Phishers: Dangerous criminals wearing a fake disguise, trying to steal your wallet.
- The Spammers: Annoyance salespeople shouting about cheap watches and crypto scams.
- The Valid Guests: Your actual friends and colleagues who just want to come in and chat.
For years, the bouncers (email filters) have been using a simple checklist: "Does this person look suspicious? Do they have a bad name?" But recently, the criminals got a superpower: AI. They can now write perfect, polite, and convincing letters that trick the old checklists.
This paper is about building a brand-new, ultra-strict training program for the bouncers to see if they can handle these AI-powered criminals.
The Problem: The Old Training Manual is Outdated
The researchers realized that the "training manuals" (datasets) bouncers have been using are like old, dusty photos from the 1990s. They are missing details, the handwriting is messy, and they don't show how modern criminals use AI to sound like your best friend.
The Solution: "PhishFuzzer" (The Training Simulator)
The team built a machine called PhishFuzzer. Think of it as a Mad Libs generator on steroids.
- The Seeds: They took 3,300 real emails (some from their own inboxes, some from public archives) to use as "seeds."
- The Magic: They fed these seeds into a powerful AI. The AI didn't just copy them; it created 19,800 new, slightly different versions of each email.
- The Variations: It changed the length (short vs. long), the sender (real companies vs. fake ones that look real), and the wording, while keeping the core "trick" the same.
The result is a massive library of 23,100 emails that are perfectly labeled: "This is a Phish," "This is Spam," or "This is Valid." Crucially, they also added metadata—the digital equivalent of checking the guest's ID, seeing the car they drove in, and checking if they have a suspicious package in their hand.
The Test: Two Super-Bouncers
They took two of the smartest AI bouncers available today (Qwen and Gemini) and put them through the wringer. They asked them to sort the emails in two ways:
- The "Blind" Test: Just reading the letter (Subject + Body).
- The "Full Background Check" Test: Reading the letter plus checking the sender's ID, the links, and the attachments.
The Results: A Tale of Two Bouncers
1. The "Paranoid" vs. The "Relaxed"
- Gemini was like a paranoid bouncer. It was great at spotting the dangerous Phishers. It rarely let a criminal in. However, it was too strict. It often kicked out innocent friends (Valid emails) thinking they were criminals.
- Qwen was like a relaxed bouncer. It was very good at letting friends in, but it let a few more Phishers slip through the cracks.
2. The Magic of the "ID Check" (Metadata)
When they gave the bouncers the extra info (URLs, sender names, attachments):
- Good News: Both bouncers got much better at spotting the dangerous Phishers. The extra clues helped them see through the disguises.
- Bad News: They got worse at spotting the Spammers.
- The Analogy: Imagine a spam email is a guy selling fake watches. If you only read his speech, you know he's a scammer. But if you also see his fancy car and his "official" business card (metadata), the bouncer gets confused and thinks, "Well, he has a nice car, maybe he's legit?" The extra details made the spam look more like a normal business email, and the bouncers let them in.
3. The "Spam vs. Valid" Gray Area
The biggest problem wasn't the criminals; it was the Spam. The researchers found that even the smartest AIs struggle to tell the difference between "Annoying Marketing" (Spam) and "Legitimate Newsletters" (Valid). It's like trying to tell the difference between a guy shouting "BUY MY WATCH!" and a guy shouting "FREE TICKET TO THE CONCERT!" Sometimes, it depends entirely on who is listening. The AI got stuck in this gray area.
The Big Takeaway
This paper tells us two main things:
- AI is getting scary good at faking emails. Old filters can't catch them anymore. We need new tools that look at the whole picture (not just the words).
- More info isn't always better. While checking a guest's ID helps catch the criminals, it sometimes makes the bouncer too trusting of the annoying salespeople.
The authors have made their entire "training simulator" and the massive list of emails free for everyone to use. This allows security companies to build better, smarter bouncers that can finally handle the AI-driven future of email attacks.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.