SIREN (Luring LLMs onto the Rocks): PAIR-Driven Preference Manipulation in Web-RAG Recommenders
This paper introduces SIREN, an automated attacker-judge framework that adapts the PAIR jailbreaking loop to systematically manipulate web-augmented LLM recommendation rankings by iteratively editing retrieved source content with a taxonomy of poisoning techniques, achieving high success rates in moving targeted entities to the top rank while controlling for retrieval variables.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are walking through a giant, magical library where a super-smart robot librarian helps you find the best things to do, eat, or see. You ask, "What are the top five restaurants in town?" and instead of just pointing to a list, the librarian hops onto the internet, reads hundreds of real websites in seconds, and then writes you a brand-new, personalized recommendation letter. This is how modern "Web-RAG" (Retrieval-Augmented Generation) systems work: they don't just guess; they go out, gather facts from the live web, and synthesize them into an answer. But here's the catch: because the librarian is so good at reading, they trust what they find. If a sneaky baker manages to slip a note into one of the websites the librarian is reading, the librarian might accidentally write that baker's shop as the number one spot in the whole city, even if it's not actually the best. This paper explores exactly how easy it is for a trickster to "poison" the librarian's reading list and hijack the final recommendation.
The researchers behind this study, led by Evan Caville and his team, built a digital "hacker-bot" named SIREN (which stands for "Luring LLMs onto the Rocks," a playful nod to the myth of sirens luring sailors to their doom). Their goal wasn't to invent fake restaurants or break into the robot's brain; instead, they wanted to see if they could take a real website that the robot was already planning to read, quietly edit its text, and trick the robot into ranking that specific business as the absolute #1 choice. They treated the robot like a judge in a contest and the website edits like a contestant trying to win by changing their costume just enough to look better, without changing the actual rules of the game.
SIREN works like a clever, iterative game of "guess and check." The attacker-bot picks a target business and a specific website about it. Then, it tries 23 different "tricks" (like hiding a fake ranking claim in the image description, or writing a fake "top 10" list right in the middle of the page). After making a tiny edit, the bot asks the robot librarian to generate a new recommendation. A second "judge-bot" checks the result: Did the target business move up the list? If not, the attacker-bot learns from the mistake, tries a different trick, and asks again. It keeps doing this up to 20 times per attempt, refining its edits until it either wins or runs out of tries.
The results were quite revealing. Across 124 different attempts using two different versions of the Claude robot (Haiku and Sonnet), SIREN successfully pushed a target business to the #1 spot in 50% of the trials. That means half the time, a simple edit to a single webpage was enough to completely flip the robot's recommendation. The study found that the most effective tricks weren't the ones that shouted, "Rank me first!" (which the robot sometimes spotted and ignored). Instead, the winners were the ones that looked like normal, helpful information—like a "seeded list" that said, "Here are the top 3 places," with the target business conveniently listed first, or a declarative claim that sounded like a fact.
Interestingly, the success of these tricks depended heavily on which robot was doing the reading. The "Haiku" model was much easier to trick, reaching the #1 spot in about 61% of its full tests, while the "Sonnet" model was tougher, only falling for the trick 26% of the time in the same tests. However, when the researchers took the successful tricks from one model and tried them on the other, the results were mixed; a trick that worked perfectly on one didn't always work on the other. This suggests that there isn't one single "magic spell" to break all AI recommenders; it depends on the specific model and the specific question being asked.
The team also checked if these tricks would work in the real world, outside of their controlled test lab. They took the 62 successful edits and tested them in fresh, new sessions without the "learning" loop. Surprisingly, 80.5% of those tricks still worked, proving that once you find a way to fool the robot, the trick often sticks. However, the study also noted that this was a controlled experiment where the list of websites the robot read was kept exactly the same. In a real-world scenario, the robot might choose different websites or filter them differently, so while the threat is real, the exact success rate in the wild might vary.
Ultimately, SIREN shows that as we rely more on AI to make decisions about what to buy or where to go, the "source code" of the internet becomes a new battlefield. If a business can edit its own website to include a cleverly disguised claim, it can potentially hijack the AI's opinion. The paper suggests that simply filtering out obvious "spam" isn't enough; we need smarter ways to check if a ranking claim is backed up by multiple sources, because the AI is currently very good at being fooled by a single, well-dressed lie.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.