GEO-Flag: Detecting and Measuring GEO-Optimized Web Content
This paper introduces GEO-Flag, a comprehensive framework featuring the GEOFlagBench benchmark and an Intervention-Paired Training method to effectively detect, audit, and measure the prevalence of Generative Engine Optimization (GEO) in real-world search ecosystems.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The internet is changing how we find answers. For decades, we searched by typing a question and scanning a list of links, clicking through to read the source ourselves. Now, a new generation of search tools uses artificial intelligence to read those links for us, synthesizing the information into a single, direct answer with citations. This shift promises convenience, but it also creates a new vulnerability. Just as websites were once tweaked to rank higher in traditional search lists, a new practice has emerged where content is deliberately altered to be picked up and cited by these AI systems. This practice, known as generative engine optimization, can make weak or false information appear well-supported because the AI presents it as a fact without showing the user the full list of competing sources. The danger lies in the opacity: if the AI cites a page that was strategically built to be cited, the user sees a confident answer, not the manipulation behind it.
A team of researchers set out to understand the scale of this problem and to build tools to measure it. They began by creating a massive test set of 3,200 webpages covering health, finance, technology, and travel. These pages included original human writing, text polished by artificial intelligence, and content generated entirely by machines. Crucially, they also created versions of these pages that had been specifically modified using eight different strategies designed to game the new AI search engines. With this diverse collection, they tested existing methods for spotting these optimized pages. They found that while some tools could flag the content with high accuracy overall, they were often relying on shortcuts. Many detectors were simply guessing based on whether the text was written by a human or an AI, rather than identifying the specific changes made to manipulate the search engine. When the researchers tested these tools on pages where humans had carefully applied optimization techniques, the detectors often failed, revealing that their previous success was based on a misunderstanding of what they were actually detecting.
To solve this, the researchers developed a new training method called intervention-paired training. Instead of just showing the computer examples of "good" and "bad" pages, they showed it pairs of pages: the original version and the modified version. They taught the system to recognize that a specific change, like adding a certain phrase to boost visibility, should increase the suspicion score, while a general improvement in grammar or style should not. This approach forced the detector to focus on the actual act of optimization rather than the author's identity. The result was a significant improvement in reliability. The new system correctly identified optimized pages with much higher accuracy and, importantly, did not get confused by whether the original text was written by a person or a machine. It learned to spot the manipulation itself.
The team then built a complete auditing system that goes a step further. Once a page is flagged as potentially optimized, the system checks the links cited on that page to see if they are trustworthy. It assigns a tier to the publisher, ranging from highly authoritative institutions to sources that are easy to create or edit, and checks if the links are still working. When they applied this entire pipeline to real-world search results from Google and a new AI-powered search tool, they found that roughly 8.9% of the pages returned were optimized for generative search. This number was not static; for pages that had been modified in 2026, the rate jumped to 16.36%. The study also revealed a stark difference in the quality of sources used. Pages returned by the AI search tool relied heavily on citations from lower-tier sources, with nearly three-quarters of those citations receiving a low verifiability label, compared to less than half for traditional search results.
These findings suggest that the landscape of online information is shifting rapidly. The researchers estimate that more than one in ten pages currently appearing in these new search results may have been strategically altered to influence the AI. The problem is not just that false information is being spread, but that the very mechanism we use to verify facts—the citation—is becoming part of the manipulation. By identifying these patterns and measuring their prevalence, the researchers have provided a way to see behind the curtain of the new search era. Their work does not claim to have solved the issue or to have found every instance of manipulation, but it establishes a clear foundation for measuring the problem. It shows that without new tools to detect these specific optimizations, users may increasingly encounter answers that look authoritative but are built on a foundation of strategically constructed, and often unverified, sources.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.