← Latest papers
🤖 AI

LiveBrowseComp: Are Search Agents Searching, or Just Verifying What They Already Know?

This paper reveals that current LLM-based search agents often rely on intrinsic knowledge rather than genuine web searching, prompting the introduction of LiveBrowseComp, a dynamic benchmark using recent, non-salient facts to effectively evaluate true evidence-driven discovery capabilities.

Original authors: HuiMing Fan, Xiao Wang, Zheng Chu, Qianyu Wang, Zhuoyao Wang, Ming Liu, Bing Qin, XingYu

Published 2026-05-28
📖 5 min read🧠 Deep dive

Original authors: HuiMing Fan, Xiao Wang, Zheng Chu, Qianyu Wang, Zhuoyao Wang, Ming Liu, Bing Qin, XingYu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Question: Are They Exploring or Just Checking Their Homework?

Imagine you hire a super-smart detective (an AI agent) to find a specific, obscure fact. You give them a library card (internet access) and tell them, "Go find the answer."

The paper asks a suspicious question: Is this detective actually searching the library to find new information, or are they just opening a book they already memorized, finding the answer in their head, and then using the library card only to say, "See? I told you I was right"?

The authors call this behavior Intrinsic Knowledge Dependence (IKD). It's like a student who has already memorized the answers to a practice test. When they take the real test, they don't need to study; they just write down what they remember and use the "open book" rule to double-check that their memory is correct.

The Problem with Current Tests (The "Static" Benchmarks)

Currently, we test these AI detectives using "Static Benchmarks" (like BrowseComp). Think of these as old, dusty trivia quizzes.

  • The Issue: Because these quizzes have been around for a while, the AI models have likely already "read" the questions during their training.
  • The Result: The AI gets a high score, but it's not because it's good at searching. It's because it's good at remembering. It guesses the answer from its memory and uses the internet just to feel confident.

The Three "Lie Detector" Tests

To prove their suspicion, the researchers ran three simple experiments on existing AI models:

  1. The "No-Tools" Test: They took away the internet.
    • Result: The AIs still got about 44% of the answers right. This proved they were relying on their internal memory, not the search tool.
  2. The "Broken Library" Test: They let the AI use the internet, but they secretly removed all the pages that actually contained the correct answer. The AI could only see irrelevant junk.
    • Result: The AIs' scores crashed. Instead of ignoring the junk and sticking to their memory, they got confused and gave wrong answers. This showed they weren't using the search to discover truth; they were using it to confirm what they already guessed.
  3. The "Trail of Clues" Test: They watched the AI's search history.
    • Result: They found that over 50% of the search questions the AI asked were based on its own guesses, not on things it actually found on the web. It was chasing its own tail.

The Solution: Introducing "LiveBrowseComp"

To fix this, the researchers built a new test called LiveBrowseComp.

The Analogy:
If the old tests were like a Jeopardy! episode from last year (which the AI might have seen on TV), the new test is like a live news broadcast happening right now.

  • Freshness: The questions are based on facts published in the last 90 days. The AI couldn't have memorized these because they didn't exist when the AI was "born" (trained).
  • Obscurity: The facts aren't famous headlines (like "Who won the Super Bowl?"). They are "long-tail" facts, like "What was the specific earthquake magnitude in a small town in Peru three weeks ago?" or "What is the name of a niche video game released yesterday?"
  • The Goal: To force the AI to actually hunt for information, rather than just guessing from memory.

What Happened When They Ran the New Test?

The results were shocking:

  1. Memory Fails: When the AI tried to answer these new questions without using the internet (Closed-Book), their score dropped to less than 2%. They knew almost nothing about these fresh facts.
  2. Search Scores Dropped: Even when allowed to use the internet, the AI scores dropped by 25–40 points compared to the old tests.
  3. Rankings Changed: The AI models that were "winners" on the old tests (because they had good memories) suddenly became average or poor on the new test. The models that were good at actually searching rose to the top.

The Human Comparison

The researchers also asked real humans to solve these questions.

  • The Finding: Humans took about the same amount of time to solve the "Old" questions and the "New" questions.
  • The Meaning: The new questions aren't "harder" or "smarter." They are just new. The only reason the AI struggled was that it couldn't rely on its memory shortcut.

The Bottom Line

The paper concludes that current AI evaluations are flawed because they reward memory instead of searching.

  • Old Way: "Can you guess the answer and then find a website to back it up?" (AI gets an A).
  • New Way (LiveBrowseComp): "Can you find an answer you have never seen before?" (AI gets an F).

The authors argue that to truly test if an AI is a good "Search Agent," we must test it on things it doesn't already know, forcing it to do the hard work of discovery rather than just verification.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →