K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts
This paper introduces K-BrowseComp, a new web-browsing agent benchmark grounded in Korean contexts comprising 400 manually verified and synthetic problems, which reveals that even frontier and Korean-specific large language models struggle significantly with these tasks, achieving accuracy rates between 0% and 45.67%.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart, super-fast librarian who has read almost every book in the world. You ask them a simple question, and they answer instantly. But what happens when you ask them a tricky question that requires them to walk through a specific neighborhood, knock on specific doors, check a few different ledgers, and piece together clues that only exist in that local area?
That is exactly what this paper, K-BROWSECOMP, is about.
The Problem: The "Local Neighborhood" Gap
For a long time, we've tested AI by asking it to solve logic puzzles or follow instructions. The AI has gotten very good at that. But the real world isn't just about logic; it's about context.
The authors argue that while AI is great at the "global" internet (mostly English), it struggles when it has to navigate the Korean internet. It's like giving a brilliant chef a recipe written in a language they don't speak, using ingredients found only in a specific village market. Even if the chef knows how to cook, they might get lost trying to find the specific ingredient or misunderstand the local instructions.
Currently, there was no standard "test" to see how well AI could handle these specific, local Korean web searches. The existing tests were too easy or didn't require the AI to actually browse the web to find the answer.
The Solution: A New "Obstacle Course"
The researchers built a new benchmark called K-BROWSECOMP. Think of this as a specialized obstacle course for AI agents (robots that can browse the web).
- The Course: It contains 400 tricky questions.
- The Rules: To answer, the AI can't just guess or rely on what it memorized. It must:
- Go to the Korean web.
- Search for specific, hard-to-find facts (like "What is the title of the 13th poem in a specific poet's 10th book?").
- Connect clues from different websites.
- Give a single, correct answer.
They split the test into two parts:
- The Verified Set (300 questions): These were written and checked by real humans to ensure the answers are correct and the clues exist on public Korean websites.
- The Synthetic Set (100 questions): This is the clever part. The researchers used an AI to create new, even harder questions by looking at where other AIs failed. It's like a video game designer watching players get stuck on a level, then designing a new level specifically to trap them in the same way.
The Results: The AI Got Lost
The results were surprising. Even the most advanced AI models in the world (the "superstars" of AI) struggled mightily.
- The Global Giants: The best models (like GPT-5.5) only got about 45% of the verified questions right. That means they failed more than half the time.
- The Local Heroes: Korean AI models, which are supposed to be experts on Korean culture, did even worse, scoring between 0% and 10%.
- The Synthetic Trap: On the AI-generated "trap" questions, the best model only got 26% right.
Why Did They Fail? (The "Traffic Jam" Analogy)
The paper doesn't just say "they failed." It explains how they failed. Imagine the AI is a detective trying to solve a case.
- Getting the Clue, Losing the Thread: The AI often finds the right website (the right clue), but then it forgets the specific rule it was looking for. It's like finding the right house but forgetting to check the mailbox for the specific letter.
- Picking the Wrong Door: The AI finds a list of candidates (like a list of K-pop groups) but picks the first one that looks good, without checking if it fits all the rules.
- The "I Don't Know" Moment: Sometimes the AI finds the information but gets confused trying to write down the final answer, so it just gives up or guesses.
The authors found that the problem isn't that the AI can't find the information. The problem is that it can't keep track of all the different pieces of information while it's searching. It's like trying to solve a puzzle while someone keeps shuffling the pieces around; you find the right pieces, but you can't put them together in the right order.
The Takeaway
This paper is a wake-up call. It shows that just because an AI is smart and knows a lot of facts, it doesn't mean it can act like a helpful assistant in a specific local environment.
The researchers released this test and the code for free, hoping that developers will use it to build better "Korean web agents" that don't just search, but actually understand how to navigate, remember, and connect the dots in the Korean digital world.
In short: We built a difficult maze for AI to run through in a Korean context. Even the fastest runners got lost because they couldn't keep their bearings, not because they couldn't see the path.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.