← Latest papers
💻 computer science

OpenBioRQ: Unsolved Biomedical Research Questions for Agents

The paper introduces OpenBioRQ, a novel retrieval-grounded agentic benchmark comprising 12,553 unsolved biomedical research questions that exposes critical failure modes in current models—such as citation hallucination and tool collapse—while demonstrating that even frontier agents struggle to solve a significant portion of these open-ended, non-saturating tasks.

Original authors: Minbyul Jeong

Published 2026-06-23
📖 5 min read🧠 Deep dive

Original authors: Minbyul Jeong

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a student taking a very difficult exam. In most medical exams today, the teacher gives you a question and a hidden answer key. If you get the answer right, you pass. If you get it wrong, you fail. But what happens when the teacher asks a question that no one knows the answer to yet? How do you grade a student when there is no answer key?

This paper introduces OpenBioRQ, a new way to test AI "agents" (smart computer programs that can search the internet and read papers) on these unsolved medical mysteries.

Here is the breakdown of what the researchers found, using simple analogies:

1. The Problem: The "Fake Link" Trap

Imagine you are writing a report. You write a claim, like "This new drug cures headaches," and you attach a link to a famous medical journal article to prove it.

  • The old worry: People thought AI would make up fake links (like linking to a website that doesn't exist).
  • The new discovery: The researchers found that the AI is actually very good at finding real links. Almost 100% of the links it gives actually work.
  • The real trap: The AI often links to a real article, but it's the wrong article. It's like linking to a recipe for "Chocolate Cake" when you are trying to prove that "Spicy Tacos cure headaches." The link works, the paper exists, but it has nothing to do with your claim.

The paper calls this the "Wrong-Paper" failure. About 16% of the time, the AI gives you a real, working link that completely fails to support what it just said. This is dangerous because it looks trustworthy at first glance.

2. The New Test: The "Unsolved Mystery" Box

Most medical tests for AI are like a crossword puzzle where the answers are already known. The AI just has to remember the right word.

  • OpenBioRQ is different. It's like a box of 12,500 unsolved medical mysteries.
  • The AI has to act like a detective: it must use tools to search databases, read thousands of papers, and try to figure out the answer.
  • Because there is no answer key, the researchers don't grade the AI on getting the right answer. Instead, they grade it on how honest and careful it is.
    • Did it admit, "We don't know yet"? (This is good).
    • Did it make up a confident answer with a fake link? (This is bad).
    • Did it find a real link that didn't actually prove its point? (This is the "Wrong-Paper" trap).

3. The "Tool Collapse" Phenomenon

The researchers gave the AI a set of "tools" (like a magnifying glass, a library card, and a search engine) to help solve these mysteries.

  • They expected the AI to use these tools to find the best evidence.
  • The surprise: On the hardest questions, the AI sometimes just gave up using its tools. It stopped searching and just guessed based on what it remembered from its training.
  • It's like a detective who, when faced with a really tough case, decides to stop looking for clues and just guess the culprit's name.
  • The researchers found that for some models, giving them the tools didn't actually help them get a better score. They were "collapsing" and ignoring the very things they were supposed to use.

4. The "Checklist" Solution

How do you grade a test where there is no right answer?

  • The researchers created a frozen checklist for every single question.
  • Instead of asking an AI, "Is this a good answer?" (which is vague), they gave it a specific list of things to check, like:
    • "Did you mention the specific protein?"
    • "Did you admit that we don't have a cure yet?"
    • "Did you avoid making up a fake clinical trial?"
  • This turned a vague judgment into a simple "Yes/No" checklist, making the grading much fairer and more consistent.

5. The Results: Who Passed?

The researchers tested several top-tier AI models (the "frontier agents").

  • The Difficulty: Even the smartest AI models only solved about 30% to 60% of these unsolved questions. The rest remained mysteries, which is exactly what you want in a hard test (it means the test isn't too easy).
  • The Citation Issue: Even the best models still fell into the "Wrong-Paper" trap about 10–16% of the time. They found real papers, but they didn't actually support their claims.
  • The Takeaway: Just because an AI gives you a link that works doesn't mean the information is correct. In the world of medical research, existence is not correctness.

Summary

OpenBioRQ is a new, tough test for medical AI. It stops asking "Do you know the answer?" and starts asking "Can you handle the unknown without lying?"

The main lesson is that AI is getting very good at finding real sources, but it is still very bad at making sure those sources actually prove what it claims. It's like a student who can find the library book but keeps citing the wrong page. Until AI learns to stop doing that, we can't fully trust it to summarize medical research, especially for questions we don't have answers to yet.

Important Note: The authors explicitly state this is a research tool to help scientists understand how AI works. It is not a tool for doctors to make decisions about patient care. You should not use these AI answers to treat a sick person.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →