← Latest papers
🤖 AI

EpiBench: Verifiable Evaluation of AI Agents on Epigenomics Analysis

EpiBench is a verifiable benchmark for short-horizon epigenomics analysis that evaluates AI agents on 106 realistic workflow tasks across four assay types, revealing that even top-performing models like GPT-5.5 fail to achieve a majority success rate due to difficulties in applying deep, assay-specific scientific judgment despite often identifying correct files and intermediate results.

Original authors: Harihara Muralidharan, Reema Baskar, Soo Hee Lee, Tim Proctor, Kenny Workman

Published 2026-06-12
📖 3 min read☕ Coffee break read

Original authors: Harihara Muralidharan, Reema Baskar, Soo Hee Lee, Tim Proctor, Kenny Workman

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a team of very smart, highly trained robot assistants. You want to see if they can act like expert scientists when analyzing complex biological data. Specifically, you want to know if they can look at a messy pile of genetic "receipts" (data from experiments) and figure out the correct story the data is telling, just like a human expert would.

This paper introduces a new test called EpiBench to see how good these AI robots are at that specific job.

The Test: A "Pop Quiz" for AI Scientists

Think of EpiBench as a series of 106 pop quizzes. Each quiz gives the AI a snapshot of a real scientific experiment that is halfway done. The AI has to:

  1. Look at the files and data provided.
  2. Make a specific decision (like "Which samples should we compare?" or "What is the correct number here?").
  3. Give a final answer.

The test covers four main types of genetic experiments (like checking how DNA is packaged, how genes are turned on, or chemical tags on DNA). The answers are graded automatically: it's either right or wrong. There is no "close enough."

The Results: The Robots Are Still Learning

The researchers tested 16 different combinations of AI models and coding tools. Here is the bottom line: None of them passed the majority of the quizzes.

  • The Best Performer: The top team (a model called GPT-5.5 paired with a specific coding tool) got the right answer only 45% of the time. That means they failed more than half the time.
  • The Rest: The other teams did even worse, with pass rates ranging from about 12% to 39%.

It's like taking a driving test where the best driver only passes 45 out of 100 times. They aren't ready to drive alone yet.

Why Did They Fail?

The paper found that the robots aren't failing because they can't read the files or run the software. In fact, they often get the first half of the job right.

  • The "Smart but Wrong" Problem: The AI can find the right files and do the math, but then it gets confused about what the math means.
  • The Analogy: Imagine a robot chef who can perfectly chop vegetables and boil water (the technical skills). But when it comes to tasting the soup and deciding if it needs salt, the robot guesses based on what it thinks soup should taste like, rather than tasting the actual pot in front of it.
  • Specific Mistakes: The robots often:
    • Used the wrong unit of measurement (like measuring in inches instead of centimeters).
    • Applied a rule from a textbook that didn't fit the specific data in front of them.
    • Ignored the actual evidence in the files because they relied on a "familiar" story they knew from their training data.

The Verdict

The paper concludes that while AI agents are getting better at doing the work (running the code, finding the files), they are still very unreliable at judging the work. They struggle to make the specific, nuanced scientific decisions required to interpret epigenomic data correctly.

In short: The robots have the hands to do the job, but they don't yet have the "scientific gut feeling" to know if the result makes sense. They need more training before they can be trusted to analyze this type of data on their own.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →