← Latest papers
💬 NLP

Medmarks: A Comprehensive Open-Source LLM Benchmark Suite for Medical Tasks

This paper introduces Medmarks, a comprehensive open-source benchmark suite comprising 30 diverse medical tasks and 61 evaluated models, which reveals that frontier reasoning models outperform others in accuracy and token efficiency while highlighting the susceptibility of smaller models to answer-order bias.

Original authors: Benjamin Warner, Ratna Sagari Grandhi, Max Kieffer, Aymane Ouraq, Saurav Panigrahi, Geetu Ambwani, Kunal Bagga, Nikhil Khandekar, Arya Hariharan, Nishant Mishra, Manish Ram, Shamus Sim Zi Yang, Ahmed
Published 2026-05-05
📖 5 min read🧠 Deep dive

Original authors: Benjamin Warner, Ratna Sagari Grandhi, Max Kieffer, Aymane Ouraq, Saurav Panigrahi, Geetu Ambwani, Kunal Bagga, Nikhil Khandekar, Arya Hariharan, Nishant Mishra, Manish Ram, Shamus Sim Zi Yang, Ahmed Essouaied, Adepoju Jeremiah Moyondafoluwa, Robert Scholz, Bofeng Huang, Molly Beavers, Srishti Gureja, Anish Mahishi, Sameed Khan, Maxime Griot, Hunar Batra, Jean-Benoit Delbrouck, Siddhant Bharadwaj, Ronald Clark, Ashish Vashist, Anas Zafar, Leema Krishna Murali, Harsh Deshpande, Ameen Patel, William Brown, Johannes Hagemann, Connor Lane, Paul Steven Scotti, Tanishq Mathew Abraham

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world of Artificial Intelligence as a massive, bustling hospital. For a long time, we've been trying to figure out which AI "doctors" are actually good at their jobs. But the tests we've been using to grade them have become a bit like a broken scoreboard: they're either too easy (so every AI gets an A+), too secret (so no one can check the work), or they only test if the AI can memorize a textbook rather than actually think through a patient's problem.

Enter MEDMARKS, a new, fully open-source "medical school" for AI, created by a team of researchers. Think of it as a giant, transparent gym where 61 different AI models (from tiny, budget-friendly ones to massive, super-computer versions) come to take 30 different physical and mental tests.

Here is what the paper found, explained simply:

1. The Test Suite: A Variety of Challenges

Instead of just asking "What is the capital of France?" (which is easy for AI), MEDMARKS throws a mix of challenges at the models:

  • The Multiple Choice Quiz (MEDMARKS-V): Like a standard board exam where the AI picks the right answer from a list. This includes tricky math problems and long, complicated patient stories.
  • The Open-Ended Interview (MEDMARKS-OE): Here, the AI has to chat with a "patient" or write a treatment plan. Since there's no single right answer, the researchers used a "Judge AI" (a smart referee) to grade the conversation, much like a human teacher grading an essay.
  • The "Training Wheels" (MEDMARKS-T): A special subset of these tests is designed to be used as a gym for AI. Researchers can use these to actually train the AI to get better, similar to how a coach uses drills to improve an athlete.

2. The Results: Who Won the Medals?

The researchers ran the tests 71 different times with different settings to see who came out on top.

  • The Heavyweights Win (Mostly): The biggest, most powerful AI models from companies like OpenAI (GPT-5.1/5.2) and Google (Gemini 3) generally got the highest scores. They are like the Olympic athletes of AI.
  • The "Specialist" vs. The "Generalist": Some AI models were trained specifically on medical books and notes. These "specialist" models often beat much larger "generalist" models that know a little bit about everything. It's like a local doctor who knows their neighborhood better than a famous celebrity doctor who visits once a year.
  • The Efficiency Gap: Here's a surprising twist. The top-tier, closed-source AI models (the ones you pay for) were like Ferraris: they got the best results but used very little fuel (computer power). The open-source models (free to download) were like trucks: they could sometimes get the job done, but they burned through a massive amount of fuel to do it. Some open models needed 5 times more computer power to get a similar score.

3. The Quirks and Flaws

The paper also found some funny and concerning habits in these AI "doctors":

  • The "Overthinkers": When the AI got a question wrong, it often talked more than when it got it right. It was like a student rambling nervously when they didn't know the answer, hoping to stumble onto the right one.
  • The "Order Bias": If you shuffled the order of the multiple-choice answers (A, B, C, D), some smaller AI models got confused and changed their answers, even if the content was the same. It's like a student who picks "C" just because it's in the middle, not because they know the answer.
  • The "Tool" Trouble: When the researchers gave the AI a calculator tool to help with math, many models actually got worse. They either ignored the calculator, used it wrong, or got confused by the instructions. It's like giving a chef a new knife, and they accidentally cut the wrong ingredient because they weren't used to the new tool.

4. The Big Picture

The main takeaway is that while AI is getting very good at medical tasks, it's not perfect yet.

  • Frontier models (the newest, biggest ones) are the current leaders.
  • Specialized training helps, but it doesn't guarantee a win over the biggest models.
  • Efficiency is a huge problem; the free models are often too "expensive" to run in terms of computer time.
  • Reliability is still an issue; models can be easily tricked by how questions are formatted or can get confused when using tools.

The authors built MEDMARKS to be a "living leaderboard." Just like a sports league updates its rankings every week, this benchmark is designed to constantly test new AI models as they are released, ensuring we always know who the current "best" medical AI is, and exactly where they are still struggling. They made all their code and data public so anyone can run the tests and verify the results, bringing transparency to the field.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →