← Latest papers
📊 statistics

Benchmarking AI Performance on End-to-End Data Science Projects

This paper introduces a benchmark of 40 end-to-end data science projects and an automated grading pipeline to evaluate generative AI models, finding that while recent models perform well on structured tasks, they exhibit significant variability and require human verification for judgment-heavy components.

Original authors: Evelyn Hughes, Rohan Alexander

Published 2026-02-17
📖 4 min read☕ Coffee break read

Original authors: Evelyn Hughes, Rohan Alexander

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are running a cooking school. For years, you've tested your students by asking them to chop onions, boil water, or bake a single cookie. These are great tests for specific skills, but they don't tell you if a student can actually run a whole restaurant, manage a dinner rush, and serve a complete, delicious meal to a hungry customer.

This paper is about testing AI chefs to see if they can cook a full, five-course meal from scratch, rather than just chopping a single vegetable.

Here is the story of what the researchers found, explained simply:

1. The Challenge: The "Full Meal" Test

The researchers (Evelyn Hughes and Rohan Alexander) wanted to know: Can AI do a complete data science project?

A real data science project isn't just writing code. It's a long journey that looks like this:

  • Asking the right question: "What story is hidden in this data?"
  • Gathering ingredients: Downloading messy data from the internet.
  • Cleaning the kitchen: Fixing errors and organizing the data.
  • Cooking: Analyzing the numbers and making charts.
  • Serving the dish: Writing a report that explains what happened, why it matters, and citing where the ingredients came from.

They created a "test kitchen" with 40 real projects that university students had already done. They graded these projects with a strict checklist (a rubric) that looked for everything from code quality to how well the story was told.

2. The Contestants: The AI Chefs

They asked seven different AI models (the "chefs") to recreate these projects. Some were the latest, most expensive models (like Claude Opus 4.6 and GPT-5.2), and some were older models (like GPT-4o).

3. The Results: Who Passed?

The grading was done automatically by another AI (acting as the head chef), and the results were surprising:

  • The Top Chefs (The "A-Students"): The newest models, like Claude Opus 4.6, scored an 85%. They were almost as good as a smart, hard-working university senior. They could follow instructions perfectly, write clean code, and organize their files beautifully.
  • The Middle Chefs: Models like GPT-5.2 and Gemini 3 Flash scored in the high 70s. They were good, but had some rough edges.
  • The Struggling Chefs: Older models like GPT-4o scored a 32%. They basically failed the test. They couldn't finish the meal; they forgot to clean the kitchen or write the menu.

4. The Catch: The "Follow the Recipe" vs. "Taste the Food" Problem

This is the most important part of the paper. The AI models were amazing at following recipes, but terrible at tasting the food.

  • The Good News (Recipe Following): If you told the AI, "Write a title," "Make a list of files," or "Draw a graph," they did it perfectly. They are great at structure. They can build the skeleton of a house very quickly.
  • The Bad News (Taste & Judgment): When the task required thinking, the AI stumbled.
    • The "Abstract" Problem: Writing a short summary of a long story is hard. The AI often wrote summaries that sounded fancy but didn't actually say anything important.
    • The "Measurement" Problem: Explaining how a real-world thing (like "homelessness") becomes a number in a spreadsheet requires deep understanding. The AI often just guessed or gave vague answers.
    • The "Citation" Problem: The AI sometimes made up fake book titles or got the names of software packages wrong. It's like a chef saying, "I used a secret spice," but not knowing what the spice was called.

5. The Verdict: AI is a Great Intern, Not a Boss

The paper concludes that AI is ready to be a very helpful intern.

  • It can do the boring, repetitive work (cleaning data, writing code templates, making charts) faster than a human.
  • However, you cannot just let it run the show alone.

If you let an AI run a data project without a human looking over its shoulder, it might serve you a meal that looks beautiful on the plate but tastes like cardboard. It might follow the rules perfectly but miss the point of the story.

The Bottom Line:
AI can now do the work of a good college student on routine tasks, but it still needs a human "Head Chef" to taste the food, check the ingredients, and make sure the story makes sense before serving it to the world. We can use AI to speed up the work, but we cannot yet trust it to do the thinking for us.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →