← Latest papers
💻 computer science

Rethinking Artifact Evaluation for Software Engineering in the Age of Generative AI

This position paper argues that in the era of generative AI, where polished narratives are easily produced, software engineering peer review must shift its focus toward treating artifact evaluation as a first-class component to better assess the scientific substance of research.

Original authors: Christoph Treude, Christopher M. Poskitt, Rashina Hoda

Published 2026-04-21
📖 4 min read☕ Coffee break read

Original authors: Christoph Treude, Christopher M. Poskitt, Rashina Hoda

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Problem: The "Polished Wrapper" Trap

Imagine you are a food critic at a busy restaurant. You have 50 plates to review, but you only have 30 minutes.

In the past, if a dish looked messy or the description was written in bad handwriting, you might assume the chef didn't care about the food. But now, imagine a robot (Generative AI) has appeared in the kitchen. This robot can instantly write a beautiful, poetic menu description and make the plate look perfectly arranged with fancy garnishes.

Suddenly, every dish looks amazing on the outside. The "wrapper" is perfect.

The problem? The critic still has to actually taste the food to know if it's good. Tasting takes time, requires a trained palate, and is hard work. But because the menu descriptions are now so easy to fake, the critic spends all their time reading the pretty words instead of tasting the food. They might give a 5-star rating to a dish that tastes like cardboard, just because the menu was written by a genius robot.

This is exactly what is happening in Software Engineering research.

The Current Situation: Too Much Fluff, Not Enough Substance

In the world of academic research, scientists submit papers to be reviewed by other experts (peers).

  • The "Narrative": This is the story, the writing, the introduction, and the motivation. Generative AI has made it incredibly easy to write a perfect story.
  • The "Artifact": This is the actual "food." In software engineering, this means the code, the data, the experiments, and the tools the researchers built. This is the hard part. You can't just "AI" a working software system or a valid scientific experiment; it requires real human effort and expertise to build and check.

The authors of this paper argue:
Right now, reviewers are spending too much time checking if the story sounds good (which is easy to fake) and not enough time checking if the code actually works or if the data proves the point (which is hard to fake).

The Solution: Make the "Code" the Star of the Show

The paper suggests a major shift in how we judge research. They want to treat Artifacts (the code and data) as "First-Class Citizens."

Here is what that means in plain English:

  1. Stop judging the book by its cover: Don't let a beautifully written paper with a perfect AI-generated introduction fool you.
  2. Taste the food first: If a researcher claims their new software is faster or safer, the proof isn't in the paragraph where they say it's faster. The proof is in the code they wrote.
  3. Change the scoring system: Right now, if you write a great story but your code is broken, you might still get published. The authors say: If the code (the artifact) doesn't hold up, the paper shouldn't get published, no matter how good the story is.

Why This Matters Now

The paper points out a dangerous imbalance:

  • Writing a story is getting cheaper and faster (thanks to AI).
  • Building and checking the science is still expensive, slow, and requires human brains.

If we keep reviewing papers the old way, we will end up with a library full of beautifully written books that contain no useful information. The "signal" of a good paper (how well it's written) is getting weaker, while the "signal" of a good paper (does the math/code work?) is getting harder to ignore.

The Analogy of the "Blueprint"

Think of a research paper like a blueprint for a new type of bridge.

  • The Narrative is the sales pitch: "This bridge will save lives, it's beautiful, and it's the future!" (AI can write this pitch in seconds).
  • The Artifact is the actual engineering calculations and the stress-test data.

If a reviewer only reads the sales pitch, they might approve a bridge that will collapse. The authors are saying: "Stop reading the sales pitch. Hand me the engineering calculations. If the math doesn't work, the bridge doesn't get built, regardless of how pretty the pitch was."

The Bottom Line

The authors aren't saying "stop writing well." They are saying: "We need to stop letting the writing hide the lack of real work."

By making the evaluation of code and data (the artifacts) the most important part of the review process, we ensure that software engineering research remains trustworthy, even as AI makes it easier to write fancy stories. It's about shifting our attention from how the research sounds to what the research actually does.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →