An Audit of Machine Learning Experiments on Software Defect Prediction
This paper audits 101 software defect prediction studies published between 2019 and 2023, revealing widespread inconsistencies in experimental design and reporting that severely limit reproducibility and highlight a critical need for improved research practices.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a massive, bustling kitchen where thousands of chefs (researchers) are trying to invent a new recipe to predict which ingredients in a giant warehouse are about to go bad (software defects). They all claim to have the best "Machine Learning" recipe. But how do we know if their recipes actually work, or if they just look good on paper?
This paper is like a health inspector who walks into that kitchen to audit the last five years of cooking. Instead of tasting the food, the inspectors checked the recipes (the experiments) to see if the chefs followed basic rules of science, reported their ingredients honestly, and left enough notes for someone else to cook the same dish later.
Here is what the audit found, explained simply:
1. The "Recipe" Chaos
The inspectors found that every chef was doing things differently.
- The Ingredients: Some chefs used just one dataset (a pile of ingredients), while others used 365 different ones. It was a wild mix.
- The Tools: They used anywhere from 1 to 34 different "learners" (cooking methods).
- The Scorecards: They measured success using different scorecards. Some used a metric called "F1" or "Accuracy," which the inspectors warned are like judging a race by how many people finished, ignoring whether they actually ran the right distance. These scores can be misleading, especially when the "bad" ingredients are rare.
2. The "Blind Taste Test" Problem
In a fair experiment, you must test your recipe on new ingredients you haven't cooked with before. This is called "Out of Sample" validation.
- The Issue: About 35% of the chefs didn't do this. They tested their recipe on the exact same ingredients they used to learn it. It's like a chef tasting their soup while they are still adding the salt. Of course, it tastes good! But it doesn't mean it will taste good to a customer.
- The Result: Many studies claimed their recipes were great, but they might just be memorizing the ingredients rather than learning to cook.
3. The "Missing Notes" (Reproducibility)
If you find a great recipe in a magazine, you expect to be able to cook it yourself. The inspectors checked if the papers left enough notes (code, data links, specific settings) for others to reproduce the results.
- The Score: The average "reproducibility score" was only about 52%. This means if you tried to cook these dishes, you'd be missing half the instructions.
- The Paywall: Almost half the papers were behind a "paywall" (like a locked door). You had to pay to see the recipe. The inspectors noted that if you can't see the recipe, you can't check if it's real.
- Broken Links: Even when links to the data were provided, about 35% of them were broken (like a link to a grocery store that no longer exists).
4. The "Paper Mill" Suspicion
The inspectors found two very strange papers. The authors used weird, made-up phrases like "novel insects" instead of "new bugs" and "arranging of deformities" instead of "analyzing defects."
- The Metaphor: It's like a chef writing a recipe that says "add a pinch of fluffernutter" instead of "salt." This is a red flag for "paper mills"—factories that churn out fake or low-quality scientific papers just to get published, often by translating text back and forth between languages until it sounds nonsense.
5. The "Magic Number" Trap
Many chefs ran their experiments dozens of times and looked for any result that looked "statistically significant" (like finding a winning lottery ticket by buying enough tickets).
- The Issue: They didn't adjust their rules for doing so many tests. This is like flipping a coin 100 times and claiming you found a "magic pattern" just because you got heads 10 times in a row once. The inspectors found that many studies made this mistake, making their "discoveries" likely just luck.
6. Journal vs. Conference (The Magazine vs. The Flyer)
The inspectors noticed a difference between papers published in long, detailed Journals and shorter Conference papers.
- The Finding: Journal papers were slightly better. They had more complete notes, fewer missing links, and fewer mistakes. It's like a cookbook with a full recipe vs. a flyer with just a list of ingredients. The longer space and stricter review process of journals seemed to help.
The Bottom Line
The audit concluded that while there is a lot of effort being put into this field (over 1,500 studies in five years), the quality is shaky.
- The Good News: Most of the problems aren't hard to fix. They just need chefs to write down their steps clearly, use better scorecards, and test on new ingredients.
- The Bad News: Currently, if you try to repeat these experiments, you will likely fail because the instructions are missing or the math is flawed.
In short: The field is full of potential, but right now, it's like a kitchen where half the chefs forgot to write down their recipes, and the ones who did, wrote them in a language that doesn't make sense. The inspectors are asking everyone to clean up the kitchen so the rest of us can actually eat the food.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.