Guidelines for Empirical Studies in Software Engineering involving Large Language Models
This paper presents a collaborative framework of seven study types and eight mandatory or recommended guidelines to enhance the reproducibility and rigor of empirical software engineering studies involving Large Language Models, supported by an applicability matrix, reporting checklist, and a living online resource.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the world of software engineering research as a massive, bustling kitchen where chefs (researchers) are trying to invent new recipes (tools and theories) to help cooks (software engineers) make better meals (software).
For a long time, these chefs followed strict rules to ensure that if someone else tried to make the same dish, it would taste exactly the same. But recently, a new, magical ingredient has entered the kitchen: Large Language Models (LLMs). Think of LLMs as a super-smart, but slightly unpredictable, sous-chef.
This new sous-chef is amazing. It can chop vegetables (code), write recipes (documentation), and even taste-test dishes (review code) incredibly fast. However, it has three weird quirks:
- It's a bit moody: If you ask it to make the same dish twice, it might add a pinch more salt the second time, even if you gave it the exact same instructions.
- It forgets its history: We don't always know exactly what ingredients it learned from in the past, and the "recipe book" it uses changes constantly without us knowing.
- It's a black box: Sometimes, we can't see how it decided to add that extra salt.
Because of these quirks, if a researcher says, "I made a great cake using this magical sous-chef," other chefs can't replicate the cake. They don't know exactly which version of the sous-chef was used, what the exact instructions were, or if the cake would taste different tomorrow.
This paper is a "New Kitchen Handbook" written by a team of 22 expert chefs. Their goal is to teach everyone how to use this magical sous-chef without ruining the science of cooking. They've organized the kitchen into 7 different ways the sous-chef is used (like "The Taster," "The Recipe Writer," or "The Fake Customer") and created 8 Golden Rules to follow.
Here are the 8 rules, explained with simple analogies:
1. 📢 The "Honest Chef" Rule (Declare Usage)
The Rule: You must admit you used the sous-chef.
The Analogy: If you put a secret spice in your soup, you have to tell the food critics. Don't just say "I made this soup." Say, "I used the Magic Sous-Chef v2.0 to chop the onions." If you hide it, people can't judge if the soup is actually good or just lucky.
2. 📝 The "Exact Receipt" Rule (Report Version & Settings)
The Rule: Write down the exact model name, the date you used it, and the settings (like "temperature" or "creativity level").
The Analogy: Imagine a recipe that just says "add some flour." That's useless! You need to say "2 cups of King Arthur Flour, measured on a Tuesday." Since the magical sous-chef changes its mind often, you need to record the exact moment and settings so someone else can try to copy your result later.
3. 🏗️ The "Blueprint" Rule (Report Architecture)
The Rule: Don't just talk about the sous-chef; explain the whole kitchen setup around it.
The Analogy: The sous-chef doesn't work in a vacuum. It's connected to a fridge, a timer, and a list of instructions. If you built a robot chef, you need to draw the blueprints of the robot, not just say "it uses a brain." How does the robot talk to the oven? What happens if the internet cuts out? Show the whole machine.
4. 🗣️ The "Script" Rule (Report Prompts & Logs)
The Rule: Share the exact questions you asked the sous-chef and its answers.
The Analogy: If you ask a genie, "I wish for a million dollars," and it gives you Monopoly money, the problem might be how you asked. You need to show the exact script you gave the genie. Also, keep a diary of the conversation. If the genie changes its mind next week, you need to know what it said today.
5. 👨🍳 The "Human Taste-Test" Rule (Human Validation)
The Rule: Don't trust the robot blindly; have a human taste the food.
The Analogy: The sous-chef might think a dish is perfect because it looks pretty, but a human might say, "This tastes like cardboard." You need to compare the robot's work with a human expert's work to make sure the robot isn't just hallucinating. If the robot says "This code is safe," a human needs to double-check it.
6. 🆓 The "Open Source" Rule (Use an Open Baseline)
The Rule: If you use a fancy, expensive, closed-door sous-chef (like a secret company's AI), also test a free, open-source one.
The Analogy: If you say, "My secret sauce is the best!" but you won't let anyone see the recipe, we don't believe you. Show us that your secret sauce is better than a simple, homemade sauce we can all make ourselves. This proves your results aren't just because of a magic trick hidden behind a paywall.
7. 📏 The "Fair Ruler" Rule (Use Good Metrics)
The Rule: Use the right tools to measure success, and explain why.
The Analogy: You can't measure the "tastiness" of a soup with a ruler. You need a taste-test score. Don't just say "It worked!" Explain how you measured it. Did you count how many bugs were fixed? Did you ask users if they liked it? And remember, sometimes the ruler itself is broken (bad benchmarks), so admit that too.
8. 🚧 The "Honest Mistakes" Rule (Report Limitations)
The Rule: Admit where your experiment might have gone wrong.
The Analogy: No chef is perfect. Maybe your kitchen was too hot, or maybe the ingredients were old. Be honest: "The robot worked great, but only on Tuesdays," or "We only tested this on pizza, not on sushi." By admitting the flaws, you help others avoid them.
The Big Picture
This paper is essentially a quality control manual for the future of software research. It says: "The magical sous-chef is here to stay, and it's powerful. But if we don't write down exactly how we use it, we'll never know if the results are real or just a fluke."
By following these rules, researchers can ensure that their discoveries are solid, reproducible, and actually helpful to the world, rather than just a one-time magic trick that disappears when the model updates.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.