Are Large Language Models Ready for Quantum Software Engineering? A Multivocal Literature Review
This Multivocal Literature Review synthesizes evidence from 24 sources to conclude that while Large Language Models show promise for specific, code-centric Quantum Software Engineering tasks like synthesis and repair, they currently function primarily as bounded assistive tools rather than robust, lifecycle-spanning agents due to significant limitations in semantic accuracy, backend execution, and domain coverage.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Picture: A New Apprentice in a High-Stakes Lab
Imagine Quantum Computing as a brand-new, incredibly complex laboratory. The scientists here are trying to build machines that work on the laws of physics we barely understand (like atoms being in two places at once). Building software for these machines is called Quantum Software Engineering (QSE). It's notoriously difficult because the tools are new, the instructions are confusing, and if you make a tiny mistake, the whole experiment fails.
Enter Large Language Models (LLMs). Think of these as super-smart, fast-talking apprentices who have read almost every book in the library. In regular software development (like building a website or an app), these apprentices are already very popular. They can write code, fix bugs, and explain things to humans.
The Big Question: The authors of this paper asked, "Are these super-smart apprentices ready to work in our high-stakes Quantum Laboratory?"
To find out, they didn't just look at academic textbooks. They also looked at "grey literature"—which includes tech blogs, company reports, and forum discussions—because in fast-moving fields, the real-world experience often happens outside of formal papers. They reviewed 24 different sources (a mix of academic studies and industry reports) to get a complete picture.
What They Found: The Apprentice is Good at One Thing, But Struggles with the Rest
The researchers mapped the findings onto the "lifecycle" of building quantum software. Think of this lifecycle like building a house: you need to plan the design, lay the foundation, build the walls, install the plumbing, and finally inspect the work.
Here is the state of the apprentice (the LLM) in each stage:
1. The "Bricklaying" Phase (Implementation) 🧱
- Status: Very Active.
- The Analogy: This is where the apprentice is most useful. They are great at "bricklaying"—writing the actual code or circuits. If you ask them, "Write me a quantum circuit to do X," they can generate a draft quickly.
- The Catch: Just because they can lay bricks doesn't mean the wall is straight. The paper found that while they generate code well, that code often has "hallucinations" (making things up) or doesn't actually work when run on real quantum hardware.
2. The "Inspector" Phase (Analysis & Repair) 🔍
- Status: Getting Started.
- The Analogy: Here, the apprentice tries to find cracks in the wall or fix broken pipes. They are being used to refactor (reorganize) old code or explain what a complex circuit does.
- The Catch: Their explanations can be shallow, and their fixes might introduce new errors. They are helpful, but you can't trust them to do the inspection alone.
3. The "Blueprint" and "Safety Check" Phases (Requirements, Architecture, Testing) 🏗️🛡️
- Status: Almost Empty.
- The Analogy: This is where the apprentice is barely showing up. Very few studies looked at using them to design the overall architecture of the system, figure out what the customer actually needs (requirements), or run rigorous safety tests.
- The Reality: The field is so focused on just "writing code" that the bigger picture of how to build a reliable, safe quantum system is being ignored.
The Tools They Are Using: The "Brand Name" Problem
The paper noticed a heavy reliance on proprietary models (like GPT-4 from OpenAI).
- The Analogy: It's like every construction crew in town is using the exact same brand of power drill because it's the most famous one.
- The Problem: Because these tools are owned by private companies, other scientists can't always see how they work or run the exact same experiment later. This makes it hard to verify if the results are real or just a fluke. While some open-source models (like LLaMA) are being tried, they are used much less frequently.
The Main Warnings: Why We Can't Trust Them Yet
The authors identified several "red flags" that suggest these tools aren't ready to work alone in the quantum lab:
- The "Fake Fact" Problem (Correctness): The apprentice often sounds confident but is wrong. They might write code that looks perfect but fails immediately when you try to run it on a real quantum computer.
- The "Sensitive Ears" Problem (Prompt Dependency): The apprentice is very sensitive to how you ask questions. If you phrase your request slightly differently, the output changes completely. This makes it hard to get consistent results.
- The "Small Library" Problem (Data Coverage): The apprentice was trained mostly on classical software. They haven't read enough "quantum books" yet. When they encounter a complex, unique quantum problem, they don't have enough data to give a good answer.
- The "Toy Test" Problem (Evaluation): Many studies only tested the apprentice on simple, toy problems. We don't know if they can handle the messy, complex reality of real-world quantum engineering.
The Final Verdict
Are Large Language Models ready for Quantum Software Engineering?
Not quite.
The paper concludes that LLMs are currently assistive tools, not autonomous engineers.
- Think of them like a spell-checker: They are great at catching typos or suggesting a better word, but you wouldn't let them write your entire novel without you reading it first.
- In Quantum terms: You can use them to generate a draft of a quantum circuit, but a human expert must verify, test, and fix it before it ever touches a real quantum machine.
The authors suggest that for these tools to become truly reliable, researchers need to focus less on just "generating code" and more on building verification systems (safety checks) and creating open, reproducible models that everyone can trust and test. Until then, the quantum apprentice is a helpful intern, but not yet a master builder.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.