← Latest papers
💻 computer science

Separating AI-assisted authoring from governed execution through specification-driven composition A design framework and industrial experience report for explainable automation in regulated data transformation

This paper proposes a specification-driven composition framework that enables the safe use of AI in regulated data workflows by separating AI-assisted authoring from governed execution, ensuring that only validated, versioned, and fully traceable artifacts are composed and executed.

Original authors: Rostislav Markov

Published 2026-09-25
📖 6 min read🧠 Deep dive

Original authors: Rostislav Markov

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the world of regulated industries, such as medicine and finance, data does not simply flow from one place to another; it must be transformed with absolute precision. When a pharmaceutical company prepares data for a government review, every number, date, and category must follow strict rules. If a computer program makes a mistake in this process, the consequences can be severe, ranging from rejected applications to safety risks. For decades, engineers have built these programs by writing code line by line, ensuring every step is checked, tested, and approved before it ever runs. This process is slow but reliable because humans remain in control of every decision.

Today, artificial intelligence offers a new way to write these programs. These systems can read a description of a task and generate the necessary code in seconds. This speed is incredibly useful for exploring ideas or drafting initial versions of a program. However, in high-stakes environments, speed cannot come at the cost of safety. A computer cannot be allowed to simply guess the right answer and run it immediately. The core challenge for modern engineers is how to use the speed of artificial intelligence without losing the ability to explain, verify, and control the final result. They need a way to let the machine draft the work while keeping a human in the loop to verify that the work is safe, correct, and follows the rules before it is ever allowed to execute.

This is the problem Rostislav Markov, a researcher at Amazon Web Services, addresses in his recent work on regulated data transformation. He proposes a new way of working that separates the act of writing or drafting code from the act of running it. Instead of letting an artificial intelligence model generate a program and then hoping it works, the new framework forces the AI to produce a detailed plan first. This plan, called a specification, describes exactly what needs to be done, what data it will use, and what rules it must follow. The plan itself is not the program; it is a set of instructions for a separate, strictly controlled system called a composer.

The composer acts as a gatekeeper. It looks at the plan created by the AI and checks it against a library of pre-approved, tested, and verified building blocks. These blocks are specific functions that engineers have already validated to work correctly. The composer does not invent new logic or guess how to solve a problem. It simply matches the steps in the plan to the approved blocks in the library. If the plan asks for a step that has no matching approved block, or if the data types do not match the rules, the composer stops and rejects the plan. It will not run the code until every single piece has been verified against the library and the rules. This ensures that even if the artificial intelligence makes a mistake in its draft, that mistake is caught before any real data is touched.

Markov tested this idea using a real-world scenario involving clinical data, where patient information must be organized into a standard format for government submission. In his experiments, he asked an artificial intelligence to write the code for these transformations under different conditions. When the AI was given only a simple text description, it often produced code that looked correct but failed to fit into the larger system or violated hidden rules. When the AI was given a structured plan with clear rules and references to the approved building blocks, the results were much better. The key finding was not that the AI became smarter, but that the system around it became more robust. The artifacts—the plans, the tests, and the approved blocks—did the heavy lifting of ensuring safety, not the AI model itself.

The framework introduces a concept called "explainable by construction." In many discussions about artificial intelligence, people worry about understanding why a model made a specific decision, often trying to look inside the "black box" of the algorithm. This approach takes a different path. It argues that for regulated work, you do not need to understand the inner thoughts of the AI. Instead, you need to understand the engineering decisions that led to the final result. Because the system only allows approved blocks to be used and keeps a complete record of every step taken to build the final program, anyone can look at the final result and trace it back to the specific plan, the specific approved blocks, and the specific rules that were followed. The explanation is not a guess about what the AI was thinking; it is a factual record of what was approved and how it was assembled.

To prove that this system works, the researcher built a small, working version of the framework. He showed that if you give the system the same plan and the same library of blocks twice, it produces the exact same result every time. This consistency is vital for regulated environments where reproducibility is required. He also demonstrated that if a new version of a building block is added to the library, the system does not accidentally use it for an old plan; it sticks to the specific version that was approved at the time the plan was made. This prevents changes in the background from unexpectedly altering the outcome of a process. The system generates a detailed log, or trace, for every program it builds. This log links the final program back to the original plan, the specific version of every building block used, and the validation checks that were passed.

The research suggests that this separation of duties is essential for the future of artificial intelligence in regulated fields. It allows organizations to use the speed of AI to draft specifications and write code, but it keeps the authority to execute in the hands of a controlled, deterministic system. The AI becomes a powerful assistant that helps humans write better plans, but it is not the authority that decides what runs. The final decision to run a program rests on the existence of a valid plan, the presence of approved components, and the successful completion of a verification process. This approach does not eliminate the need for human oversight; rather, it gives humans better tools to manage that oversight. By focusing on the artifacts—the plans, the rules, and the approved blocks—the framework creates a boundary where artificial intelligence can be useful without becoming a risk.

In the end, the work demonstrates that the path to safe automation is not about making the artificial intelligence more explainable, but about making the process of building the automation more transparent. When the steps are clear, the rules are strict, and the components are verified, the result is a system that can be trusted. The researcher's experience shows that while artificial intelligence can generate plausible code, it is the surrounding structure of specifications, tests, and approved libraries that ensures the code is actually correct and safe for use. This shift in focus, from the intelligence of the model to the governance of the process, offers a practical way forward for industries that cannot afford to be wrong.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →