From Regulation to Implementation: A Critical Evaluation of LLM-Assisted Regulatory Compliance in Industry
This paper critically evaluates the effectiveness of Large Language Models in generating regulatory compliance artifacts for the EU's Digital Product Passports and Data Protection Impact Assessments, revealing that while strict formatting guidelines ensure consistency, they may increase hallucinations, whereas less structured requirements demand more detailed prompts to maintain output quality.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the modern industrial world, companies face a growing mountain of rules designed to protect the planet and people's privacy. To prove they are following these rules, businesses must create specific documents, often called compliance artifacts. Think of these as detailed reports that act as a passport for a product or a safety audit for a computer system. One major European rule requires a digital passport for every battery sold, listing exactly what materials are inside and how to recycle them. Another rule demands a risk assessment for any system that handles personal data, ensuring that people's private information is safe. Creating these documents by hand is difficult because the data needed is scattered across many different computer systems and comes in messy, inconsistent formats. As companies struggle with this paperwork, they have turned to artificial intelligence, specifically large language models, to write these reports for them. These computer programs are excellent at reading text and generating new content, but it remains unclear whether they can be trusted to follow strict legal instructions without making things up or missing critical details.
A team of researchers set out to test exactly how well these artificial intelligence tools perform when asked to write these mandatory reports. They focused on two very different types of documents to see if the nature of the instructions changed the outcome. The first was a Digital Battery Passport, a document with a very clear, rigid structure defined by regulators. The second was a Data Protection Impact Assessment, a document required for privacy laws that is much more open-ended and vague in its instructions. The researchers fed five different artificial intelligence models a series of tasks. In some cases, they gave the models a simple, brief instruction to write the report. In other cases, they provided a long, detailed guide that explained exactly what information was needed and how to format it. They ran these tests multiple times to see if the models would produce the same result every time, and then they checked if the generated reports contained all the legally required information.
The results revealed a sharp divide between how the models handled the two types of documents. When asked to create the Digital Battery Passport, the artificial intelligence systems were remarkably consistent. Because the rules for this document were so specific, the models produced nearly identical outputs regardless of whether the instructions were short or long. They successfully extracted the necessary data about battery materials and formatted it correctly almost every time. However, the story changed completely when the models tackled the Data Protection Impact Assessment. Here, the lack of a rigid structure meant that the models struggled to stay consistent. When given only the basic regulation text without extra guidance, the models often produced reports that missed key sections or varied wildly from one attempt to the next. The researchers found that for these vague tasks, providing the model with more context and specific examples was essential to get a reliable result. Without that extra help, the models tended to leave out important details, such as how to handle data storage limits or how to assess risks to individuals.
One of the most striking findings was how the models reacted to the level of detail in the instructions. For the strict battery passport, adding more details to the prompt did not significantly change the quality of the output, but it may lead to more hallucinations in the output. For the privacy assessment, however, the opposite was true. The models needed that extra context to perform well. When the instructions were too vague, the models often failed to include required sections, suggesting that the current legal language is too broad for these tools to interpret reliably on their own. The study also highlighted that different artificial intelligence models behaved differently; some were consistently reliable across all tasks, while others were prone to dropping important information when the instructions were not crystal clear.
Ultimately, the research suggests that the success of using artificial intelligence for legal compliance depends heavily on the type of document being created. For highly structured tasks like the battery passport, the technology is ready to be used with minimal setup. But for complex, open-ended tasks like privacy assessments, simply asking the computer to follow the law is not enough. The researchers concluded that to make these tools work effectively for the vague parts of regulation, companies must invest time in crafting detailed, specific prompts that guide the artificial intelligence through the necessary steps. Until regulations themselves become more specific about what these documents must contain, or until the way we instruct these machines improves, the use of artificial intelligence for compliance will remain a mix of high reliability for some tasks and significant risk for others.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.