Evidence Constrained Agentic Retrieval Augmented Generation for Substation Civil Engineering Preliminary Design Documents in a Single Project Study
This paper proposes EC-ARAG, an evidence-constrained agentic retrieval-augmented generation framework that significantly enhances the reliability, consistency, and efficiency of substation civil engineering preliminary design drafting by integrating regulation-aware retrieval, conflict resolution, and dual verification loops to outperform conventional RAG systems.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of large-scale infrastructure, the preliminary design document is the blueprint that turns a concept into a buildable reality. For a power substation, this text-heavy report must weave together project specifics, strict government regulations, standard engineering practices, and lessons learned from past projects into a single, coherent narrative. Every number, from the width of a road to the depth of a foundation, must be consistent throughout the entire document, and every claim must be backed by a specific, current rule. For decades, engineers have assembled these documents by manually cross-referencing hundreds of pages of standards and historical reports, a process that is slow, prone to human error, and difficult to keep perfectly consistent.
The rise of artificial intelligence offered a potential shortcut. Large language models, the same technology behind advanced chatbots, can write fluent technical prose and follow complex instructions. However, when applied to engineering, these models face a critical flaw: they rely on their internal memory, which can be outdated, hallucinate facts, or fail to distinguish between a rule that applies to a specific building type and one that does not. Simply asking an AI to "write a design" often results in text that sounds professional but is legally or physically invalid. To solve this, researchers have developed systems that force the AI to look up information before it writes, a method known as retrieval-augmented generation. Yet, even these systems struggle when the information they find is contradictory, outdated, or irrelevant to the specific conditions of the project at hand.
A team of engineers from State Grid Shanghai Electric Power Design Co., Ltd. has now introduced a more sophisticated approach to this problem, specifically tailored for the high-stakes environment of substation civil engineering. They developed a system called EC-ARAG, which acts less like a simple writer and more like a rigorous, multi-step editor. Instead of just retrieving text and passing it to the AI, this system breaks the design task down into tiny, verifiable claims. For every single statement the system generates—such as the required thickness of a concrete wall or the width of a drainage channel—it first checks if the supporting evidence is valid for that specific project, if the source of the rule is authoritative, and if the rule conflicts with any other part of the document.
The researchers tested their system on a real-world scenario involving the preliminary design for a specific substation project. They fed the system a massive library of knowledge, including thirty-six national and industry standards, twelve company-specific rules, eighteen typical design guides, and sixty-four historical design documents from past projects. The system then attempted to draft the entire civil engineering report, section by section. Unlike standard AI tools that might generate a paragraph and move on, this system operates in two loops. The first loop plans what needs to be written and hunts for the right evidence, filtering out any rules that are obsolete or do not apply to the specific site conditions. The second loop writes the text, then immediately checks its own work. If the system finds that a claim lacks support, contradicts a previous section, or uses an outdated standard, it stops and repairs that specific sentence before moving forward.
The results of this experiment were significant. The system successfully retrieved the correct, applicable regulations 94.1% of the time, a marked improvement over other methods that often retrieved irrelevant or outdated clauses. More importantly, the system ensured that 95% of the engineering claims it made were backed by solid evidence, and it reduced the rate of "unsafe claims"—statements that were unsupported or contradictory—to just 3.4%. Perhaps most striking was the system's ability to maintain consistency. In a long document where a building's name or a road's width might appear in ten different sections, the system kept these parameters consistent 97% of the time, whereas other methods often let these details drift or change unintentionally.
While the system required more computer processing time to perform these checks and repairs, it actually saved time overall. Because the draft it produced was so much cleaner and more accurate, human engineers spent 31% less time reviewing and fixing the text. The total time to complete the document dropped by 21% compared to the next best automated method. The study also highlighted where the system still needs human help. It struggled most with sections that rely heavily on physical geometry, such as road gradients and drainage directions, because the system was working only with text and did not have access to the actual maps or 3D models of the site. In cases where critical information was missing, the system learned to pause and flag the issue for a human engineer rather than guessing, a behavior that proved safer than making up a solution.
This work demonstrates that for complex, regulation-heavy tasks, the goal of artificial intelligence is not to replace the engineer's judgment but to act as a tireless, hyper-vigilant assistant. By treating the design process as a series of small, verified steps rather than a single act of generation, the researchers showed that machines can handle the heavy lifting of information gathering and consistency checking. The system does not claim to be an autonomous designer capable of approving a building; rather, it serves as a powerful tool that ensures the draft it produces is traceable, consistent, and grounded in the correct rules, leaving the human expert free to focus on the final decisions and the complex physical realities that text alone cannot capture.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.