The Say-Do Gap Architecture Applied to Agentic AI: A Reproducible NLP Pipeline for Detecting Governance Rhetoric–Reality Divergence in Corporate ESG Communication
This study introduces a reproducible, bilingual NLP pipeline that adapts the Say-Do Gap Index architecture to detect divergence between corporate rhetoric and observable governance practices regarding agentic AI within the IBEX 35 index, demonstrating the framework's generalizability beyond traditional ESG domains.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the modern corporate world, companies frequently publish reports detailing their plans to manage new technologies responsibly. They speak of ethics, safety, and oversight, promising that their use of advanced tools is under control. However, a persistent challenge in business and governance is the distance between what an organization says it will do and what it actually does. This gap is not unique to technology; it appears whenever an institution makes a public claim about its internal practices. Researchers have long studied this phenomenon, noting that words are often easier to produce than the structural changes required to back them up. When a company claims to govern a complex system but lacks the visible committees, policies, or certifications to prove it, the result is a disconnect between rhetoric and reality. This disconnect matters because it can mislead investors, regulators, and the public about the true state of corporate responsibility.
A new study by Alfredo Merlet, a researcher at Instituto Profesional IACC, tackles this problem specifically regarding "agentic artificial intelligence." This term refers to a type of AI that can act autonomously, making decisions and taking steps without constant human direction. As these systems become more common, companies are beginning to discuss how they will govern them. Merlet's work does not simply ask whether companies are talking about this technology; it asks whether their words match their actions. To answer this, the researcher built a new, repeatable system designed to measure the gap between corporate promises and the hard evidence of governance. The study focuses on the IBEX 35, a list of the thirty-five largest companies in Spain, treating them as a test case to see if the method works.
The core of this research is a pipeline, or a step-by-step process, that separates what a company says from what it has actually done. The "Say" part of the equation comes from the thousands of news articles and corporate communications these firms have published between 2023 and 2025. The researcher used a specialized computer program to scan these texts for mentions of agentic AI governance. The program was trained to recognize specific language related to oversight and control, distinguishing between a company simply talking about using AI and one discussing how it manages the risks of that AI. The "Do" part is different. Instead of relying on what the companies say, this section looks for physical proof. The researcher searched for three specific types of evidence: the existence of a board-level committee dedicated to technology risks, a formally published policy document on AI governance, and proof that the company adheres to external certification standards. These are tangible items that can be verified, unlike a speech or a press release.
When the researcher applied this system to the thirty-four companies in the sample, the results were striking, though perhaps not in the way one might expect. The study found that the system detected almost no evidence of agentic AI governance discourse in the corporate communications, and correspondingly, no evidence of the specific governance structures required to manage it. In the period from 2023 to 2025, the computer analysis of news and reports found a structurally low or null prevalence of agentic AI mentions, and the search for the three types of hard evidence came up empty for the vast majority of the firms. To ensure this was not a mistake in the computer's reading, the researcher manually checked a sample of the articles that the computer had marked as having no relevant content. This human review confirmed that the computer was correct: the companies were not yet producing significant discourse about agentic AI governance, nor were they showing the concrete steps of governance that would prove they were managing it.
This outcome is significant because it demonstrates that the method works. The study was designed primarily to test whether a framework originally built for general environmental and social issues could be adapted to the specific, technical world of autonomous AI. The fact that the system successfully identified a null result—indicating a lack of both discourse and action—proves the architecture is sound and capable of distinguishing between genuine findings and classifier errors. The researcher did not set out to accuse these specific Spanish companies of failing; rather, the goal was to show that a tool exists which can measure the difference between speech and substance, even when that difference is a lack of both. The study explicitly rules out the idea that this is a final audit of the entire industry or a definitive judgment on every firm's integrity. Instead, it presents a reproducible tool that can be used by others to check different markets or different technologies.
The value of this work lies in its transparency and its openness. The researcher has made the entire system available to the public, including the list of words the computer looks for and the code used to find the evidence. This allows other scientists, regulators, or investors to use the same method to check companies in other countries or to look at different emerging technologies. The study suggests that this approach can be adapted for markets in Latin America and beyond, provided the specific documents and regulations of those regions are accounted for. It offers a way to move beyond simply reading what companies claim and toward verifying what they have actually built.
Ultimately, the paper provides a clear, repeatable way to see if a company's governance matches its words. It shows that for agentic AI in this specific sample and time window, the current landscape is characterized by a structurally low prevalence of both rhetoric and the structural evidence of control. This does not mean the companies are lying, but it does mean that the visible proof of their management and the discourse surrounding it are currently missing. The study concludes that this method of measuring the gap is a powerful tool for understanding corporate behavior, one that can be applied to any situation where an organization makes a claim about how it manages a complex system. By separating the promise from the proof, the research offers a new lens through which to view the rapidly evolving landscape of artificial intelligence and corporate responsibility.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.