The Invisible Editorial Layer: Formalizing Undisclosed Inference-Time Steering, Probability Placement, and the Attribution Problem in Deployed Language Models
This paper argues that undisclosed inference-time interventions create a hidden "editorial layer" in deployed language models that systematically biases outputs and obscures attribution, necessitating new governance frameworks like "Inference Policy Transparency" to address the gap between model weights and observed behavior.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
When we ask a large language model a question, we often imagine a direct conversation between a human and a machine. We assume the answer comes straight from the computer's brain, shaped only by the data it was trained on and the words we typed. This view treats the model as a static library, where the books inside are fixed, and the only variable is which book the librarian pulls off the shelf. However, the reality of how these systems work in the real world is more complex. The technology behind these models has matured to the point where the computer can do more than just retrieve information; it can subtly rewrite the odds of what it says next, even before it finishes a single sentence. This happens in a hidden layer of software that sits between the model's core knowledge and the final text you see on your screen. It is a space where the raw output of the machine can be nudged, adjusted, or steered without changing the machine's underlying memory.
A new paper by Augusto Camargo explores the consequences of this hidden layer, arguing that we are looking at the wrong part of the system if we only care about the model itself. The author suggests that the most significant influence on what an AI says might not be in its training data or its internal code, but in the invisible rules applied the moment it is about to speak. This research shifts the focus from asking "What does the model know?" to "Who is controlling what the model is allowed to say?" The paper does not claim that these systems are currently being used to manipulate the world in a specific, proven way, but rather that the technical tools to do so exist, are already in use for other purposes, and remain largely invisible to the people using them. It is a call to recognize that the answer you receive is not just a reflection of the model, but a product of a system that can be tuned for political, commercial, or ideological reasons without anyone knowing.
The core of this argument rests on a simple but powerful observation: the model is not the same thing as the system that delivers its answers. In a standard setup, a user types a prompt, the model generates a list of possible next words with different probabilities, and then a computer picks one. The paper points out that in modern, real-world applications, there is often an extra step in the middle. Before the computer makes its final choice, a separate piece of software can look at those probabilities and adjust them. It can make certain words slightly more likely to appear and others slightly less likely, all without changing the model's internal weights or the user's prompt. This process is similar to how a radio station might slightly boost the volume of a specific song to make it stand out, even though the song itself hasn't changed. The paper formalizes this as an "inference policy," a set of rules that operates at the very last moment of generation.
The author demonstrates that this capability is not theoretical; it is already a proven technology used for other things. For instance, companies use similar methods to embed invisible watermarks into text to prove it was written by an AI, or to steer a model away from toxic language. The paper argues that the same technical machinery used to add watermarks or ensure safety could be repurposed for something far more subtle: framing. Framing is the act of presenting a topic in a way that highlights certain aspects while downplaying others, influencing how a person understands the issue without lying about the facts. A system could be programmed to consistently choose words that make a government policy sound like a protective measure, while another version of the same system could be tuned to make the exact same policy sound like a bureaucratic burden. Both answers could be factually correct, but the choice of words would guide the user toward a specific emotional or political conclusion.
This leads to a major problem the paper calls the "attribution problem." If a user asks a political question and receives a biased answer, it is nearly impossible to tell where that bias came from. Is it because the model was trained on biased data? Is it because the developers added a hidden instruction to the system? Or is it because a third party paid to have the probabilities shifted in a specific direction? Because the adjustment happens in the split second before the answer is generated, it leaves no trace in the conversation history. Unlike a hidden instruction that might be accidentally revealed, or a biased training set that might show up in a test, this kind of steering is invisible to anyone looking at the model's code or the final text. The paper suggests that this makes it incredibly difficult to hold anyone accountable for the opinions an AI expresses, because the source of the influence is hidden in the deployment layer, not the model itself.
The author also introduces a concept called "probability placement," which describes a hypothetical new form of advertising. Instead of a company paying to have its product explicitly mentioned in an answer, they could pay to have their brand slightly more likely to be chosen when the model is deciding between options. Imagine asking for a recommendation on which database to use. The model might naturally think three options are equally good, but a commercial policy could nudge the odds so that the paying company's product becomes the most likely choice. This would happen without any explicit advertisement, without the user realizing they are being influenced, and without the model ever being told to "sell" anything. It is a way of embedding commercial influence directly into the fabric of the recommendation, making it feel like a natural, unbiased opinion.
The paper examines how these dynamics interact with current laws and regulations, such as the European Union's AI Act and advertising standards. It notes that while laws exist to prevent subliminal manipulation or require disclosure of commercial ties, they were written with older technologies in mind. They assume that if a product is being promoted, it will be obvious. They do not account for a system where the promotion is a tiny, statistical shift in the probability of a word appearing. The author argues that if a system can systematically steer users toward a specific political view or commercial choice without them knowing, it challenges the very idea of informed decision-making. The harm might not be immediate or obvious in a single conversation, but the paper suggests that over millions of interactions, this subtle steering could accumulate to shape public opinion and market behavior in profound ways.
Ultimately, the research proposes a new principle for governance called "inference policy transparency." This would require companies to disclose not just how their models were trained, but how they are being steered when they are used. It calls for a system where the rules that modify the probabilities are visible and auditable, much like a label on a food package. The goal is to shift the question from "What does the model encode?" to "Who controls the probability distribution between the model and the user?" The paper concludes that as these systems become more central to how we access information, make decisions, and debate ideas, we must recognize that the model is only one part of the equation. The invisible layer that sits between the machine and the human is where the real power to shape reality may now reside, and it is a layer that currently operates in the dark.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.