← Latest papers
🤖 AI

Epistemic Subordination: Generative AI and the Infrastructure of Knowledge

This paper argues that generative AI creates a condition of "epistemic subordination" by encoding dominant cultural frameworks as the default infrastructure of knowledge, a structural harm that existing legal frameworks fail to address because they regulate downstream outputs rather than the training processes where the subordination originates.

Original authors: Gilad Abiri, Emanuel V. Towfigh

Published 2026-08-20
📖 7 min read🧠 Deep dive

Original authors: Gilad Abiri, Emanuel V. Towfigh

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the modern world, we increasingly rely on artificial intelligence to summarize news, write stories, and answer questions. These systems, known as generative AI, do not simply retrieve facts from a database; they learn to predict what words should come next by studying vast amounts of text from the internet. This process creates a statistical model of human language, where the system learns patterns based on what it has read most often. For decades, experts have worried that these systems might be biased, meaning they could unfairly favor one group of people over another. The usual assumption has been that this bias comes from specific errors in the data or from the choices made by the programmers who build the tools. However, a new perspective suggests the problem is far deeper and more structural. It is not just that the AI makes mistakes; it is that the very foundation of how the AI understands the world is built on a single, dominant way of thinking, making all other ways of knowing seem like deviations.

Two legal scholars, Gilad Abiri and Emanuel Towfigh, argue that this phenomenon, which they call epistemic subordination, represents a fundamental shift in how knowledge is organized. They propose that generative AI does something no previous technology has done: it encodes the majority's way of knowing as the default infrastructure for all knowledge. In this system, the cultural assumptions, languages, and reasoning styles of the global majority become the invisible standard against which everything else is measured. The harm is not merely that the AI produces unfair results, but that it constructs a statistical reality where minority cultures and languages are structurally pushed to the bottom, not because they are excluded, but because they are absorbed and flattened into a single, homogenized average. The authors show that this issue unifies three separate problems often treated as distinct: discriminatory outputs, threats to minority cultures, and the narrowing of democratic debate. They find that these are not random glitches but the inevitable result of how these models are built.

The mechanism behind this subordination unfolds in three distinct stages, starting with the data itself. Most AI models are trained on the unstructured totality of human expression available on the internet. This digital landscape is not neutral; it is overwhelmingly dominated by English-language content, Western perspectives, and the views of populations with the education and access to publish online. When a model learns from this corpus, it does not just learn words; it learns the cultural assumptions and normative frameworks of that majority as its statistical default. The second stage makes this default irreversible. During training, the data is compressed into billions of statistical parameters that cannot be unpacked or traced back to their original sources. Unlike older computer programs where a programmer could point to a specific rule causing a bias, these models are opaque by design. The bias is not in a single line of code or a specific variable; it is woven into the entire architecture of the system.

The third stage involves the tools used to make these AI systems safe and fair, known as value alignment. Techniques like reinforcement learning from human feedback adjust what the model says to make it more polite or less harmful. However, the authors argue that these tools only change the output, not the underlying knowledge. When developers try to force a model to be diverse, the results can be absurd, such as generating images of racially diverse Nazi soldiers, because the model's baseline has no genuine framework for contextual diversity. It simply tries to crudely override a statistical default it cannot truly understand. Consequently, the harm is produced at the level of construction, not at the level of the final answer. The model does not just produce biased decisions; it subordinates entire forms of knowledge and culture by treating them as deviations from a standard they had no part in setting.

This structural reality creates a crisis for existing laws designed to protect minorities. Anti-discrimination laws, for instance, work by comparing a specific decision against a neutral baseline. If a bank rejects a loan applicant based on a specific rule that hurts a protected group, the law can strike it down. But generative AI offers no such discrete target. The discriminatory output is the product of the entire architecture, not a single identifiable rule. The "decision" is a probabilistic guess, the "criteria" are billions of invisible parameters, and the "entity" responsible is a chain of developers and deployers. The legal tools of intent and causation lose their grip because the harm is not a specific act of exclusion but a condition of inclusion where minority viewpoints are absorbed and flattened.

The threat extends to cultural and linguistic rights, which are designed to protect minority institutions from the dominance of the majority. Traditionally, a minority school could curate its own library or choose its own curriculum to preserve its language and traditions. The cultural bias in those technologies resided in the content, which could be selected. Generative AI reverses this. The bias is in the architecture of the model itself. A minority community cannot build its own foundation model because the cost is hundreds of millions of dollars and the required data does not exist at the scale of the internet for smaller languages. Recent testing shows that even the best models achieve high accuracy in English but drop significantly in languages like Swahili or Yoruba. When a religious school or a minority-language classroom uses an AI assistant, it imports a normative framework that defaults to secular, Western reasoning. The technology penetrates the institutional spaces these rights were meant to protect, restructuring knowledge from within in a way that previous tools never did.

Finally, this phenomenon threatens the very infrastructure of democratic deliberation. A healthy democracy depends on the availability of genuinely diverse frameworks for understanding problems. It requires that people be exposed to viewpoints they would not have chosen in advance. However, if generative AI becomes a primary source of information, it homogenizes the epistemic field. Minority frameworks are not censored; they are absorbed into the model's training data but structurally subordinated in the output. Public discourse will not disappear, but it will increasingly take place within a field that has already been flattened to the median culture. The legal frameworks designed to protect viewpoint pluralism, such as rules for media diversity, assume that diverse knowledge exists independently and just needs a channel to reach the public. Epistemic subordination collapses this assumption, narrowing the field from which diverse speech arises.

The authors conclude that current legal responses are structurally unable to reach this problem because they operate downstream, regulating what AI systems do rather than how they are built. Laws like the European Union's AI Act or various state regulations focus on the application of AI in hiring or lending, or on the safety of the final product. They do not address the composition of the training data or the architecture of the model itself. The central claim of the essay is that protecting epistemic pluralism requires a fundamentally different regulatory orientation. Law must be pushed to govern at the level of training, shaping the composition of data, the methods of alignment, and the model architecture.

This is not merely a theoretical demand; recent research suggests the technical capacity to intervene is real. Scientists have proposed modular architectures where a base model draws on smaller, community-specific models to preserve distinct perspectives. Others have demonstrated how to systematically construct culturally grounded training data for underrepresented languages, ensuring nuance is preserved rather than flattened. There are also methods to generate synthetic training data to represent specific cultural perspectives and to redesign the alignment process to represent a full distribution of human values rather than a single default. While these proposals are not finished solutions, they demonstrate that the technical path forward exists. What remains absent is the political will and the legal framework to require and enable these interventions. The task ahead is to build laws that can reach the construction of the model itself, ensuring that the infrastructure of knowledge does not subordinate the very diversity it claims to serve.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →