← Latest papers
💬 NLP

Words, Spaces and Generative AI: Layers of language in contemporary architecture

This paper argues that language functions as a critical design material in contemporary architecture, particularly with the rise of generative AI, by analyzing the interplay of discourse, programming languages, and annotations to propose a research agenda connecting linguistic theories to computational design practices.

Original authors: Anca-Simona Horvath

Published 2026-08-26
📖 6 min read🧠 Deep dive

Original authors: Anca-Simona Horvath

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

For centuries, architects have treated words as mere instructions, a way to tell builders what to construct. But a growing body of thought suggests that language is actually a building material in its own right, as essential to design as brick or steel. This idea rests on a simple observation: we do not just describe the world with words; we use words to shape how we see it. When we call a crowded neighborhood a "slum," we frame it as a problem to be removed; when we call it a "community," we frame it as something to be preserved. This framing happens before any solution is found, guiding our thinking in ways we often do not notice. Now, as artificial intelligence tools that turn text into images and 3D models become common in architecture, this relationship between language and space has become critical. These tools do not just follow orders; they interpret the messy, metaphorical way humans speak to generate new forms. Understanding how these machines read our words is no longer just a theoretical exercise; it is the key to understanding how the built environment of the future will be conceived.

Anca-Simona Horvath's chapter, "Words, Spaces and Generative AI," explores this new frontier by examining how three distinct layers of language intertwine when architects use these powerful new tools. The first layer is the natural language we speak and write every day, including the specific jargon architects use to discuss their work. The second layer is the programming code that runs the computers, an artificial language built on strict logic rather than human nuance. The third layer is the annotations—the short labels and descriptions attached to the millions of images used to train these AI systems. Horvath argues that when an architect types a prompt like "a sustainable library with a green roof," the AI does not simply look up a definition. Instead, it navigates a complex web where human metaphors, rigid computer code, and the biased categories of its training data collide to produce a result.

To understand how this works, the paper traces the history of how we think about language. It looks back to the early 20th century, when philosopher Ludwig Wittgenstein suggested that communication is like swapping mental pictures; we use words to trigger images in each other's minds. Later, researchers like George Lakoff and Mark Johnson showed that our minds are wired with metaphors, using concrete experiences to understand abstract ideas. For instance, we understand time as something that "flows" or "flies," treating it like a river or a bird. These metaphors are not just poetic flourishes; they are the very structure of our thinking. In the 1970s, urban planner Donald Schön applied this to design, arguing that the way we frame a problem determines the solutions we can see. If we view a city block as a "machine," we design for efficiency; if we view it as a "garden," we design for growth. These ideas form the backbone of how humans interact with design, but they create a unique challenge for machines.

The paper then turns to the specific mechanics of how text-to-image and text-to-3D tools operate. Recent studies show that these systems can generate building plans, bridge structures, and even vernacular housing styles based on text descriptions. However, the results often reveal a gap between human intent and machine output. For example, when researchers asked AI to generate traditional houses from Turkey, Japan, and Mexico, the models successfully copied the visual style—the shapes of roofs and the arrangement of windows—but they often missed the deeper cultural logic and social meaning behind those forms. The AI produced visually plausible images that were semantically thin, lacking the rich context that a human architect would understand. Another study found that when architects used these tools without a structured method, the outputs were often just random collections of shapes that looked like buildings but made no sense functionally. The researchers suggest that to get better results, architects need to treat prompting not as a simple command, but as a sophisticated act of communication that requires understanding the hidden rules of the machine.

A significant portion of the paper focuses on the hidden layer of language: the annotations that feed these systems. The most famous dataset used to train image-generating AI, called ImageNet, contains over 14 million images organized into more than 20,000 categories. These categories were not chosen randomly; they were based on a dictionary of words called WordNet, which organizes concepts hierarchically. This means the AI learned to see the world through a specific, human-made lens that includes our biases and assumptions. The paper highlights that this dataset classified people in offensive ways, grouping them under psychiatric diagnoses or slurs, because the original human labels contained those prejudices. When an AI generates an image of a "family" or a "professional," it is drawing on these flawed categories. The machine does not know what a family is; it only knows how the word "family" was statistically linked to certain images in its training data. This reveals that the AI is not a neutral observer but a mirror reflecting the specific, often biased, way humans have organized their knowledge.

The author concludes that for architecture to move forward with these tools, the field must take language seriously as a design material. This means moving beyond the idea that a prompt is just a simple instruction. Instead, architects need to understand the "embodiment" of the machine—the fact that it processes information differently than a human brain does. While humans experience the world through their bodies and senses, AI systems experience the world through data and patterns. The paper suggests that future research should look closely at how different languages and cultures shape design thinking, and how these differences might be lost or distorted when translated into machine code. It calls for a new approach that combines the precision of computer science with the nuance of human communication. By studying how words shape our mental images and how those images are translated into code, architects can begin to harness the full potential of generative AI, ensuring that the buildings of the future are not just visually striking, but deeply meaningful and culturally resonant.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →