← Latest papers
📄 social_science

Computational spatiality: virtual production, generative media and the politics of screen space

This paper introduces the concept of "computational spatiality" to analyze how virtual production and generative media, despite their distinct technical origins, converge on screen to transform spatial construction and shift creative authority from frame composition to the management of parameters, assets, and proprietary infrastructures.

Original authors: Rui Gao, Yiwei Zhao, Weice Yang

Published 2026-09-04
📖 8 min read🧠 Deep dive

Original authors: Rui Gao, Yiwei Zhao, Weice Yang

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

For decades, the magic of cinema relied on a simple division: the camera captured a real world, and editors stitched those pieces together to tell a story. Even when filmmakers built fake worlds with models or painted backgrounds, the final image was the only thing that mattered. The space behind the actors was either a physical place or a flat picture. Today, that division is dissolving. Two powerful technologies are reshaping how screen worlds are built, but they do so in fundamentally different ways. One approach, known as virtual production, treats a digital environment like a physical set that can be walked through and filmed in real time. The other, generative media, uses artificial intelligence to create images that look real but are built from patterns and probabilities rather than fixed blueprints. Understanding the difference between these two methods is no longer just a technical detail for engineers; it changes who gets to decide what a scene looks like, how much control a director truly has, and what kind of reality we are seeing on our screens.

A new study by researchers Rui Gao, Yiwei Zhao, and Weice Yang explores this shift by introducing a concept they call "computational spatiality." Instead of asking whether a scene is real or fake, they ask how the space itself is organized and controlled. The researchers argue that we are moving away from a system where the final image is the only goal. In its place, we now have a system where the space is a living, editable field of possibilities that exists long before the camera starts rolling. This field is not neutral. It is shaped by software, by the people who build the tools, and by the rules embedded in the code. The study suggests that while these technologies offer filmmakers incredible freedom to change a scene instantly, they also create a new kind of dependence on the platforms and systems that make those changes possible.

To understand this, one must look at how these two technologies actually work, because they operate on opposite principles. Virtual production, famously used in films like The Lion King and the Mandalorian, relies on what the researchers call "persistent geometry." Imagine a digital mountain. In this system, that mountain is built from a fixed set of coordinates and 3D shapes. It exists as a solid object in a database. When a camera moves, the computer recalculates the view of that mountain from the new angle, but the mountain itself does not change. Its shape, its distance, and its relationship to the ground remain constant. If a director asks to see the mountain from a different side, the system simply shifts the viewpoint. The world is consistent because it is built on a stable map. This allows actors and cameras to move freely, with the background reacting instantly and accurately to their position.

In contrast, generative media, which includes the latest video creation tools, works without a fixed map. Instead of a pre-built mountain, the system creates an image based on what it has learned from millions of other pictures. When a user asks for a street scene, the system does not pull up a stored 3D model of a street. Instead, it guesses what a street should look like, frame by frame, based on patterns it has seen before. It might create a convincing storefront, a wet pavement, and a parked car. However, because it is guessing rather than retrieving a fixed object, the details can shift. A car might gain a wheel in the next frame, or a doorway might slide to a different spot. The image looks real, but the space behind it is not stable. It is a "probabilistic appearance," meaning it looks like a place because it follows the rules of how places usually look, not because it is a specific, unchanging place.

The researchers found that this difference changes the nature of filmmaking and the power dynamics on a set. In the world of virtual production, the challenge is to keep the digital world perfectly aligned with the physical camera. If the tracking system fails, the background might slide out of sync with the actor, breaking the illusion. The work required here is about precision and coordination. Teams must ensure that the lighting, the camera movement, and the digital assets all match up perfectly. The authority to change the scene lies with the people who manage these assets and the engine that runs them. A director can ask for a sunset, but a technical artist must ensure the digital sun is available and can be rendered quickly enough to keep up with the shoot.

With generative media, the challenge is different. The work is not about keeping things aligned, but about selecting and repairing. Because the system generates a new version of the scene every time, the filmmaker must sift through many options to find the one that works. If a character walks through a door that disappears in the next shot, the artist has to fix it. The power here lies in the ability to choose which version becomes the final image. However, the researchers note that this freedom comes with a hidden cost. The "space" the filmmaker is working in is defined by the training data and the rules of the software. If the software has never seen a specific type of village, it cannot easily create one. The filmmaker is free to ask for anything, but the system decides what is possible to generate.

This leads to a central finding of the study: the image is no longer just a surface to be read; it is the public face of a long chain of decisions. In the past, a director decided the look of a scene, and the crew built it. Now, the look of a scene is determined by a combination of the director's choice, the software's capabilities, the availability of digital assets, and the rules of the platform. The researchers call this "computational spatiality." It is the field where space is organized, changed, and governed before it ever reaches the screen. This field is political because it determines who gets to decide what a world looks like. If a filmmaker wants to change a valley, they might be limited by what the software library contains or what the artificial intelligence model has learned.

The study also highlights how this shift affects the labor of filmmaking. In virtual production, teams must build massive digital libraries and maintain complex tracking systems before a single frame is shot. In generative workflows, the labor shifts toward selecting, repairing, and managing the outputs of the AI. The researchers point out that while these tools make it easier to create a single striking image, they make it harder to maintain a consistent world over a long movie. A generated street might look perfect for two seconds, but if the scene needs to be filmed from a different angle later, the system might create a street that doesn't match the first one. The filmmaker then has to do extra work to make the two versions fit together.

The researchers are careful not to say that one method is better than the other. Instead, they show that they are two distinct ways of making a screen world hold together. One uses a stable map to ensure consistency, while the other uses patterns to create the appearance of a world. Both methods move the power of creation away from the final frame and into the systems that build the space. This means that the freedom to change an image instantly often comes with a stronger dependence on the companies that own the software, the models, and the data. A director might have the power to change a sky in seconds, but they cannot change the engine that makes the sky possible.

The study concludes by suggesting that we need to pay attention to the invisible rules that govern these digital spaces. Just as a physical set has limits based on its construction, a digital space has limits based on its code and its data. These limits are not always obvious. They might appear as a glitch where a building changes shape, or they might appear as a subtle bias where certain types of landscapes are easier to create than others. By understanding "computational spatiality," we can see that the screen is not just a window into a story, but a window into a system of control. The world we see on screen is the result of a negotiation between human creativity and the rigid structures of the software that makes it possible. The researchers argue that to truly understand modern cinema, we must look beyond the image and ask who built the space, how it is kept together, and who gets to decide what survives.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →