Data as Transcendental Stabilizations: Symbolic Mediation and the Pre-Semantic Conditions of Data
This paper argues that data are not ontological primitives or mere technical artifacts but emerge through historically variable symbolic stabilizations that transform differential configurations into operationally available differences, thereby reframing issues like algorithmic bias as disputes over competing regimes of intelligibility rather than simple technical management problems.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are holding a bucket of sand. To a geologist, it's a sample of silica and feldspar. To a child, it's a castle-building material. To a poet, it's a metaphor for time slipping through fingers. But before any of those people can say "this is sand," something invisible has to happen first. The chaotic, shifting grains have to be caught, named, and turned into something that can be counted, shared, and remembered. This is the secret life of data.
In the world of science and technology, we often treat data like little bricks of reality that we just pick up off the floor. We assume that if we have a thermometer, the temperature is just "there," waiting to be written down. But a new paper by Rodrigo Garrido suggests this is a bit of a trick. It argues that data isn't something we find; it's something we make. To understand this, we need to look at three big ideas. First, symbolic mediation: the idea that we can't touch reality directly; we always have to use tools, words, or numbers to translate the world into something we can understand. Second, stabilization: the process of taking something messy and changing it so it stays the same long enough to be useful. And third, projection: like shining a flashlight on a 3D object to see a 2D shadow, we have to choose which parts of reality to highlight and which to ignore to make a "datum." Why does this matter? Because if we think data is just "raw truth," we might miss the fact that the way we choose to measure things actually decides what we can see, what we can know, and even what counts as "real" in our digital world.
So, what is this paper actually saying? It proposes that a "datum" (a single piece of data) is not a tiny, pre-existing piece of the universe. Instead, it is a transcendental stabilization. That's a fancy way of saying it's a snapshot of reality that has been frozen, packaged, and made repeatable by a specific set of rules and tools.
Think of reality as a giant, continuous, humming river of information. It's always flowing, changing, and full of infinite differences. You can't just scoop up the river and put it in a jar; it would just spill out. To get "data," you have to build a dam. You have to decide where to cut the river, what to measure, and how to label it. The paper argues that this act of building the dam is what creates the data. Before the dam is built, there is just the river (ontological variation). Once the dam is built and the water is channeled into a specific pipe, that is the data.
The author uses some cool metaphors to explain how this works. Imagine you are trying to describe a song. The sound waves in the air are the "river." If you want to turn that song into a file on your computer (a datum), you have to chop the sound into tiny, discrete chunks. This is called tokenization in the world of computers. You can't feed the whole continuous song into a machine; you have to break it into little pieces (tokens) that the machine can hold. But here's the kicker: the way you chop it up changes the song. If you chop it by the beat, you get one kind of data. If you chop it by the pitch, you get a totally different kind. The paper says the "data" isn't the song itself; it's the specific way you decided to chop it up.
The paper also argues that this isn't just a one-time thing. Once you chop the song up, you have to keep those pieces from falling apart. This is where technical sedimentation comes in. Imagine you write a recipe on a piece of paper. That paper is a "technical support." It holds the recipe so you can read it tomorrow. But the paper itself has rules: it has lines, it has a certain size, and maybe it's written in a language you know. Those rules from the past (the paper, the language) are "sedimented" into your recipe. They decide what you can write and how you can write it. The paper suggests that all our data is like this. It's built on layers of old tools, old rules, and old ways of thinking that we often don't even notice. A database designed in the 1980s might still be deciding what kind of data we can collect today, simply because of how it was built.
The author is very clear about what data is not. It is not a "raw" piece of reality waiting to be discovered. You can't just find a "datum" in nature. A temperature change exists, sure, but it only becomes a "temperature datum" when a thermometer measures it, a computer records it, and a scientist agrees on how to write it down. The paper also pushes back against the idea that data is just a neutral building block for truth. Instead, it suggests that the way we stabilize data (the rules we use) actually creates the truth we see. If you use a different set of rules (a different symbolic form), you get a different kind of data, and a different kind of truth.
The paper doesn't claim to have found a magic formula to fix all our data problems. Instead, it suggests that we need to change how we think about data. We shouldn't treat data as if it's just "out there." We should realize that data is a product of our choices, our tools, and our history. When we talk about things like "algorithmic bias" or "data privacy," the paper suggests we shouldn't just look at the numbers. We should look at the dam we built. We should ask: Who decided where to cut the river? What rules did we use to chop up the song? And what old, dusty rules from the past are still controlling what we can see today?
In short, the paper invites us to stop looking at data as a mirror of reality and start seeing it as a painting. The painting isn't the reality itself; it's a specific, stabilized, and stabilized version of reality, created by the artist's brush, the canvas, and the history of art that came before. And just like a painting, it tells us as much about the artist and the tools they used as it does about the world they are trying to show.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.