← Latest papers
🧬 biology

Beauty is in the ELBO of the Beholder: A Variational Account of Processing Fluency in Face Perception

This paper demonstrates that human judgments of facial attractiveness align with the evidence lower bound (ELBO) of variational autoencoders trained without supervision, providing empirical support for the theory that aesthetic pleasure arises from the processing fluency of statistically regular, prototypical faces.

Original authors: Francisco M. López, Jochen Triesch

Published 2026-08-26
📖 4 min read☕ Coffee break read

Original authors: Francisco M. López, Jochen Triesch

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

For decades, scientists have wondered why some faces captivate us while others fade into the background. Theories of human perception have long suggested that beauty is not just a matter of personal taste, but is deeply rooted in how easily our brains can process what we see. When a face is symmetrical, average, or familiar, our minds seem to recognize it with less effort, a phenomenon known as processing fluency. This ease of understanding is thought to generate a feeling of pleasure, leading us to rate those faces as more attractive. While this idea has been tested in countless behavioral experiments with human subjects, it has remained unclear how this mental ease actually emerges inside the complex machinery of learning. Can a computer, trained only to see faces without ever being told which ones are beautiful, develop a similar sense of preference?

To answer this, researchers Francisco López and Jochen Triesch turned to a type of artificial intelligence called a variational autoencoder. These are machines designed to learn the underlying structure of data, such as thousands of photographs of human faces, without any human guidance or labels about beauty. The system works by trying to compress a face into a simplified internal code and then reconstructing it from that code. To do this successfully, the machine must balance two competing goals: it needs to keep enough detail to recreate the face accurately, but it also needs to keep the internal code simple and organized. The researchers proposed that the mathematical score used to measure how well the machine balances these two goals is a direct, computational version of human processing fluency. If a face is easy for the machine to understand and reconstruct, it should score highly on this measure.

The team trained four separate AI models on different large collections of face images, ensuring that none of the models ever saw a human rating of attractiveness. They then tested these models on a standard set of 597 faces from the Chicago Face Database, for which human attractiveness ratings were already known. The results were striking. The faces that humans rated as most beautiful were consistently the same faces that the machines found easiest to process. In the mathematical space where the machines organize faces, the direction of increasing beauty aligned almost perfectly with the direction of increasing processing ease. This alignment held true regardless of which dataset the machine was trained on or which specific demographic group the faces belonged to. The machine did not need to be taught what beauty was; it discovered a path toward it simply by learning to see faces efficiently.

Furthermore, the researchers found that this preference was not just a single number but a specific direction in the machine's internal map. When they compared the maps created by different models, they found that the "beauty direction" was remarkably consistent. Even though the models were built from scratch with different starting points and different training data, they all organized the faces in a way that pointed toward the same concept of attractiveness. This suggests that the human sense of beauty is not a random cultural construct but is tied to a fundamental, reproducible structure in how faces are represented. The study also confirmed that attractive faces tend to be more "prototypical," meaning they look more like the average face of their group, both in their physical shape and in the machine's internal code. However, the machine's ability to predict beauty went beyond simple averages, capturing subtle details that a basic geometric average might miss.

The findings offer a new bridge between the old theories of aesthetics and modern artificial intelligence. They suggest that the pleasure we feel when seeing a beautiful face may stem from the same efficiency that allows a machine to compress and reconstruct an image with minimal error. The machine's success in predicting human preferences without any human instruction implies that our attraction to certain faces is deeply connected to the ease with which our minds can make sense of them. While the study does not claim that human brains work exactly like these computer models, it provides strong evidence that the principles of efficient representation are a key part of the story. The research concludes that beauty, in the context of face perception, is indeed in the eye of the beholder, but that eye is guided by a deep, shared logic of how information is processed and understood.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →