← Latest papers
🤖 AI

A Large-scale Evaluation of Text-guided Models for Facial Editing

This paper presents the first large-scale evaluation of six text-guided models for facial editing, introducing a new dataset of 169 attributes and revealing that while these models excel at hair and accessory modifications, they struggle with pose changes and exhibit significant demographic biases in over-editing dark-skinned male and older faces.

Original authors: Rahul Nair, Saurav Pandit, Hannah Kerner

Published 2026-09-01
📖 4 min read☕ Coffee break read

Original authors: Rahul Nair, Saurav Pandit, Hannah Kerner

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a digital artist who can change a person's hair color, add a pair of glasses, or turn their head to look at the horizon, all by typing a simple sentence into a computer. This is the promise of text-guided image editing, a rapidly growing field where artificial intelligence learns to understand human language and apply those instructions to photographs. For years, researchers have built systems capable of these feats, but they often struggled with a specific challenge: keeping the person in the photo looking like themselves while making the requested changes. Older methods could either change the face too much, making it unrecognizable, or they could only make very limited adjustments, like changing the angle of the head. Now, a new generation of tools has emerged that claims to solve these problems, offering the ability to make complex, realistic edits based on text prompts. But as these tools become more powerful, a critical question remains: do they work equally well for everyone, or do they have hidden flaws that treat different groups of people unfairly?

A team of researchers set out to answer this question by putting six of the most popular text-guided editing models to the test. They did not just ask the computers to make a few random changes; they conducted a massive, systematic evaluation involving nearly one million image edits. The goal was to see how well these models could handle a sequence of four different instructions in a row, such as changing a person's hair color, then adding a hat, then changing the hat, and finally turning their head. To do this, the researchers created a new, extensive library of 169 specific editing tasks focused on hair, accessories, and head position. They tested the models on two large collections of celebrity photos, ensuring a diverse mix of men and women, young and old, and people with light and dark skin tones.

The results revealed a clear picture of where these technologies stand today. The models proved quite capable of handling changes to hair and accessories. When asked to change a hairstyle or add a pair of sunglasses, most of the systems performed well, successfully following the instructions while keeping the person's face recognizable. However, the same models struggled significantly when asked to change a person's pose, such as making them look to the left or right. In these cases, the systems often failed to follow the command or altered the face in ways that made the person unrecognizable.

Perhaps more concerning was the discovery that all the models tended to "over-edit." This means that when given a specific instruction, the computer would often change things that were not asked for. For example, if a user asked the model to change a person's hairstyle to a bob cut, the model might also change the hair color to brown, even though the user never requested a color change. The researchers found that this happened frequently, with the models making about 25 extra, unwanted changes for every 100 instructions given. These errors were not random; they often stemmed from the models learning incorrect associations, such as believing that a specific hairstyle always goes with a specific hair color.

The study also uncovered significant biases in how these models treated different people. The over-editing was not distributed evenly across all faces. The models were much more likely to make unwanted changes when editing the faces of men, particularly those with dark skin tones, and when editing the faces of older people. For instance, when asked to edit a dark-skinned man's hair, the system was far more likely to accidentally change other features, like adding facial hair or altering the skin tone, compared to when it edited a light-skinned woman. Similarly, the models were less successful at preserving the identity of older faces, often making them look younger or changing their features more drastically than those of younger people.

These findings suggest that while text-guided editing tools are becoming powerful enough for real-world use, they are not yet perfect. They work well for simple tasks like adding a hat or changing hair color, but they still struggle with more complex movements and carry hidden biases that can lead to unfair or inaccurate results for certain groups. The researchers emphasize that future work must focus on fixing these over-editing habits and ensuring that the technology treats every face with the same level of accuracy and respect, regardless of age, gender, or skin tone. Until these issues are resolved, users should be aware that these digital artists are still learning, and they may sometimes change more than they are asked to.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →