SIG: A Synthetic Identity Generation Pipeline for Generating Evaluation Datasets for Face Recognition
This paper introduces the Synthetic Identity Generation (SIG) pipeline, a method for creating ethically sourced, balanced, and controllable synthetic face datasets to overcome the logistical and privacy challenges of traditional data collection, and validates its effectiveness through the release of the open-source ControlFace10k dataset for evaluating face recognition algorithms and bias.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the modern world, cameras and computers are increasingly asked to do the work of human eyes, identifying people in airports, stadiums, and at borders with speed and reliability. For these systems to work fairly and accurately, they must be tested against a wide variety of faces, representing different ages, skin tones, and genders, and captured from many different angles. However, gathering such a diverse collection of real human photos is fraught with difficulty. Collecting images with the consent of every person involved is slow and expensive, while scraping photos from the internet without permission raises serious ethical and legal concerns. This creates a gap: researchers need balanced, high-quality data to test if their systems are biased, but they cannot easily get it from the real world.
To bridge this gap, a team of researchers at Southern Methodist University has developed a new way to create the data they need from scratch. They built a pipeline, which they call SIG, that uses advanced image-generation technology to produce realistic pictures of people who do not exist. These synthetic faces are not random; the system allows the creators to specify exactly who the person should look like, controlling their age, race, gender, and even the angle of their head. By generating these images, the team has produced a new dataset called ControlFace10k, which contains over 10,000 images of more than 3,300 unique synthetic individuals. This dataset is perfectly balanced, with equal numbers of people from different racial backgrounds, ages, and genders, offering a clean, ethical tool to test how well face recognition systems perform across different groups.
The process begins with a digital architect that writes detailed descriptions, or prompts, for the image generator. Instead of just asking for a "face," the system combines specific names and attributes to create a unique identity. It then uses a powerful image-generation engine, which works by starting with a noisy, static-like image and gradually refining it until a clear face emerges, guided by the text description. To ensure the faces appear in specific poses—looking left, right, or straight ahead—the system uses a control mechanism that acts like a skeleton, mapping the desired body position onto the new image. This allows the researchers to generate the same person looking in three different directions, ensuring that the system is tested on how well it recognizes a single identity despite changes in posture.
The result is ControlFace10k, a dataset containing 3,336 unique synthetic identities. Each identity is represented by images of a person at three different ages—25, 50, and 65 years old—and across four major racial groups: African, Asian, Caucasian, and Indian. The dataset is also split evenly between men and women. Unlike many previous attempts to create fake faces, which often resulted in images that looked slightly off or failed to maintain a consistent identity across different photos, this pipeline produces hyper-realistic images with ideal lighting and focus. The researchers emphasize that the goal was not to create perfect training data, but to create a controlled environment where they could measure exactly how a face recognition system behaves when faced with specific, balanced variables.
To see if this synthetic data is useful, the team tested it against two of the most advanced face recognition systems currently available. They fed the images into these systems to see how the machines scored the similarity between different faces. The results showed that the synthetic faces behaved very much like real human faces. When the systems compared two different people, the scores were low, indicating the systems correctly identified them as strangers. When the systems compared three different photos of the same synthetic person, the scores were high, showing the system recognized the identity. However, the tests also revealed that the systems still struggled when the same person was shown from different angles, a common challenge in real-world scenarios.
The analysis also highlighted subtle differences in how the systems treated different racial groups. For instance, one of the tested systems gave slightly higher similarity scores to faces of African and Indian individuals compared to others, suggesting a potential bias in how that specific algorithm processes those features. Because the dataset was perfectly balanced, these patterns stood out clearly, whereas they might have been hidden in a messy, real-world dataset where some groups are underrepresented. This demonstrates the value of the synthetic approach: it provides a clear, unbiased mirror to reflect the strengths and weaknesses of the technology.
The researchers conclude that while their synthetic dataset is not a replacement for all real-world data, it is a powerful tool for evaluation. It allows scientists to test specific scenarios, such as how a system handles aging or different poses, without the ethical and logistical hurdles of collecting real photos. By making this dataset freely available, the team hopes to encourage a more rigorous and fair approach to developing face recognition technology, ensuring that these systems work reliably for everyone, regardless of their appearance or the angle at which they are viewed.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.