← Latest papers
💬 NLP

DialogPII: A multilingual dataset of synthetic dialog transcripts to detect personal information

The paper introduces DialogPII, a multilingual dataset of synthetic dialog transcripts and corresponding speech-derived resources covering 11 languages and 19 entity types across eight interaction scenarios, designed to support the development and evaluation of automatic de-identification systems for protecting privacy in conversational data.

Original authors: Roland Roller, Vera Czehmann, Derya Erman, Luke Flanagan, Ibrahim Baroud, Frédéric Blain, Viviana Cotik, Eletta Giusto, Akhil Juneja, Mariana Neves, Maria Słowińska, Christine Hovhannisyan, Aaron Loui
Published 2026-06-30
📖 4 min read☕ Coffee break read

Original authors: Roland Roller, Vera Czehmann, Derya Erman, Luke Flanagan, Ibrahim Baroud, Frédéric Blain, Viviana Cotik, Eletta Giusto, Akhil Juneja, Mariana Neves, Maria Słowińska, Christine Hovhannisyan, Aaron Louis Eidt, Lisa Raithel, Sebastian Möller, Maija Poikela

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a giant library of conversations—people talking about their health, calling the police, or chatting with customer support. These talks are gold mines for researchers trying to teach computers how to understand human speech. But there's a catch: these conversations are full of "secret ingredients" like real names, phone numbers, home addresses, and medical conditions. If you share the library without hiding these secrets, you accidentally expose people's private lives.

This paper introduces DialogPII, a new, massive "training gym" designed to teach computers how to spot and hide those secret ingredients before anyone shares the data.

Here is how they built it, explained simply:

1. The "Fake" but Realistic Library

You can't just grab real people's private phone calls and share them; that's illegal and unethical. So, the researchers used a "digital chef" (a Large Language Model) to cook up synthetic conversations.

Think of this like a TV writer's room. Instead of filming real people, they wrote scripts for 8 different types of scenes:

  • Emergency calls (like 911).
  • Doctor's office visits.
  • Therapy sessions.
  • Police reports.
  • Customer support chats.
  • And more.

They created 1,029 conversations (about 147 for each of the 11 languages they support). To make sure the "actors" didn't sound like robots, human editors stepped in to fix awkward phrasing, change repetitive names, and make sure the cultural details (like street names or local customs) felt real for each country.

2. The "Red Pen" Game (Annotation)

Once the scripts were written, the team played a game of "Spot the Secret." They went through every line and highlighted anything that could identify a person.

  • Names? Highlighted.
  • Phone numbers? Highlighted.
  • "My cousin Bob" (a relationship)? Highlighted.
  • "I work at the local hospital" (a job)? Highlighted.

They found 19 different types of secrets to look for. They did this in 11 languages (including English, German, Arabic, Hindi, and Turkish), making sure the "Red Pen" rules worked everywhere.

3. The "Robot Voice" Challenge

Here is the clever part. Most privacy tools only look at text. But people speak!

  1. The team took their written scripts and used a computer to read them out loud (Text-to-Speech).
  2. Then, they used a different computer program to listen to that audio and write it back down (Speech-to-Text).

This created a "messy" version of the conversation. Just like real speech-to-text software, this version had typos, missed words, and weird pauses. The researchers then took their "Red Pen" highlights from the clean text and tried to map them onto this messy audio transcript. They manually fixed the mistakes to create a perfect "answer key" for training computers to find secrets even in noisy, spoken data.

4. The "Test Drive"

To prove their new gym works, they built a basic computer brain (a model) and trained it on these fake conversations.

  • The Result: The computer got really good at finding secrets in the clean text (about 90% accuracy).
  • The Harder Test: When they tested it on the messy, speech-derived transcripts, the accuracy dropped a bit (to about 82%), which is expected because the audio transcription makes the job harder.
  • The Real-World Check: They even tested it on a small set of real (but anonymized) phone calls from a different database. The computer still did a decent job, proving it learned general rules, not just memorized the fake scripts.

Why This Matters

Think of DialogPII as a safe, legal, and diverse "flight simulator" for privacy.

  • Before: Researchers had to use real, sensitive data (risky) or very small, boring datasets (not useful).
  • Now: They have a massive, multilingual, multi-scenario dataset that is 100% safe to share.

This allows developers to build better "privacy shields" for AI systems. Once these shields are trained on DialogPII, they can be used to protect real people's data in hospitals, call centers, and therapy apps, ensuring that when we share data for science, we don't accidentally share our secrets.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →