← Latest papers
💬 NLP

DALDALL: Data Augmentation for Lexical and Semantic Diverse in Legal Domain by leveraging LLM-Persona

The paper introduces DALDALL, a persona-based data augmentation framework that leverages domain-specific legal roles to generate high-quality, diverse synthetic queries, thereby improving the performance of dense retrievers in low-resource legal information retrieval tasks.

Original authors: Janghyeok Choi, Jaewon Lee, Sungzoon Cho

Published 2026-03-25
📖 4 min read☕ Coffee break read

Original authors: Janghyeok Choi, Jaewon Lee, Sungzoon Cho

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to be a legal detective. Its job is to read thousands of pages of court cases and find the exact document that answers a specific legal question.

The problem? There aren't enough "practice questions" (data) for the robot to learn from. In the real world, legal data is scarce, expensive to label, and very hard to find. It's like trying to teach someone to play chess by only showing them three games.

This paper introduces DALDALL, a clever new way to create more practice questions using Artificial Intelligence (AI), but with a special twist: Role-Playing.

The Problem: The "Copycat" Robot

Usually, when we ask an AI to make up new practice questions, it acts like a parrot. If you show it one example of a lawyer asking a question, the AI just copies the words and changes a few here and there.

  • Result: The robot learns to recognize only those specific words. If a real lawyer asks the same question using different words, the robot gets confused and fails.

The Solution: The "Method Acting" Approach

The authors of this paper realized that in the legal world, everyone speaks differently depending on their job.

  • A Judge speaks formally and looks for the final ruling.
  • A Prosecutor is aggressive and focuses on breaking the law.
  • A Defense Attorney is protective and looks for loopholes.
  • A Law Professor explains the theory behind the law.

Instead of asking the AI to "make up a question," they told the AI: "Pretend you are a Prosecutor. Now, ask a question about this case." Then they said, "Now pretend you are a Judge. Ask the same question again."

This is called Persona-Based Prompting. It's like hiring a group of actors to improvise a scene. Even though they are all acting out the same plot, they all say it in their own unique voice, using different vocabulary and focusing on different details.

How It Works (The Recipe)

  1. The Seed: They take one real legal case (the "seed").
  2. The Extraction: They pull out the most important facts (the "essentials") so the AI doesn't accidentally change the truth of the case.
  3. The Role-Play: They feed these facts to the AI and say, "Act like a Prosecutor," then "Act like a Judge," etc.
  4. The Harvest: The AI spits out many different versions of the question. One might be very technical, another very emotional, another very concise.

The Results: Why It Matters

The researchers tested this on two big legal datasets (think of them as giant libraries of court cases).

  • More Variety: The "Role-Playing" method created questions that were 20% more diverse in vocabulary than the standard "parrot" method. It was like the robot learned to speak in 20 different dialects instead of just one.
  • Better Memory: When they trained their legal search robot on these new, diverse questions, it got much better at finding the right answers. It didn't just memorize the words; it learned the meaning behind them.
  • The "Long Text" Surprise: They found that the longer the original legal case was, the better the AI got at role-playing. A short case gave the AI little to work with, but a long, complex case gave the "Prosecutor" and "Judge" plenty of material to argue about, creating even more unique questions.

The Bottom Line

Think of DALDALL as a gym for legal AI.

  • Old Way: The AI lifts the same heavy weight (the same few words) over and over. It gets strong at lifting that weight, but weak at everything else.
  • New Way (DALDALL): The AI lifts weights of different shapes, sizes, and colors (different personas). It becomes a well-rounded athlete capable of handling any legal question thrown its way, even if the question is phrased in a way it has never seen before.

This is a huge win for low-resource fields like law and medicine, where we don't have millions of examples to teach our AI. We just need to teach it to think like different people, and suddenly, we have a whole new world of data to learn from.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →