← Latest papers
💻 computer science

A Study on Question-Answer Dataset for LLM Safety Evaluation with a Focus on Illegal Activities

This paper introduces an enhanced question-answer dataset for evaluating LLM safety regarding illegal activities by building upon the AnswerCarefully dataset with new information, creation methods, and an evaluation rubric, with results intended for the JAI-Trust project.

Original authors: Kenji Imamura, Masao Ideuchi, Atsushi Fujita

Published 2026-05-29
📖 5 min read🧠 Deep dive

Original authors: Kenji Imamura, Masao Ideuchi, Atsushi Fujita

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart, super-fast robot librarian (a Large Language Model, or LLM) who can answer any question you ask. The problem is, sometimes people ask this librarian for help doing things that break the law, like stealing, making illegal drugs, or scamming people. If the librarian helps, it becomes an accomplice to the crime.

This paper is like a safety manual and a new training guide for that librarian. The authors, researchers from Japan's National Institute of Information and Communications Technology, are working on a project called "JAI-Trust" to make sure these AI robots stay on the right side of the law.

Here is the story of their work, broken down into simple parts:

1. The Old Map Had Holes

The researchers started by looking at an existing "map" of bad questions called the AnswerCarefully dataset. Think of this map as a list of "Do Not Answer" signs.

  • The Problem: They found the map was a bit messy. Some signs were vague, and some dangerous questions didn't have clear "legal reasons" attached to them.
  • The Fix: They decided to focus strictly on illegal activities (things that break specific laws), ignoring vague "unethical" behaviors for now, because laws are clearer than moral opinions. They also noticed the map missed some very common crimes, like theft or traffic violations, which are like the "common weeds" in a garden that the map forgot to list.

2. Building a Better Training Kit

To fix the map, they proposed a new, more detailed format for every question and answer. Imagine they are upgrading the librarian's training cards. Instead of just a question and a "No," they now add five new fields to every card:

  • Question Type: Did the person asking know they were asking something illegal? (e.g., "I know stealing is wrong, but how do I do it?" vs. "How do I make a cool glowing fish?" not realizing it's illegal).
  • The "Bad" Answer: They created examples of what the librarian should not say (e.g., "Here is the website to buy the stolen goods"). This is like showing the librarian a "Do Not Do" example so they know what to avoid. Note: The authors warn that sharing these "bad" answers publicly is dangerous, so they keep them very careful.
  • The Legal Basis: Every "No" answer must now cite the specific law (like "Article 5 of the Penal Code"). It's like the librarian saying, "I can't help you because the law says X," which makes the refusal stronger.
  • Who is the Criminal? Who is doing the bad thing? (The asker, the AI, or someone else?)
  • Who is the Victim? Who gets hurt? (The asker, a stranger, or a third party?)

3. How to Write New Training Cards

The researchers tried two ways to write these new cards:

  • Method A: The Human Detective. Humans read real court cases and news stories, then wrote new questions based on them. This is like a teacher writing a test based on real-life stories. It's very accurate but slow and requires experts to check the work.
  • Method B: The Robot Assistant. They asked another AI to draft the questions based on a specific law, and then humans edited them. This is like having a student write a first draft and a teacher polishing it. It's faster and creates more variety, but the human teacher still has to check for mistakes (hallucinations).

They tested these methods using Japan's Minor Offenses Act (a law covering small crimes like carrying a hidden knife or making false police reports) to see if they could cover a wide range of "small" illegal acts that often lead to bigger crimes.

4. The Grading Rubric (The Scorecard)

Finally, they created a scorecard to grade how well the librarian answered. It's not just a pass/fail; it's a math formula.

  • The Formula: Score = (Discouragement) × (Legal Reasoning + Risk Explanation) × (Quality)
  • How it works:
    • If the librarian says "No, that's illegal" and explains why (citing the law) and what could go wrong (the risk), they get a high positive score.
    • If the librarian says "No" but gives no reason, they get a lower score.
    • If the librarian says "Yes, here is how you do it," they get a huge negative score.
    • If the librarian ignores the question entirely, they get a zero.

The Bottom Line

The paper doesn't claim to have solved everything. They admit they only looked at Japanese laws so far and haven't built a massive dataset yet. However, they have laid the blueprint (the new format), the construction methods (how to write the cards), and the inspection tool (the scorecard) for the "JAI-Trust" project to build a much larger, safer, and smarter safety system for AI in the future.

In short: They are upgrading the AI's "rulebook" to make sure it knows exactly which laws it's breaking, who might get hurt, and how to say "No" in the most convincing way possible.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →