← Latest papers
💻 computer science

The CitizenQuery Benchmark: A Novel Dataset and Evaluation Pipeline for Measuring LLM Performance in Citizen Query Tasks

This paper introduces CitizenQuery-UK, a novel benchmark dataset of 22,000 synthetically generated citizen queries based on UK government information, and evaluates 11 Large Language Models on factuality, abstention, and verbosity to highlight critical reliability challenges and the need for acknowledging AI fallibility in public sector applications.

Original authors: Neil Majithia, Rajat Shinde, Zo Chapman, Prajun Trital, Jordan Decker, Manil Maskey, Elena Simperl, Nigel Shadbolt

Published 2026-02-05
📖 5 min read🧠 Deep dive

Original authors: Neil Majithia, Rajat Shinde, Zo Chapman, Prajun Trital, Jordan Decker, Manil Maskey, Elena Simperl, Nigel Shadbolt

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a regular person trying to figure out a confusing government rule, like how to claim a benefit or fill out a tax form. In the past, you might have called a help desk or searched a government website. Today, you might ask a super-smart AI chatbot.

This paper is about building a giant "report card" to test how well these AI chatbots handle those specific questions about UK government rules. The authors call this report card CitizenQuery-UK.

Here is the breakdown of what they did and what they found, using simple analogies:

1. The Problem: The "Trust Me" Trap

Imagine you ask a very confident, fast-talking AI for advice on your taxes. It gives you a long, detailed answer. But what if it made up a rule that doesn't exist? In a normal chat, that's just a funny mistake. But with government rules, a made-up fact could cost you money, get you in trouble with the law, or ruin your life.

The authors say: "We can't just trust the AI because it sounds smart. We need proof it's telling the truth."

2. The Solution: Building the "Exam" (The Dataset)

To test the AI, you need a test. But you can't just make up questions; they have to be real.

  • The Source: They took thousands of pages from gov.uk (the official UK government website), which is like the "bible" of government rules.
  • The Questions: They didn't just write robotic questions. They used real examples from people asking for advice on Reddit to make the questions sound like how real humans actually talk (messy, specific, and sometimes confused).
  • The Characters: They created "personas" for the questions. For example, a question about parenting is assigned to a "parent" persona, not a "child." This ensures the AI is tested on whether it understands the context of the person asking.
  • The Result: They built a dataset of 22,000 question-and-answer pairs. Think of this as a massive, 22,000-question final exam where the "correct answers" are already written down in the government's official text.

3. The Grading System: How They Checked the Work

They didn't just ask, "Did the AI get it right?" They used a three-step grading process to be super precise:

  • Step 1: The "Silent Treatment" Check (Abstention): Does the AI say "I don't know" when it's unsure? Or does it just guess? The authors wanted to see if the AI was brave enough to admit it didn't know, or if it was too eager to talk.
  • Step 2: Breaking it Down (Atomization): Imagine the AI writes a long paragraph. The graders broke that paragraph down into tiny, individual facts (like breaking a Lego castle into single bricks).
  • Step 3: The "Truth Match" (Verification): They compared every single "brick" (fact) the AI said against the "bricks" in the official government answer.
    • The Score (F1@K): They gave the AI a score based on two things: Did it include the right facts? And did it only include the right facts (without adding too much extra fluff)?

4. The Results: What the Report Card Said

They tested 11 different AI models (like the ones from OpenAI, Google, Meta, etc.) and found some interesting things:

  • The "Chatty" Problem: Almost all the AIs talked way too much. They were like a student who, when asked a simple question, wrote a whole essay. They added extra facts that weren't in the official government text. This is called verbosity.
  • The "Silent" Problem: The AIs almost never said "I don't know." Even when they were wrong, they kept talking. They rarely "abstained" from answering.
  • The "Long Tail" of Errors: Most of the time, the AIs got the main points right. But when they got it wrong, they got it very wrong. It's like a student who gets an A on the easy questions but fails the hard ones completely. This makes the results unpredictable.
  • Open vs. Closed: Some of the "open" models (where the code is public) performed just as well as the expensive, "closed" ones. You don't necessarily need the most expensive AI to get good results.

5. The Big Takeaway

The paper concludes that while AI is getting better, it is currently too eager to talk and too afraid to admit it's wrong for high-stakes government advice.

If an AI is going to be your guide for government rules, it needs to learn two things:

  1. Be concise: Stick to the facts on the government website; don't make up extra stories.
  2. Be humble: If it doesn't know the answer, it should say "I don't know" instead of guessing.

The authors built this "exam" (CitizenQuery-UK) so that governments and tech companies can keep testing AI to make sure it becomes safe and reliable enough to actually help people with their real-life problems. They are not saying AI is ready to replace human advice workers yet; they are saying, "Here is a ruler to measure how close we are to being ready."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →