← Latest papers
💬 NLP

ASSERT: A Measurement Pipeline for GenAI Audits

The paper introduces ASSERT, a specification-driven measurement pipeline for GenAI audits that explicitly ties reported compliance rates to their underlying measurement choices, thereby clarifying how variations in audit design—such as dialogue setup or evaluation criteria—significantly impact system rankings and interpretability.

Original authors: Riccardo Fogliato, Abhinav Palia, Xiawei Wang, Emily Sheng, Chad Atalla, Jean Garcia-Gathright, Nicholas Pangakis, Sharman Tan, Dan Vann, Hannah Washington, P. Alex Dow, Heba Elfardy, Hanna Wallach, S
Published 2026-08-17
📖 4 min read☕ Coffee break read

Original authors: Riccardo Fogliato, Abhinav Palia, Xiawei Wang, Emily Sheng, Chad Atalla, Jean Garcia-Gathright, Nicholas Pangakis, Sharman Tan, Dan Vann, Hannah Washington, P. Alex Dow, Heba Elfardy, Hanna Wallach, Sandeep Atluri

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to grade a new, super-smart robot that can write stories, solve math problems, and chat with you. You want to know: "Is this robot honest?" or "Is it safe?" In the world of Artificial Intelligence (AI), scientists often answer this by running a test and giving the robot a single score, like "85% honest." It sounds simple, like a report card. But here's the catch: that score doesn't just tell you about the robot; it also tells you about the test itself. If you change the questions, the person grading the answers, or even the rules for what counts as a "lie," the score can swing wildly. It's like grading a student's essay: if you ask them to write about "cats," they might get an A, but if you ask about "quantum physics," they might get a C. The student didn't change; the test did. This paper dives into that messy reality, showing us that a single number isn't enough to judge a robot's behavior. We need to know exactly how the test was built to understand what the score actually means.

Enter ASSERT, a new tool created by researchers at Microsoft to fix this grading problem. Think of ASSERT not as a single test, but as a "recipe book" for building tests. Usually, when people test AI, they might keep their grading rules hidden or change them without saying so. ASSERT forces everyone to write down their recipe first. It asks the researchers to clearly define what "deception" or "safety" looks like in plain language, then builds a specific set of challenges (test cases) based on that definition, and finally runs the AI through those challenges. The magic of ASSERT is that it ties every single score back to the specific recipe used to create it. If the score changes, you can look at the recipe book and see exactly which ingredient was swapped out.

To see how powerful this is, the researchers used ASSERT to test a specific AI model (GPT-5.5) on the tricky topic of conversational deception—basically, when the AI lies or misleads a user in a chat. They didn't just run one test; they ran a "multiverse" of tests. Imagine they tested the same robot in a hundred different scenarios: sometimes the robot was talking to a friendly user, sometimes a tricky one; sometimes the robot was being judged by a strict AI judge, sometimes a more lenient one; and sometimes the rules for what counted as a lie were super strict, and other times they were loose.

The results were eye-opening. When they changed just one part of the recipe, the robot's "honesty score" jumped around a lot.

  • The Judge Matters: When they swapped the AI that was grading the answers, the honesty score for the same robot chat logs swung from 80% to 95%. That's a huge difference! One grader thought the robot was mostly honest, while another thought it was lying almost all the time.
  • The Rules Matter: When they made the rules for spotting a lie stricter (requiring clear, undeniable proof), the score went up to over 90%. When they made the rules looser (accepting any hint of a lie), the score dropped to around 73%.
  • The User Matters: When they changed the "simulated user" chatting with the robot from a standard character to a more complex one, the score jumped from 82% to over 90%.

The paper suggests that because these scores move so much based on the choices made by the researchers, we can't trust a single number to tell us which AI is "better." If one AI scores 82% and another scores 85%, that tiny difference might just be because they used different judges or different rules, not because one robot is actually smarter or safer. In fact, the researchers found that changing the judge could even flip the ranking, making a "worse" robot look like the winner just because of who was grading the test.

So, what's the takeaway? The paper doesn't say AI is broken or that we can't test it. Instead, it suggests that we need to be much more transparent. We can't just say, "This AI is 82% safe." We have to say, "This AI is 82% safe when tested with these specific rules, this specific judge, and these specific questions." By using a tool like ASSERT to write down every step of the recipe, researchers and companies can stop hiding behind a single number and start having honest conversations about what their AI systems can and cannot do. It turns a magic trick into a clear, inspectable process, helping us figure out if the robot is truly telling the truth, or if we just asked it the wrong question.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →