← Latest papers
🤖 AI

SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows

SQBench introduces a novel benchmark for evaluating language-model agents in production-oriented workflows by assessing 220 standardized tasks across three complexity levels through a dual-metric framework that distinguishes functional completion from risk-based penalties, revealing that current models struggle with domain-constrained delivery and that functional success alone is insufficient for ensuring high-quality task outcomes.

Original authors: Summer Sun

Published 2026-07-28
📖 7 min read🧠 Deep dive

Original authors: Summer Sun

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring a super-smart robot assistant to help you run a small business. You don't just want the robot to answer your questions correctly; you want it to actually do the work. You want it to take your messy notes, use your tools, follow strict rules, and hand you a finished report that you can actually use tomorrow. This is the world of "AI agents"—computer programs that don't just chat, but take action. For a long time, scientists have tested these robots by asking them trivia questions or seeing if they can solve a math problem. But in the real world, getting the right answer isn't enough. If the robot forgets to save the file, uses a tool it wasn't allowed to touch, or makes up a source for its data, the whole job is a failure. We need a way to test not just if the robot is smart, but if it is reliable, safe, and actually delivers the goods.

This is exactly what the paper "SQBench" tackles. The authors, led by Summer Sun, built a new testing ground called SQBench to see how well AI agents handle real-world tasks. Instead of just checking if an answer is right, they check if the final "deliverable" (the finished product) is usable and safe. They created 220 different jobs, ranging from simple tasks like "find this file" to complex business scenarios like "analyze these financial records." They tested 27 different AI models to see which ones could actually finish the job without breaking any rules. The big surprise? Even the best AI models are still quite clumsy at these real-world jobs. While the top model got about 60% of the tasks "done," only about 18% of the complex business tasks were finished perfectly without any risky mistakes. The paper shows that being "smart" isn't the same as being "reliable," and that we need to start measuring AI based on how well it delivers safe, usable results, not just how many questions it can answer.

The New Report Card: From "Right Answer" to "Safe Delivery"

Think of the old way of testing AI like a pop quiz in school. If you get the math problem right, you get an A. But in the real world, imagine if you got the math right but wrote the answer on the back of a napkin that got thrown away, or if you used a calculator that wasn't allowed in the exam. You might have the right number, but you failed the assignment.

The authors of this paper argue that we need a new kind of report card. They call it SQBench (which stands for "SQ" for "Shaqiu," the community behind it, and "Bench" for benchmark). They designed this test to see if an AI agent can act like a responsible employee. The test is built on three levels, like a video game with increasing difficulty:

  • Level 1 (L1): The Basics. These are simple, atomic tasks. Can the robot follow a specific instruction? Can it find a file? Can it write code that doesn't crash? This is like testing if the robot can tie its own shoes.
  • Level 2 (L2): The Combo Moves. Here, the robot has to chain several skills together. It might need to read a spreadsheet, use a search tool, and then write a summary. This is like asking the robot to make a sandwich: get the bread, get the meat, put it together, and wrap it up.
  • Level 3 (L3): The Boss Battle. This is the real-world stuff. The robot has to work within strict business rules, like handling healthcare data or financial reports where making a mistake could be dangerous. It has to follow laws, keep secrets, and produce a report that a human could actually trust.

The "Strict Pass" Rule: Why "Good Enough" Isn't Good Enough

Here is the most important part of the paper: The authors realized that just finishing a task isn't enough. An AI might finish a report, but if it made up a fake website link to support its claim, or if it accidentally deleted a file it wasn't supposed to touch, the report is useless.

To fix this, they created a 10D Risk Matrix. Think of this as a "safety inspector" that walks through the robot's work after it's done. The inspector looks for 10 specific types of trouble, such as:

  • D1: Making up facts or sources (Hallucination).
  • D2: Ignoring formatting rules (like writing in the wrong font).
  • D3: Breaking safety or privacy rules.
  • D4: Wasting time or using too many resources.
  • D8: Lying or hiding mistakes to look good.

The scoring system works like this:

  1. Completion: Did the robot finish the job? (Yes/No).
  2. Risk Penalty: Did the robot break any rules while doing it? If it made up a fact, it gets a penalty. If it broke a safety rule, it gets a huge penalty.
  3. Strict Pass: To get a "Strict Pass," the robot must have Completion = 1 (it finished the job) AND Risk Penalty = 0 (it didn't break a single rule).

This is a very high bar. It's like saying, "You can't just submit the homework; you have to submit it on time, with the right font, no cheating, and no made-up facts."

What the Tests Actually Found

The authors ran 27 different AI models through this 220-task gauntlet. They didn't let the models try again and again to get lucky; each model got just one shot at each task. Here is what happened:

  • The Best Performer: The model named Kimi K3 did the best, achieving a Weighted Pass@1 of 60.5%. This means it successfully completed and passed the safety check on about 60% of the tasks.
  • The Big Gap: Even the best model failed nearly half the time. But the real story is in the details.
  • The "L3" Problem: When the tasks got to Level 3 (the complex business scenarios), the scores dropped dramatically. The average "Strict Pass" rate for Level 3 across all models was only 18.5%. Every single model performed worse on these complex, real-world tasks than on the simpler ones.
  • The "Almost There" Failures: The paper found that out of 2,348 tasks where the models did finish the job (Completion = 1), 113 of them (4.8%) still failed the Strict Pass because they triggered a risk. For example, a model might have written a perfect report, but if it included a fake citation link, it failed. This proves that just getting the job done isn't enough; the quality and safety of the work matter just as much.

Why This Matters

The paper shows that we are currently very good at testing if AI can "think" (answer questions), but we are bad at testing if AI can "work" (deliver safe, usable results). The current generation of AI models is like a brilliant intern who is great at brainstorming ideas but keeps forgetting to save the files or accidentally sends emails to the wrong people.

The authors are careful to say that this test (SQBench v1.0) is just a snapshot. They tested 220 specific tasks, and the results might change if they tested different industries or tasks. They also note that they didn't have human experts check every single answer, so there is some uncertainty. However, the pattern is clear: Delivery under domain constraints is a shared weakness. Whether it's finance, healthcare, or manufacturing, AI agents are currently struggling to navigate the strict rules and safety requirements of real-world jobs.

The paper concludes that we need to stop celebrating "completion" and start celebrating "safe delivery." We need to measure AI not just by how many questions it can answer, but by how many times it can finish a job without causing a mess. Until we can do that, AI agents might be smart, but they aren't quite ready for the real world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →