TaxPraBen: A Scalable Benchmark for Structured Evaluation of LLMs in Chinese Real-World Tax Practice
The paper introduces TaxPraBen, the first scalable benchmark designed to evaluate Large Language Models on 13 Chinese tax practice tasks—including real-world scenarios like risk prevention and strategy planning—revealing significant performance disparities among models and highlighting the need for specialized evaluation in legally regulated domains.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a very smart, well-read robot how to be a Tax Accountant.
You might think, "No problem! This robot has read the entire internet. It knows everything about language, math, and logic." But here is the catch: Taxes are not just about reading or math; they are about navigating a minefield of rules, numbers, and real-world consequences.
This paper introduces TaxPraBen, a new "driving test" specifically designed to see if AI robots are actually ready to drive a tax car, or if they are just good at reciting the driver's manual.
Here is the breakdown of what they did, using some everyday analogies:
1. The Problem: The "Textbook" vs. The "Street"
Previously, researchers tested AI on tax questions that were like multiple-choice quizzes.
- The Old Way: "What is the tax rate for a car?" (Answer: 10%).
- The Reality: A real tax expert doesn't just know the rate. They have to look at a messy business scenario, calculate the exact tax owed, figure out if the business is breaking the law, and suggest a strategy to save money legally.
The authors realized that current AI models are like students who ace the written test but crash the car on the first day of driving. They can recite laws (Knowledge) but fail when asked to apply them to a messy, real-world situation (Application).
2. The Solution: The "TaxPraBen" Driving School
The team built a new benchmark called TaxPraBen. Think of it as a simulator that puts the AI through three levels of difficulty, based on how humans learn (Bloom's Taxonomy):
Level 1: The Flashcards (Memorization)
- The Task: "Recite the exact text of Article 111 of the Tax Law."
- The Analogy: Can the robot read the dictionary and repeat the definition of "deduction" perfectly?
- Result: Most big AI models are good at this. They have good memories.
Level 2: The Reading Comprehension (Understanding)
- The Task: "Read this 5-page news article about a new tax rule and summarize the main point."
- The Analogy: Can the robot understand the story behind the rule, not just the words?
- Result: Chinese-specific AI models (like Qwen) did better here than generic English models, likely because they understand the local "dialect" of tax laws better.
Level 3: The Real-World Crisis (Application)
- The Task: "Here is a company's messy financial data. Calculate their tax, find the risks, and tell them how to save money without getting arrested."
- The Analogy: This is the final exam. The robot must do math and write a legal strategy at the same time.
- Result: This is where everyone struggled. Even the smartest AIs (like GPT-4o) got the math wrong or gave dangerous advice. They are great at talking, but terrible at doing the actual calculation.
3. The "Magic Filter" (Structured Evaluation)
One of the biggest headaches in testing AI is that they are messy. You ask for a number, and they give you a paragraph of text with the number hidden inside.
The authors created a "Magic Filter" (a structured evaluation pipeline).
- The Analogy: Imagine you ask a chef for a recipe. Instead of getting a messy paragraph, the AI is forced to fill out a strict form (JSON).
- If the AI says, "The tax is 500 dollars," the filter extracts just "500" and checks if it's right. If the AI says, "The tax is probably around 500, maybe a bit more," the filter marks it as wrong.
- This ensures the test is fair and can be graded automatically, just like a computer grading a math test.
4. The Big Surprises
When they ran the test on 19 different AI models, they found some interesting things:
- Big is Better (but not always): The massive, closed-source models (like ERNIE-3.5 and GPT-4) generally did the best. They have huge "brains" with lots of knowledge.
- Local Knowledge Wins: AI models trained specifically on Chinese data (like Qwen2.5) beat the global models on Chinese tax tasks. It's like a local guide knowing the shortcuts better than a tourist with a map.
- Special Training Doesn't Guarantee Success: One model, YaYi2, was specifically trained on tax data. You'd think it would be the champion, but it actually performed poorly.
- Why? It's like a student who memorized the old textbook but the test was on new laws. The data they trained on wasn't diverse enough or didn't match the real-world test questions.
- The Math Gap: The hardest part for all AI was the math. They can write a beautiful essay about tax planning, but if you ask them to calculate the exact dollar amount, they often fail.
5. Why This Matters
This paper is a wake-up call. It tells us that we cannot trust AI with our taxes yet.
If you ask an AI to write a poem, it's amazing. If you ask it to calculate your tax return, it might give you a number that looks right but is actually wrong, leading to fines or audits.
TaxPraBen is the tool we need to stop the "hype" and start the "reality check." It forces AI developers to stop building models that just talk about taxes and start building models that can actually do the work.
In short: We have built a very smart robot that can read the tax code. But until it passes the TaxPraBen driving test, we shouldn't let it drive the car.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.