From Performance to Purpose: A Sociotechnical Taxonomy for Evaluating Large Language Model Utility
This paper introduces the Language Model Utility Taxonomy (LUX), a comprehensive framework organized into performance, interaction, operations, and governance domains, to standardize the evaluation of large language model utility in real-world, high-stakes applications beyond mere task-level success.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are buying a car.
For years, the only way to judge a car was by looking at its horsepower. If the engine was big and fast, it was considered "good." But today, we know that a car with a massive engine is useless if the brakes don't work, the GPS is confusing, it costs a fortune to fill up with gas, or if it's illegal to drive in your city.
This is exactly the problem with Large Language Models (LLMs)—the AI chatbots like the one you are talking to right now.
For a long time, we only judged these AIs by how well they answered trivia questions or wrote code (their "horsepower"). But as companies start using them for real, serious jobs (like diagnosing patients, writing legal contracts, or managing customer service), we realized that speed and accuracy aren't enough. We need to know if the AI is safe, affordable, easy to use, and follows the rules.
This paper introduces a new "Car Manual" for AI called LUX (Language Model Utility Taxonomy). It's a checklist to help people decide if an AI is actually useful for their specific needs, not just if it's smart.
Here is how the LUX framework breaks down, using simple analogies:
1. Performance: The Engine
This is the old-school way of judging AI. It asks: Does the car run?
- Task Validity: Is the answer actually true? (No hallucinations/fake facts). Is it complete? (Did it answer the whole question?). Is it timely? (Is the info up to date?).
- Stability: If you ask the same question twice, do you get the same answer? If you ask a slightly different way, does it still make sense? And if the AI isn't sure, does it admit it, or does it confidently lie?
2. Interaction: The Dashboard and the Driver
This asks: Is the car easy to drive, and does it talk to you clearly?
- Workflow: How well does the AI connect to your other tools? Can it talk to your database, your email, or your calendar? Or is it an island that can't do anything on its own?
- Presentation: How does the AI speak? Is it clear, easy to read, and in the right format? If you ask for a summary, does it give you a paragraph or a novel?
- Traceability: Can the AI show its homework? If it makes a claim, can it point to the source? If you ask "Why did you say that?", can it explain its thinking process?
3. Operations: The Gas Tank and the Mechanics
This asks: Can you afford to keep this car running?
- Cost: It's not just about the price of the car; it's the gas, the insurance, and the mechanic bills. For AI, this means: How much does it cost per question? Do you need expensive super-computers to run it? Is it too expensive for your budget?
- Efficiency: How fast is it? If 1,000 people ask questions at the exact same time, does the car stall, or does it keep cruising smoothly?
4. Governance: The Traffic Laws and Safety Features
This asks: Is this car legal and safe to drive in your neighborhood?
- Policy Enforcement: Does the AI follow the rules? If a company says "Never mention our competitor," does the AI obey? Can it be tricked (jailbroken) into saying something it shouldn't?
- Security: Is the car locked? Does it protect your private data? If you tell the AI a secret, does it keep it safe, or does it leak it to the internet?
The Big Takeaway
The authors built a dynamic website (a digital toolbox) that connects every part of this checklist to real-world tests.
The main message is simple:
Don't just pick the "smartest" AI. Pick the AI that fits your purpose.
- If you are writing a novel, you might care most about Presentation and Creativity.
- If you are running a hospital, you care most about Governance (safety) and Performance (accuracy).
- If you are a startup with no money, Operations (cost) is your #1 priority.
The LUX framework helps you stop looking at just the engine and start looking at the whole car, ensuring the AI you choose actually gets you where you want to go.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.