← Latest papers
💻 computer science

Economic Evaluations of Language Models

This paper introduces EconEvals, a cost-effective, open-source evaluation suite and simulation framework that assesses language models' potential to save time across US occupations, revealing that while current models could significantly impact nearly half of all jobs, their actual adoption is currently limited by privacy concerns and proprietary system barriers rather than technical capability.

Original authors: Alexander Wan, Stephane Hatgis-Kessell, Tomás Aguirre, Percy Liang, Rishi Bommasani

Published 2026-07-23
📖 5 min read🧠 Deep dive

Original authors: Alexander Wan, Stephane Hatgis-Kessell, Tomás Aguirre, Percy Liang, Rishi Bommasani

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world of artificial intelligence as a massive, bustling library where the books are getting smarter every day. For a long time, the librarians (the scientists) have been testing these new, super-smart books by asking them to solve math puzzles or write code. It's like checking if a new car can win a race on a track. But here's the problem: real life isn't a race track. Real life is a messy, complicated city where people do thousands of different jobs, from fixing pipes to managing hospitals. The big question everyone is asking is: "If these AI books are so smart, how much will they actually help us in our daily jobs? Will they save us time, or will they just sit on the shelf?"

To answer this, we need to understand a few things. First, there are "benchmarks," which are like standardized tests for AI. If an AI gets a high score, it means it's good at that specific test. Second, there's "exposure," which is a fancy word for asking, "How much of a worker's day could this AI actually take over?" Finally, there's the gap between what AI can do in a test and what it actually does in the real world. We care about this because if we don't know the answer, we can't plan for the future. We might think AI is going to change everything tomorrow, or we might think it won't change anything at all, and both guesses could be wrong.

Enter a team of researchers from Stanford and the University of São Paulo who decided to stop guessing and start measuring. They built a new, open-source toolkit called ECONEVALS. Think of this toolkit as a giant, digital map of the entire US job market. Instead of just looking at a few popular jobs like software engineers, they mapped out over 1,000 different occupations, breaking them down into nearly 19,000 tiny tasks. It's like having a microscope that can zoom in on every single thing a worker does, from "answering the phone" to "designing a bridge."

The team faced a tricky problem: they needed to know what real people actually ask AI when they are working, but most of that data is private. So, they did two things. First, they dug through millions of public conversations to find real work-related questions. Second, when they couldn't find real data for a specific job, they created "synthetic" data. Imagine a robot actor role-playing as a nurse, a teacher, or a construction worker, asking an AI for help with their specific daily tasks. They used this mix of real and role-played data to test 10 different AI models.

What did they find? The results are a bit like a "potential vs. reality" story. The researchers discovered that current AI models are actually quite good at helping with work. In fact, their simulations suggest that for 47% of all US occupations, AI could save workers a significant amount of time on at least half of their daily tasks. That's a huge chunk of the workforce!

However, here is the twist: just because AI can do the job doesn't mean people are using it to do the job. The team found that for 79% of the tasks where AI could theoretically save a lot of time, people aren't actually using tools like Claude very much. It's like having a Ferrari in your garage that can drive 200 mph, but you're only driving it to the mailbox because you're afraid of the speed or don't know how to shift gears.

So, what's stopping us from using AI more? The researchers ran a special simulation where they broke down exactly why AI couldn't save time on certain tasks. They found that the biggest bottlenecks aren't that the AI is too dumb; it's that the real world is too complicated. The main reasons AI can't help yet include:

  • Physical actions: AI can't pick up a hammer or fix a leaky pipe.
  • Privacy: You can't let an AI read your private medical records or legal files.
  • Proprietary systems: Many companies use old, secret software that AI can't connect to.
  • Live interaction: AI struggles with real-time, face-to-face conversations where things change instantly.

The paper also compared their new, detailed method to older ways of measuring AI. The old methods were like looking at a list of superpowers and saying, "This AI can do everything!" The new method is more like actually watching the AI try to do the job and seeing where it trips. They found that the old methods were way too optimistic, predicting that AI could help with 68% of jobs, while their more realistic simulation suggests it's closer to 47%.

In short, this paper tells us that AI is a powerful tool that is ready to help with a lot of our work, but we aren't using it yet because of real-world hurdles like privacy rules and the need for physical hands. The researchers have built a map and a measuring stick that can be updated as AI gets smarter, helping us understand exactly where the technology will land next, rather than just guessing. They aren't saying AI has solved the economy, but they have given us a much clearer picture of where the roadblocks are and where the open lanes might be.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →