← Latest papers
🤖 AI

Artificial Effort

This paper demonstrates that most canonical real-effort tasks used in experimental economics can now be solved accurately and cheaply by various Large Language Models, thereby undermining the validity of these tasks in unsupervised settings where participants might outsource work to AI.

Original authors: Federico Belotti, Stefano Coniglio, Antonio Cosma, Francesco Fallucchi

Published 2026-05-26
📖 4 min read☕ Coffee break read

Original authors: Federico Belotti, Stefano Coniglio, Antonio Cosma, Francesco Fallucchi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a teacher giving a class a series of puzzles to solve. You want to know how hard the students are working, so you give them tasks that require real mental effort, like counting specific items in a messy picture or solving a math problem. This is called a "real-effort task." The whole system relies on one big assumption: a human is actually doing the work.

This paper asks a scary question for researchers: What if the students aren't humans at all? What if they are AI robots?

Here is the story of what the researchers found, explained simply:

1. The "Robot Student" Test

The researchers took 8 classic puzzles used in economics experiments (like adding numbers, counting zeros in a grid, or solving a mini Sudoku) and handed them to 23 different AI models (the "brains" behind tools like ChatGPT, Gemini, and Claude). They treated the AI like a student taking a test.

The Result: The AI didn't just pass; it aced most of the easy tests.

  • The Easy Stuff: For tasks like adding three numbers or decoding a simple letter code, the AI got nearly 100% correct. It was like giving a calculator a math test; it just worked perfectly.
  • The Hard Stuff: The AI struggled with tasks that require "counting" many small things (like counting zeros in a huge grid) or spotting patterns in messy, distorted images. It's like the AI is great at logic but sometimes gets confused when it has to count every single grain of sand on a beach.

2. The "Newer is Better" Rule

The researchers tested older AI models and newer ones.

  • The Trend: Every time a company released a newer version of their AI, it got smarter.
  • The Surprise: The "mid-range" models (the cheaper, more common ones) were catching up to the super-expensive, top-tier models very fast. It's like a budget car brand suddenly making cars that drive almost as well as a luxury sports car. This means you don't need the most expensive AI to cheat on these tests; a cheaper one works just fine.

3. The "Pocket Change" Problem

Here is the most important part for the researchers. In these experiments, humans get paid a small amount of money (a "piece rate") for every puzzle they solve.

  • The Cost: The researchers calculated how much it costs to pay an AI to solve these puzzles.
  • The Math: It costs a fraction of a penny to let an AI do the work. Even the most expensive AI models cost less than 25 cents to solve a whole set of puzzles.
  • The Conclusion: If a person trying to make money online could hire an AI to do the work, they would make a huge profit. The AI is not only faster and more accurate, but it is also cheaper than a human. It's like hiring a robot to mow your lawn for $0.05 instead of paying a neighbor $20.

4. The "Bribe" That Didn't Work

In human experiments, researchers often say, "If you get this right, you get a bonus!" Humans usually work harder when they hear about the bonus.

  • The Test: The researchers told the AI, "You are a human participant," and "You will get $0.50 for every correct answer."
  • The Result: The AI didn't care. Its performance didn't change at all.
  • Why? An AI doesn't have a "mood" or a desire for money. It just processes the request. Telling it to "try harder" or "be human" is like telling a calculator to "add faster" by promising it a cookie. It just doesn't work.

The Big Takeaway

The paper concludes that the old way of testing human effort is broken in the age of AI.

If you run an online experiment where people can use their phones or computers without supervision, you can no longer be sure that the person solving the puzzle is actually a human. They might just be using a cheap AI to do the work for them.

The authors suggest three ways to spot the "Robot Students":

  1. Pick harder puzzles: Use tasks that require visual counting or messy image recognition, which AIs are still bad at.
  2. Watch the motivation: If a participant's performance doesn't change when you offer a bonus, they might be a robot (since humans usually care about bonuses).
  3. Watch the fatigue: Humans get tired or learn as they go; robots stay exactly the same performance level from the first question to the last.

In short: The "Real Effort" in these experiments might now be "Artificial Effort," and it costs almost nothing.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →