← Latest papers
📈 economics

Token Is All You Price

The paper demonstrates that token-based pricing in GenAI services is the revenue-optimal way to screen buyers with different levels of urgency by using a menu of stopping-time caps on information throughput.

Original authors: Weijie Zhong

Published 2026-04-28
📖 3 min read☕ Coffee break read

Original authors: Weijie Zhong

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are at a high-end, all-you-can-eat buffet, but there is a catch: the chef is incredibly talented, but they can only cook so many dishes per hour because the kitchen is small.

Now, imagine there are two types of customers:

  1. The "Hungry Professional": They are in a massive rush. They need to eat now to get back to a business meeting. They are willing to pay a premium to get food on the table immediately.
  2. The "Leisurely Diner": They have all afternoon. They don't mind waiting for the perfect dish to be prepared, as long as the food is good.

The Problem: How does the restaurant owner make the most money without making the "Hungry Professional" wait too long or making the "Leisurely Diner" feel ripped off?

The Old Way of Thinking (The "Menu of Different Foods" Approach)

In traditional economics, a business usually tries to separate these customers by offering different products. They might give the hungry person a quick, mediocre sandwich and the patient person a slow-cooked, gourmet steak. This is called "distorting the product." It’s not ideal because the hungry person doesn't actually want a mediocre sandwich—they want the gourmet steak, just faster.

The Paper’s Big Discovery (The "Timer" Approach)

This paper, written by Weijie Zhong, argues that for modern AI services (like ChatGPT or Claude), the smartest way to make money isn't to change what the AI says, but to change how long you are allowed to talk to it.

Instead of offering a "dumb" AI for fast people and a "smart" AI for slow people, the paper says the optimal strategy is to:

  1. Use one "Super-Chef" (The Greedy Process): Use the best possible AI model that works as fast and efficiently as possible to find the answer. This is what computer scientists call "preference-aligned"—it’s the version of the AI that is most helpful to a human.
  2. Use a "Stopwatch" (The Token Cap): Instead of changing the AI, you just give people different "time limits" or "token quotas."

The hungry professional pays a high price for a "Short Burst" pass (e.g., "You can use the Super-Chef for 5 minutes"). The patient diner pays a lower price for a "Long Session" pass (e.g., "You can use the Super-Chef for 5 hours").

Why This Matters for AI (The "Token" Connection)

If you’ve ever used an AI and seen messages like "You have 40 messages left for the next 3 hours," or if you've seen B2B companies paying for different "API Tiers" (Priority vs. Batch), you are seeing this exact math in action.

The paper explains that:

  • The AI doesn't need to be "dumbed down" for cheaper tiers. It stays high-quality.
  • The "Price" is actually a "Speed Limit." You aren't just paying for information; you are paying for the right to keep the information flowing before the system cuts you off.

The "Magic" of the Math

The most surprising part of the paper is that the "best" way to serve everyone is actually the "best" way to serve a single person who has no time limit at all.

In most businesses, the owner has to "cheat" the customer a little bit to make a profit (by giving them lower quality). But in the world of AI information, the owner can maximize their profit by being perfectly honest with the quality and only being "strict" with the clock.

In short: The paper proves that in the age of AI, you don't sell "intelligence"—you sell "time to think."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →