Fiat Elpis · AI economics note 11

The AI token price floor is moving from cost to value

A higher list price does not necessarily mean AI became more expensive. Agents can use more tokens, complete larger tasks and justify higher spend while cost per outcome continues to fall.

21 August 2026 8 min read

From 300 recent @FiatElpis posts · expanded with primary-source research

The first phase of model competition compressed the price of raw intelligence. The next phase is less about the cheapest million tokens and more about which systems can turn tokens into completed, governed enterprise work.

01 · The information delta

Usage is deepening faster than headline prices imply

OpenAI said in April that enterprise represented more than 40% of revenue, Codex had reached 3 million weekly users and its APIs processed more than 15 billion tokens per minute. By August, its enterprise research showed frontier firms generating 8.3 times as many output tokens per active user as typical firms, up from 2.6 times in January.

The mix is also shifting from chat to agents. As of June, Codex generated 64% of combined Codex and ChatGPT output tokens among enterprise customers. Agentic tasks consume more output because they search, use tools, revise work and operate over longer horizons.

Meanwhile, model economics continue to improve. OpenAI says price per million tokens fell 97% from GPT-4 to GPT-5.4. DeepSeek’s current V4 pricing remains extremely low and uses cache discounts, while its documentation reserves the right to adjust prices. The commodity input is still deflationary; the workload is expanding.

“We may have found the token bottom—but the unit that matters is becoming the completed task.”

The original Fiat Elpis market note on X

02 · Signals to track

Why spend can rise while unit economics improve

Tasks

Agents consume a larger work package

A long-running coding or research task can use far more tokens than a chat response while replacing hours of labor rather than seconds of search.

Caching

Repeated context becomes cheaper

Prompt caching lowers the marginal cost of large recurring contexts, encouraging applications to keep more company knowledge in the loop.

Control

Enterprises pay for governance

Usage analytics, permissions, auditability and spend controls become part of the product, making list-token comparisons less complete.

Scarcity

Capacity can still command a premium

At the frontier, latency, concurrency and access to the best model can be scarce even when average inference cost falls.

The new unit of account

Tokens are becoming like cloud compute: a metered input whose price matters less than the revenue or labor attached to the workload. The winner may not sell the cheapest token. It may deliver the lowest cost per reliable, completed outcome.

03 · What may be mispriced

A price floor would be an industry signal, not an end to deflation

Model providers can raise selected prices because capability, context, throughput or demand improved while still cutting the cost of an equivalent outcome. A blanket “token inflation” thesis therefore needs careful normalization by model quality and task success.

The more durable signal is that customers are increasing total spend as tools become useful enough to integrate. If usage, concurrency and paid enterprise controls expand faster than inference efficiency, industry revenue can grow even as normalized compute costs fall.

  1. Compare price per completed benchmark or workflow, not raw tokens alone.
  2. Track enterprise output tokens per active user and adoption outside engineering.
  3. Separate cached input, uncached input, reasoning and output pricing.
  4. Measure gross margin after inference, support, orchestration and reserved capacity.

04 · What would change my mind

What would restart the race to zero

Value-based pricing fails if capability becomes interchangeable faster than workloads deepen:

  • Open models match frontier outcomes while remaining dramatically cheaper to operate.
  • Enterprise usage plateaus after pilots and output per user stops growing.
  • Agents fail reliability or governance tests in high-value workflows.
  • Inference supply expands faster than demand, eliminating latency and capacity premiums.
  • Customers route workloads automatically to the cheapest equivalent model.

Bottom line

The token is cheap; dependable work is not

The raw price of intelligence remains in a long-term decline. What is changing is the amount of intelligence customers can productively consume and the value of the tasks they are willing to delegate.

That is the credible version of a token bottom: not a permanent floor under every API rate, but a point where capability and demand let leading providers capture more value per customer even as cost per outcome keeps falling.

Sources & method

Primary sources, thesis separated from fact

This note expands themes from the author’s recent X posts. Reported facts are linked to their sources; market interpretation is explicitly the author’s view. Market levels may change after publication.