Fiat Elpis · AI economics note 11
The AI token price floor is moving from cost to value
A higher list price does not necessarily mean AI became more expensive. Agents can use more tokens, complete larger tasks and justify higher spend while cost per outcome continues to fall.
From 300 recent @FiatElpis posts · expanded with primary-source research
The first phase of model competition compressed the price of raw intelligence. The next phase is less about the cheapest million tokens and more about which systems can turn tokens into completed, governed enterprise work.
01 · The information delta
Usage is deepening faster than headline prices imply
OpenAI said in April that enterprise represented more than 40% of revenue, Codex had reached 3 million weekly users and its APIs processed more than 15 billion tokens per minute. By August, its enterprise research showed frontier firms generating 8.3 times as many output tokens per active user as typical firms, up from 2.6 times in January.
The mix is also shifting from chat to agents. As of June, Codex generated 64% of combined Codex and ChatGPT output tokens among enterprise customers. Agentic tasks consume more output because they search, use tools, revise work and operate over longer horizons.
Meanwhile, model economics continue to improve. OpenAI says price per million tokens fell 97% from GPT-4 to GPT-5.4. DeepSeek’s current V4 pricing remains extremely low and uses cache discounts, while its documentation reserves the right to adjust prices. The commodity input is still deflationary; the workload is expanding.
“We may have found the token bottom—but the unit that matters is becoming the completed task.”
The original Fiat Elpis market note on X
02 · Signals to track
Why spend can rise while unit economics improve
Agents consume a larger work package
A long-running coding or research task can use far more tokens than a chat response while replacing hours of labor rather than seconds of search.
Repeated context becomes cheaper
Prompt caching lowers the marginal cost of large recurring contexts, encouraging applications to keep more company knowledge in the loop.
Enterprises pay for governance
Usage analytics, permissions, auditability and spend controls become part of the product, making list-token comparisons less complete.
Capacity can still command a premium
At the frontier, latency, concurrency and access to the best model can be scarce even when average inference cost falls.
The new unit of account
Tokens are becoming like cloud compute: a metered input whose price matters less than the revenue or labor attached to the workload. The winner may not sell the cheapest token. It may deliver the lowest cost per reliable, completed outcome.03 · What may be mispriced
A price floor would be an industry signal, not an end to deflation
Model providers can raise selected prices because capability, context, throughput or demand improved while still cutting the cost of an equivalent outcome. A blanket “token inflation” thesis therefore needs careful normalization by model quality and task success.
The more durable signal is that customers are increasing total spend as tools become useful enough to integrate. If usage, concurrency and paid enterprise controls expand faster than inference efficiency, industry revenue can grow even as normalized compute costs fall.
- Compare price per completed benchmark or workflow, not raw tokens alone.
- Track enterprise output tokens per active user and adoption outside engineering.
- Separate cached input, uncached input, reasoning and output pricing.
- Measure gross margin after inference, support, orchestration and reserved capacity.
04 · What would change my mind
What would restart the race to zero
Value-based pricing fails if capability becomes interchangeable faster than workloads deepen:
- Open models match frontier outcomes while remaining dramatically cheaper to operate.
- Enterprise usage plateaus after pilots and output per user stops growing.
- Agents fail reliability or governance tests in high-value workflows.
- Inference supply expands faster than demand, eliminating latency and capacity premiums.
- Customers route workloads automatically to the cheapest equivalent model.
Bottom line
The token is cheap; dependable work is not
The raw price of intelligence remains in a long-term decline. What is changing is the amount of intelligence customers can productively consume and the value of the tasks they are willing to delegate.
That is the credible version of a token bottom: not a permanent floor under every API rate, but a point where capability and demand let leading providers capture more value per customer even as cost per outcome keeps falling.
Sources & method
Primary sources, thesis separated from fact
- OpenAI — the next phase of enterprise AI
- OpenAI — enterprise agents from assistance to execution
- OpenAI — managing AI investments in the agentic era
- DeepSeek — current models and API pricing
- Anthropic — 2026 state of AI agents report
This note expands themes from the author’s recent X posts. Reported facts are linked to their sources; market interpretation is explicitly the author’s view. Market levels may change after publication.