The Price of Cognition
In July the price sheet became the most-read document in AI. A million tokens of context is table stakes; the real question is what a unit of thinking costs — and who can afford to run it all day.
Every frontier release in July shipped with the same second page: the meter. Claude Opus 5 prices at $5 per million input tokens and $25 per million output, with batch processing at half price [1]. Gemini 3.1 Pro undercuts at $2 per million in and $12 out for prompts up to 200K tokens [1]. The capability race gets the headlines; the pricing table is where the industry actually competes.
The line item nobody budgeted
For a business running agents continuously, these numbers stop being API trivia and become cost of goods sold. An agent that reads a full million-token context pays $5 before it does anything; every long answer bills at output rates. The engineering conversation flips from ‘can the model do it’ to ‘what is worth the tokens’ — and procurement starts asking why cognition is the one input nobody forecasts.
The pressure on that floor is structural. Within days of the closed-lab releases, DeepSeek-V4-Flash and Qwen 3.8 Max shipped as open-weights alternatives [2] — and every open release resets the ceiling on what closed labs can charge for the middle of the market. The premium narrows to the frontier edge: agentic reliability, desktop hands, verified benchmarks.