Four labs, one month: what it means that frontier inference now costs half
2026·09·27 · 3 min read

Foto: Kevin Ache (@kevinache) · Unsplash
Esta entrada todavía no está traducida a este idioma — se muestra la versión original.
In September 2026 the interesting releases were not about capability, they were about the invoice. Anthropic, OpenAI, Sakana and Meta all shipped something cheaper in the same four weeks, and that changes which ideas are worth prototyping.
The month the price moved, not the benchmark
Frontier model releases used to be announced with a chart. September 2026 was announced with a price list. Four labs shipped in the same four weeks, and in every case the headline was cost per token, not a new capability.
What actually shipped
Anthropic — Claude Opus 5.5, September 22. Anthropic says it performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5. Input and output are $4 and $20 per million tokens, 20% below Opus 5, and cache reads — which the announcement points out are the majority of the cost in agentic and coding work — dropped to $0.20 per million, 60% less. It also generates output more than 30% faster. (Anthropic)
OpenAI — GPT-6 Sol and Luna, September 22. Two lighter members of the GPT-6 family, with API prices cut by 50% against the GPT-5.6 promotional pricing. Sol's output went from $20 to $10 per million tokens; Luna's from $1.20 to $0.50, with input at $0.10. OpenAI attributes the cut to caching and inference improvements it is passing through rather than to a smaller model. (OpenAI)
Sakana AI — Fugu Max, September 11. A different bet: instead of one cheaper model, an orchestration layer that routes each task to the leanest model in a pool of open-weight and specialized models. $2 per million input tokens and $6 per million output, which Sakana puts 40–60% below Sonnet 5, GPT-5.6 Terra and Kimi K3 on output. (Sakana AI)
Meta — Muse Spark 1.3, September 2. No price tag here, but the same axis from the other side: in comparisons by Meta engineers it used about 20% fewer tool calls and 25% fewer tokens than 1.2 for the same work. (Meta AI Research)
Cheaper per token is not the whole story
The part worth noticing is that three of the four releases talk about tokens spent on a task, not the price of a token. Anthropic's 40% figure is the net of a lower unit price and fewer tokens per task. Meta's number is purely efficiency. Sakana's entire product is deciding not to send a lookup to a multi-trillion-parameter model.
That distinction matters if you are the one paying. A 20% unit discount is a line item. A model that finishes the same refactor in half the turns changes what you are willing to let it attempt unattended.
What it changes for someone building on these APIs
For the kind of work this site runs on — a chat widget, a translation draft, a summarizer — cost stopped being the constraint some time this month. At $0.10 per million input tokens, the budget ceiling is no longer what decides whether a feature ships.
What decides it now is everything that was always the harder problem: whether the output is good enough to put in front of a person without review, whether the failure mode is safe, and whether you can tell afterwards what it did. None of those got cheaper in September.
The honest summary of the month: the excuse "the model is too expensive for that" expired. The excuse "I can't verify what it produced" did not.