Logbook: letting Fugu Max decide which model gets the job
2026·09·27 · 3 min read

Photo: Dmytro Demidko (@wildbook) · Unsplash
An orchestrator that routes each task to the cheapest model that can handle it is either a very good idea or a very expensive abstraction. Here is the test I set up to find out which, and what I refuse to conclude from it.
Status: setup, not results. No runs, no numbers yet. This entry exists to fix the method before the results can talk me into a conclusion.
The premise
Sakana AI's Fugu Max (September 11, 2026) is not a model, it is a router. One API call goes in; behind it, a pool of open-weight and specialized models, and a layer that decides which one is the leanest that can do the job. Priced at $2 per million input tokens and $6 per million output, which Sakana puts 40–60% below Sonnet 5, GPT-5.6 Terra and Kimi K3 on output.
The framing in the announcement is the part I agree with before testing anything: "A system that deploys a multi-trillion-parameter model to execute a simple data lookup is not intelligent, but wasteful."
What I want to know
Not whether it is cheaper on Sakana's benchmarks — it is, that is what the post is for. What I want to know is where the routing decision is wrong, because that is the number that decides whether it goes into a product.
A router has a specific failure mode: it sends a task to a model that is almost good enough. The output comes back plausible and subtly wrong, at a third of the price, and you find out two steps later.
The test set
Four task shapes drawn from things this site actually does, each with a known right answer:
| Shape | Example from here | Right answer known because |
|---|---|---|
| Trivial extraction | Pull the tags out of a post body | Deterministic, checkable |
| Short generation | Draft the excerpt of an entry | Compared against the hand-written one |
| Translation | The EN↔ES draft the CMS already produces | Existing output to diff against |
| Reasoning over context | Answer a question from the knowledge base | Fixed question set with graded answers |
Each shape runs twice: once through Fugu Max, once against the strongest model I would otherwise have called directly. Recorded per run: cost, latency, and whether the output is acceptable without edits.
The rule I am setting now, before seeing anything
Cheaper only counts when the output is acceptable. A run that costs a tenth and needs a human to fix it costs more than the expensive one. So the comparison is cost-per-accepted-output, not cost per call, and the acceptance call gets made against the criteria above rather than by how the text feels when I read it.
The interesting result would be a shape where delegating clearly wins and a shape where it clearly loses. If everything comes out roughly even, the honest conclusion is that the experiment was too small to say anything — and that is what will be written here.