Logbook: the bench I am setting up for GPT-6 Sol against Opus 5.5
2026·09·27 · 3 min read

Foto: cottonbro studio · Pexels
Esta entrada todavía no está traducida a este idioma — se muestra la versión original.
Two frontier models launched on the same day with price cuts. Before believing either set of numbers, I want one real refactor from this repository run through both. This entry is the setup and the rules — no results yet.
Status: setup, not results. This entry describes an experiment that has not been run yet. When it is, the numbers go in this same entry and the status line disappears. There are no measurements below.
Why bother
On September 22, 2026 both Anthropic and OpenAI shipped a model whose headline was cost. Anthropic claims Opus 5.5 costs 40% less to run than Opus 5; OpenAI claims a flat 50% cut against GPT-5.6 promotional pricing.
Both numbers are true and both are useless to me. They are measured on the vendor's tasks. What I want to know is the cost of one task I actually have.
The task
A single real refactor from this repository, not a toy. The candidate is the
one that has been sitting in the backlog: three admin list pages that still
render their own inline <table> instead of using the generic DataTable<T>
that already exists. It qualifies because it has properties a synthetic prompt
does not:
- A correct answer exists and is mechanically checkable — the build, the type checker, the linter and the existing test suite all have an opinion.
- It requires reading code the model did not write, and matching a convention that is documented in one place and demonstrated in another.
- It is boring. Nobody will be impressed by the result, which is exactly why it is a fair test.
What gets measured
| Measure | How |
|---|---|
| Wall clock | From first prompt to a diff that builds |
| Real cost | The provider's own usage reporting for the session, not an estimate |
| Turns | Number of round trips, and tool calls within them |
| Quality | Review of the diff against the convention, by hand |
| Damage | Anything it changed that it was not asked to change |
The rules, written before the run
Both models get the same prompt, the same starting commit, and the same acceptance criteria. No hints after the first message. If either one asks a clarifying question, it gets the same answer, and the fact that it asked is recorded — after Muse Spark 1.3 made "asks when ambiguous" a headline feature, that behaviour deserves a column of its own.
If a run ends with a diff that does not build, it is a failure, not a do-over. Retrying until the model gets it right measures my patience, not the model.
What would make the result meaningless
Worth writing down before I fall for it: one task is one data point. A single refactor on a single repository in a single language says nothing about either model in general, and I will not be drawing a chart from n=1. The only claim this experiment can support is what happened to this task on this day, at these prices.