
The most expensive model can make for the cheapest job
Translated by AI from the Swedish original · Read the Swedish original
A new AI model can cost more per token and still come out cheaper, if it does the job in fewer steps. A job that took the previous model over 20 hours took Opus 5.5 under three.
Anthropic's new model Claude Opus 5.5 was given a codebase of 200,000 lines to audit and fix. According to one tester, it finished in under three hours. The previous model, Opus 5, took more than 20 hours and used 2.5 times as many tokens. The price per token came down as well. But the big difference was that the new model needed less work to finish.
For an agent, an AI that works on its own through a task in many steps, the price of each token matters less if the new model:
- needs fewer steps to solve the same task. GitHub tested Opus 5.5 in its coding environment and says it solved more tasks than Opus 5 in less than half as many steps.
- gets stuck less often and needs fewer retries. At Quantium, a difficult coding task used to take 38 instructions over four days. With Opus 5.5, 11 instructions over three hours were enough.
- can reuse earlier context more cheaply. An agent rereads the same material again and again, and according to Anthropic that part accounts for most of the cost of agent work. Anthropic cut that price by 60 percent. OpenAI gives a 90 percent discount on text Sol has already read.
- can keep working independently for many hours. A developer at Clio let Opus 5.5 work unattended overnight, and the model stayed on the task for more than 18 hours.
- and more often actually completes the task. Anthropic asked Opus 5.5 to cut load times on every page of a web app. According to Anthropic, it succeeded 39 times out of 40, while Opus 5 made only smaller improvements.
Opus 5.5 and GPT-6 Sol were released on the same day, 22 September. Both companies cut the price per token. But when they wanted to show what the models are worth, they pointed to what a finished job cost. OpenAI itself writes that the cost of letting agents work for a long time matters more as tasks get longer.
Anyone who only compares price lists is comparing the wrong number. When AI starts carrying out longer tasks, the interesting question is not what a token costs, but what it cost to get the job done.
Ask upplyst.ai
Why does it matter?
Anyone comparing AI models usually looks at the price per token. For an agent that works for hours, the number of steps and retries decides what the job costs, and so does how often it actually finishes. A model with a higher price per token can therefore be cheaper per finished job.
What is the background?
Claude Opus 5.5 is Anthropic's new model and GPT-6 Sol is OpenAI's. Both were released on 22 September. An agent is an AI that works on its own through a task in many steps, such as auditing and fixing code. Context is the material the agent has already read. When the agent rereads it, the companies charge a separate, lower price.
What is uncertain?
The examples from GitHub, Quantium and Clio are the companies' own accounts of their own tests. The code audit comes from a tester who is not named. The web app test is Anthropic's own. The sources do not say what the runs cost in money, only how many tokens and hours they took. The article shows no case where a model with a higher price per token was cheaper per job, only why it can happen.