Model watch
A benchmark result is not a fixed property of a model. It is the outcome of one method, one version, one configuration and one day. Epoch corrected 42 percent of the problems in FrontierMath when version 2 was released on 12 June 2026, and ARC Prize state themselves that results on their community leaderboard are self reported and unverified. That is why this page carries positions and no scores, and why the numbers link to whoever measured them.
Writing code
Checked 23 August 2026Which model should write and fix code?
- LeadsClaude Opus 5AnthropicThe highest result on SWE-bench Verified. On LMArena's coding leaderboard, however, several older Opus versions sit ahead of it, which says something about the difference between fixing a bug and being pleasant to work with.
- CloseKimi K3 MaxMoonshotThe highest placed model outside Anthropic in human preferences on coding.
The two leaderboards disagree, and that is the point. SWE-bench measures whether the bug got fixed, LMArena measures what people preferred to work with. The runner up on SWE-bench, Claude Mythos 5, is also not generally available: only selected organisations defending critical infrastructure are allowed to run it.
SWE-bench
Measures. Whether the model can fix a real bug in a real repository, measured as the share of GitHub issues where the patch it produced makes the project's own tests pass.
Does not measure. Whether the code is understandable, safe, or something a colleague would have approved. A passing test is not the same thing as a good solution.
Worth knowing. The same leaderboard can mix different agents and scaffolding around the models, so a high number may just as well be a better tool as a better model. There is a view that runs everyone with the same simple agent, and that is the one that compares models.
Current results at SWE-benchReasoning
Checked 23 August 2026Which model handles hard multi step problems?
- LeadsGPT-5.6 SolOpenAIThe highest result both on ARC-AGI-2 and on FrontierMath tier 4.
- CloseClaude Opus 5AnthropicJust behind on ARC-AGI-2.
Read from summaries rather than directly from ARC Prize and Epoch, because their own leaderboards could not be read by machine. The numbers should therefore be checked against the source before being quoted onward.
ARC-AGI
Measures. Whether the model can solve pattern problems it has never seen, meaning work out the rule itself rather than recognise it.
Does not measure. Anything resembling ordinary work. The tasks are deliberately constructed and a high score says little about usefulness in an everyday task.
Worth knowing. Results on the semi private sets are run and verified by ARC Prize. The community leaderboard is self reported and not verified, which they state themselves.
Current results at ARC-AGIFrontierMath
Measures. Mathematics at research level, problems that take a specialist hours and cannot be solved by recognising a standard move.
Does not measure. Arithmetic and the mathematics anyone actually uses at work. This is the ceiling, not the floor.
Worth knowing. Epoch released version 2 on 12 June 2026 with corrections to 42 percent of the problems. Do not compare results across the versions.
Current results at FrontierMathWriting
Checked 23 August 2026Which model writes text a person wants to read?
- LeadsClaude Fable 5AnthropicFirst both overall and on creative writing. It was cut off for eighteen days this summer after a United States export order, and has been available again since 1 July.
- CloseGemini 3.7 FlashGoogleThe highest placement outside Anthropic on creative writing.
Read on LMArena, both the overall leaderboard and the creative writing category, which point the same way here. They do not always, and the overall leaderboard is the wrong one to read when the question is specific.
LMArena
Measures. What people actually prefer. Two answers are put against each other without the sender being visible, and the ranking is built from the votes.
Does not measure. Whether the answer is true. A preference vote measures which answer appealed, not which one was correct, and a well phrased error sometimes wins.
Worth knowing. The ranking differs sharply between categories, so anyone reading the overall leaderboard is reading the wrong question. The data is published under CC-BY and can be inspected.
Current results at LMArena