Families, each fielding its best model
Each line, such as Claude Opus, GPT or Gemini Flash, is represented by its strongest model, row intact, never averaged. Open one to see every model in it.
- No. 1 · AA
- 57.6
- Claude Opus 5.5 · $8.00
- Executioner · price
- $0.17
- MiMo V2.6 Flash
- SOTA
- 19
- Score above the executioner
- Killed
- 61
- Dearer, and no better
- 1scores from Claude Opus 5.557.61504
- 2scores from GPT 6 Astra52.71477
- 3scores from Gemini 4 ArgonNew52.61525
- 4scores from Muse Spark 1.348.11494
- 5scores from Grok 4.746.41442
- 646.31480
- 7scores from Qwen 3.8 Max45.41482
- 8scores from GLM 5.344.81478
- 943.71454
- 1043.61488
- 1139.51474
- 12Executioner37.91452
Each dot is a model: further right costs more, higher up scores more. Where the dashed lines cross stands the executioner, the best value on the board. Everything below and to its right costs more and scores no higher, so it is killed. Above the line is SOTA, the top tier; below and to the left is cheaper and weaker, the low-cost picks. How the kill line is drawn →
Claude Opus 5.5
- Artificial Analysis LLM Leaderboard57.6No. 1 of 372
- Arena — Agent0.138No. 2 of 49
- Artificial Analysis Coding Agents0.660No. 2 of 22
- Arena — Code1815No. 1 of 109
- Arena — Text1504No. 4 of 306
Each rank is the board’s own: the rank it publishes, or its place in its own order where it publishes none. The bar is how much of that board the model is ahead of. Scores from different boards cannot be compared.
Artificial Analysis LLM Leaderboardbest at max
- low42.3
- medium51.2
- high53.6
- xhigh56.0
- max57.6
Tested at one setting only
- Arena — Agent · agenthigh0.138
- Artificial Analysis Coding Agents · Claude Codemax0.660
- Arena — Codemax1815
- Arena — Texthigh1504
The same model at each effort setting. A bar runs from zero to that board’s best score, so read each board on its own and never across boards.