OnlySOTA

Families, each fielding its best model

Each line, such as Claude Opus, GPT or Gemini Flash, is represented by its strongest model, row intact, never averaged. Open one to see every model in it.

No. 1 · AA
57.6
Claude Opus 5.5 · $8.00
Executioner · price
$0.17
MiMo V2.6 Flash
SOTA
19
Score above the executioner
Killed
61
Dearer, and no better
Ranking · AA 18–58 · ARENA 1415–1534
  1. 1scores from Claude Opus 5.557.61504
  2. 2scores from GPT 6 Astra52.71477
  3. 3scores from Gemini 4 ArgonNew52.61525
  4. 4scores from Muse Spark 1.348.11494
  5. 5scores from Grok 4.746.41442
  6. 646.31480
  7. 7scores from Qwen 3.8 Max45.41482
  8. 8scores from GLM 5.344.81478
  9. 943.71454
  10. 1043.61488
  11. 1139.51474
  12. 12Executioner37.91452
All 372 models on AA →
Kill line · score × price

Each dot is a model: further right costs more, higher up scores more. Where the dashed lines cross stands the executioner, the best value on the board. Everything below and to its right costs more and scores no higher, so it is killed. Above the line is SOTA, the top tier; below and to the left is cheaper and weaker, the low-cost picks. How the kill line is drawn →

Artificial Analysis Intelligence Index · starts at 02040600$0.05$0.25$1.00$5.00$20.0057.6$8.00Claude Opus 5.5ExecutionerMiMo V2.6 Flash37.9 · $0.17SOTALOW-COSTKILLEDBlended Price, USD per 1M tokens (3:1 input:output) · log10
Kill lineFrontier — nothing beats these on both price and scoreSOTALow-costKilledHollow dot or faint cross — estimated by the source

Claude Opus 5.5

anthropic/claude-opus-5-5 · closed · $8.00 · SOTA

Open its full page →
Scores on each board
  • Artificial Analysis LLM Leaderboard57.6No. 1 of 372
  • Arena — Agent0.138No. 2 of 49
  • Artificial Analysis Coding Agents0.660No. 2 of 22
  • Arena — Code1815No. 1 of 109
  • Arena — Text1504No. 4 of 306

Each rank is the board’s own: the rank it publishes, or its place in its own order where it publishes none. The bar is how much of that board the model is ahead of. Scores from different boards cannot be compared.

Scores at each effort setting

Artificial Analysis LLM Leaderboardbest at max

  1. low42.3
  2. medium51.2
  3. high53.6
  4. xhigh56.0
  5. max57.6

Tested at one setting only

  • Arena — Agent · agenthigh0.138
  • Artificial Analysis Coding Agents · Claude Codemax0.660
  • Arena — Codemax1815
  • Arena — Texthigh1504

The same model at each effort setting. A bar runs from zero to that board’s best score, so read each board on its own and never across boards.