OnlySOTA

Coding

Writing and changing code, split by who does the judging: a grader running tests and a person picking the better of two answers measure different things, and often disagree.

No. 1 · AA
66.9%
Claude Opus 5.5 · $8.00
Executioner · price
$0.20
GPT 6 Luna
SOTA
15
Score above the executioner
Killed
39
Dearer, and no better
Ranking · AA 42–67% · A-CODE 1441–1831
  1. 166.9%1815
  2. 263.1%1749
  3. 3New61.8%1680
  4. 4Legacy61.0%1625
  5. 561.0%1786
  6. 660.9%1618
  7. 7Legacy59.8%1592
  8. 859.7%1657
  9. 959.5%1658
  10. 1059.0%1623
  11. 1158.9%1570
  12. 12Legacy58.8%1542
  13. 1358.7%1446
  14. 14Legacy57.8%1620
  15. 1557.8%1638
  16. 16Legacy57.6%1689
  17. 17Legacy57.4%1532
  18. 1856.6%1583
  19. 1956.5%1788
  20. 20Legacy56.5%1620
  21. 21Legacy56.4%1695
  22. 22Legacy56.1%1512
  23. 2355.8%1758
  24. 2455.0%1519
  25. 25Legacy55.0%1552
  26. 26Executioner54.6%1579
All 122 models on AA →
Kill line · score × price

Each dot is a model: further right costs more, higher up scores more. Where the dashed lines cross stands the executioner, the best value on the board. Everything below and to its right costs more and scores no higher, so it is killed. Above the line is SOTA, the top tier; below and to the left is cheaper and weaker, the low-cost picks. How the kill line is drawn →

SciCode · starts at 020%40%60%0$0.05$0.25$1.00$5.00$20.0066.9%$8.00Claude Opus 5.5ExecutionerGPT 6 Luna54.6% · $0.20SOTALOW-COSTKILLEDBlended Price, USD per 1M tokens (3:1 input:output) · log10
Kill lineFrontier — nothing beats these on both price and scoreSOTALow-costKilledHollow dot or faint cross — estimated by the source

Claude Opus 5.5

anthropic/claude-opus-5-5 · closed · $8.00 · SOTA

Open its full page →
Scores on each board
  • Artificial Analysis LLM Leaderboard57.6No. 1 of 372
  • Arena — Agent0.138No. 2 of 49
  • Artificial Analysis Coding Agents0.660No. 2 of 22
  • Arena — Code1815No. 1 of 109
  • Arena — Text1504No. 4 of 306

Each rank is the board’s own: the rank it publishes, or its place in its own order where it publishes none. The bar is how much of that board the model is ahead of. Scores from different boards cannot be compared.

Scores at each effort setting

Artificial Analysis LLM Leaderboardbest at max

  1. low42.3
  2. medium51.2
  3. high53.6
  4. xhigh56.0
  5. max57.6

Tested at one setting only

  • Arena — Agent · agenthigh0.138
  • Artificial Analysis Coding Agents · Claude Codemax0.660
  • Arena — Codemax1815
  • Arena — Texthigh1504

The same model at each effort setting. A bar runs from zero to that board’s best score, so read each board on its own and never across boards.

Where this board’s numbers come from

Each panel is one board’s ranking as published, never averaged. When two disagree, they are usually measuring different things.

Graded

Scored against tests or a reference, with no human in the loop and no agent loop either — these measure the model's output directly.

Web development

Human preference between two built web applications. Arena's code arena is web work end to end — every category it publishes is a web application category — so this is the board it actually is, rather than a general coding board with a narrower name.