Coding
Writing and changing code, split by who does the judging: a grader running tests and a person picking the better of two answers measure different things, and often disagree.
- No. 1 · AA
- 66.9%
- Claude Opus 5.5 · $8.00
- Executioner · price
- $0.20
- GPT 6 Luna
- SOTA
- 15
- Score above the executioner
- Killed
- 39
- Dearer, and no better
- 166.9%1815
- 263.1%1749
- 3New61.8%1680
- 4Legacy61.0%1625
- 561.0%1786
- 660.9%1618
- 7Legacy59.8%1592
- 859.7%1657
- 959.5%1658
- 1059.0%1623
- 1158.9%1570
- 12Legacy58.8%1542
- 1358.7%1446
- 14Legacy57.8%1620
- 1557.8%1638
- 16Legacy57.6%1689
- 17Legacy57.4%1532
- 1856.6%1583
- 1956.5%1788
- 20Legacy56.5%1620
- 21Legacy56.4%1695
- 22Legacy56.1%1512
- 2355.8%1758
- 2455.0%1519
- 25Legacy55.0%1552
- 26Executioner54.6%1579
Each dot is a model: further right costs more, higher up scores more. Where the dashed lines cross stands the executioner, the best value on the board. Everything below and to its right costs more and scores no higher, so it is killed. Above the line is SOTA, the top tier; below and to the left is cheaper and weaker, the low-cost picks. How the kill line is drawn →
Claude Opus 5.5
- Artificial Analysis LLM Leaderboard57.6No. 1 of 372
- Arena — Agent0.138No. 2 of 49
- Artificial Analysis Coding Agents0.660No. 2 of 22
- Arena — Code1815No. 1 of 109
- Arena — Text1504No. 4 of 306
Each rank is the board’s own: the rank it publishes, or its place in its own order where it publishes none. The bar is how much of that board the model is ahead of. Scores from different boards cannot be compared.
Artificial Analysis LLM Leaderboardbest at max
- low42.3
- medium51.2
- high53.6
- xhigh56.0
- max57.6
Tested at one setting only
- Arena — Agent · agenthigh0.138
- Artificial Analysis Coding Agents · Claude Codemax0.660
- Arena — Codemax1815
- Arena — Texthigh1504
The same model at each effort setting. A bar runs from zero to that board’s best score, so read each board on its own and never across boards.
Where this board’s numbers come from
Each panel is one board’s ranking as published, never averaged. When two disagree, they are usually measuring different things.
Graded
Scored against tests or a reference, with no human in the loop and no agent loop either — these measure the model's output directly.
Web development
Human preference between two built web applications. Arena's code arena is web work end to end — every category it publishes is a web application category — so this is the board it actually is, rather than a general coding board with a narrower name.