Frontier AI model leaderboards with the kill line
Two boards averaged read as one that agrees. Nothing here is averaged — the disagreement is the reading.
N=118 · SOTA=21 · KILLED=87
Ranking · AA 17–53 · ARENA 1396–1510
- 153.4—
- 252.81480
- 350.71493
- 449.71506
- 548.2—
- 647.11483
- 744.91483
- 844.41456
- 943.81485
- 1042.31466
- 11Retired42.01481
- 1241.91475
- 1341.2—
- 14Retiredest40.71502
- 1540.31481
- 1639.9—
- 1739.81500
- 18est39.61490
- 19NewUnmatched39.5—
- 2039.11468
- 21Retiredest39.01476
- 22Retired38.61482
- 2338.41461
- 2437.51452
- 2536.31463
- 26NewUnmatchedExecutionerest35.5—
Kill line · index × price
Kill lineFrontier — nothing beats these on both axesSOTALow-costKilledHollow, or a faint cross — estimated by the source
Claude Fable 5.1
2sources
6entries
6configs
Where it places · 2 sources
—Artificial Analysis LLM Leaderboard53.4
1Arena — Agent0.139
Bars are drawn on each source's own fitted range, so position is not comparable between rows. The rank is the source's own, and a source that publishes none gets neither a number nor a bar.
Effort ladder · 6 configs
Artificial Analysis LLM Leaderboardmax wins
Arena — Agent · agentonly max ran
Six named efforts, same columns on every track; a greyed name is a configuration this source did not publish. Scores are each source's own and are not compared across tracks.