Agent
Not one answer but a task seen through, with tools. Two questions live here, and they are not the same one: which model to send in, and which harness to drive it with.
- No. 1 · AA
- 63.6%
- Claude Sonnet 5.5 · $4.00
- Executioner · price
- $0.17
- MiMo V2.6 Flash
- SOTA
- 16
- Score above the executioner
- Killed
- 39
- Dearer, and no better
- 163.6%0.6840.125
- 259.6%0.6600.138
- 359.6%0.6160.123
- 4New57.1%0.6380.076
- 556.1%0.6290.112
- 655.1%0.6220.143
- 7Legacy49.0%0.5970.087
- 8Legacy43.9%0.5670.097
- 9Legacy42.4%—0.082
- 1041.9%0.5360.024
- 11Legacy39.9%0.5460.065
- 1238.9%0.4330.025
- 1335.4%—0.004
- 1434.8%—0.033
- 1533.3%0.5430.040
- 1633.3%—0.005
- 1732.8%—-0.004
- 1826.8%—0.040
- 1925.8%0.5630.040
- 2025.3%—-0.007
- 21Executioner22.7%—-0.006
Each dot is a model: further right costs more, higher up scores more. Where the dashed lines cross stands the executioner, the best value on the board. Everything below and to its right costs more and scores no higher, so it is killed. Above the line is SOTA, the top tier; below and to the left is cheaper and weaker, the low-cost picks. How the kill line is drawn →
Claude Sonnet 5.5
- Artificial Analysis LLM Leaderboard56.0No. 2 of 372
- Arena — Agent0.125No. 3 of 49
- Artificial Analysis Coding Agents0.684No. 1 of 22
- Arena — Code1786No. 3 of 109
- Arena — Text1471No. 45 of 306
- Arena — Vision1268No. 35 of 128
Each rank is the board’s own: the rank it publishes, or its place in its own order where it publishes none. The bar is how much of that board the model is ahead of. Scores from different boards cannot be compared.
Artificial Analysis LLM Leaderboardbest at max
- low35.9
- medium40.8
- high46.8
- xhigh51.9
- max56.0
Artificial Analysis Coding Agents · Claude Codebest at max
- low0.421
- medium0.459
- high0.550
- xhigh0.629
- max0.684
Arena — Codebest at xhigh
- high1715
- xhigh1786
Tested at one setting only
- Arena — Agent · agentmax0.125
- Arena — Textxhigh1471
- Arena — Visionxhigh1268
The same model at each effort setting. A bar runs from zero to that board’s best score, so read each board on its own and never across boards.
Where this board’s numbers come from
Each panel is one board’s ranking as published, never averaged. When two disagree, they are usually measuring different things.
Models
The model itself, whatever it was driven by. Artificial Analysis grades completed tasks; Arena observes real sessions and scores tool reliability, task completion and steerability over its own traffic.
Harnesses
The same rows read the other way round: an agent is a program, and the model is what it drives. HarnessTax runs every model under every harness, so a gap between two of its rows is the harness's alone; Artificial Analysis covers more harnesses, mostly on their own vendor's models.