OnlySOTA

Agent

Not one answer but a task seen through, with tools. Two questions live here, and they are not the same one: which model to send in, and which harness to drive it with.

No. 1 · AA
63.6%
Claude Sonnet 5.5 · $4.00
Executioner · price
$0.17
MiMo V2.6 Flash
SOTA
16
Score above the executioner
Killed
39
Dearer, and no better
Ranking · AA 1–64% · AA-CODE 0.42–0.68 · A-AGENT -0.02–0.16
  1. 163.6%0.6840.125
  2. 259.6%0.6600.138
  3. 359.6%0.6160.123
  4. 4New57.1%0.6380.076
  5. 556.1%0.6290.112
  6. 655.1%0.6220.143
  7. 7Legacy49.0%0.5970.087
  8. 8Legacy43.9%0.5670.097
  9. 9Legacy42.4%—0.082
  10. 1041.9%0.5360.024
  11. 11Legacy39.9%0.5460.065
  12. 1238.9%0.4330.025
  13. 1335.4%—0.004
  14. 1434.8%—0.033
  15. 1533.3%0.5430.040
  16. 1633.3%—0.005
  17. 1732.8%—-0.004
  18. 1826.8%—0.040
  19. 1925.8%0.5630.040
  20. 2025.3%—-0.007
  21. 21Executioner22.7%—-0.006
All 115 models on AA →
Kill line · score × price

Each dot is a model: further right costs more, higher up scores more. Where the dashed lines cross stands the executioner, the best value on the board. Everything below and to its right costs more and scores no higher, so it is killed. Above the line is SOTA, the top tier; below and to the left is cheaper and weaker, the low-cost picks. How the kill line is drawn →

Terminal-Bench 4.0 · starts at 020%40%60%0$0.05$0.25$1.00$20.0063.6%$4.00Claude Sonnet 5.5ExecutionerMiMo V2.6 Flash22.7% · $0.17SOTALOW-COSTKILLEDBlended Price, USD per 1M tokens (3:1 input:output) · log10
Kill lineFrontier — nothing beats these on both price and scoreSOTALow-costKilledHollow dot or faint cross — estimated by the source

Claude Sonnet 5.5

anthropic/claude-sonnet-5-5 · closed · $4.00 · SOTA

Open its full page →
Scores on each board
  • Artificial Analysis LLM Leaderboard56.0No. 2 of 372
  • Arena — Agent0.125No. 3 of 49
  • Artificial Analysis Coding Agents0.684No. 1 of 22
  • Arena — Code1786No. 3 of 109
  • Arena — Text1471No. 45 of 306
  • Arena — Vision1268No. 35 of 128

Each rank is the board’s own: the rank it publishes, or its place in its own order where it publishes none. The bar is how much of that board the model is ahead of. Scores from different boards cannot be compared.

Scores at each effort setting

Artificial Analysis LLM Leaderboardbest at max

  1. low35.9
  2. medium40.8
  3. high46.8
  4. xhigh51.9
  5. max56.0

Artificial Analysis Coding Agents · Claude Codebest at max

  1. low0.421
  2. medium0.459
  3. high0.550
  4. xhigh0.629
  5. max0.684

Arena — Codebest at xhigh

  1. high1715
  2. xhigh1786

Tested at one setting only

  • Arena — Agent · agentmax0.125
  • Arena — Textxhigh1471
  • Arena — Visionxhigh1268

The same model at each effort setting. A bar runs from zero to that board’s best score, so read each board on its own and never across boards.

Where this board’s numbers come from

Each panel is one board’s ranking as published, never averaged. When two disagree, they are usually measuring different things.

Models

The model itself, whatever it was driven by. Artificial Analysis grades completed tasks; Arena observes real sessions and scores tool reliability, task completion and steerability over its own traffic.

Harnesses

The same rows read the other way round: an agent is a program, and the model is what it drives. HarnessTax runs every model under every harness, so a gap between two of its rows is the harness's alone; Artificial Analysis covers more harnesses, mostly on their own vendor's models.