OnlySOTA

Agent

How well a model works through a task with tools rather than answering in one turn. Two questions live here and they are not the same one: which model to drive, and which harness to drive it with.

Each panel is one source's own ranking, shown as it publishes it. Nothing on this page is averaged across panels — where two disagree, both are right about what they measured.

Models

The model itself, whatever it was driven by. Artificial Analysis grades completed tasks; Arena observes real sessions and scores tool reliability, task completion and steerability over its own traffic.

Arena — Agent

Overall as of 2026-08-19

  1. Claude Opus 5 0.125
  2. Claude Fable 5 (High) 0.116
  3. Kimi K3 0.104
  4. GPT-5.6 Sol 0.097
  5. Claude Opus 4.8 0.095
  6. GPT-5.5 0.085
  7. Claude Opus 4.7 0.081
  8. Claude Sonnet 5 0.066
  9. Claude Opus 4.6 0.066
  10. DeepSeek V4 Pro (High) (0813) 0.063

all 44 on Arena — Agent

Harnesses

The same rows read the other way round: an agent is a program, and the model is what it drives. This has one source, because it is the only board found that names both halves of the pair.

Artificial Analysis Coding Agents

Artificial Analysis Coding Agent Index by harness · fetched 2026-08-20

  1. Claude Code Claude Opus 5 0.667
  2. Codex GPT-5.6 Sol 0.666
  3. Grok Build Grok 4.5 0.644
  4. Kimi Code CLI Kimi K3 0.613
  5. Muse Code Muse Spark 1.2 0.605
  6. Opencode Gemini 3.7 Flash 0.571
  7. Antigravity SDK v0.1.8 Gemini 3.7 Flash 0.564
  8. Cursor CLI GPT-5.5 0.461
  9. Gemini CLI Gemini 3.1 Pro (high) 0.303