OnlySOTA

Coding

Writing and changing code. Split by how it is judged, because a grader running tests and a person picking the better of two answers do not measure the same thing and do not agree.

Each panel is one source's own ranking, shown as it publishes it. Nothing on this page is averaged across panels — where two disagree, both are right about what they measured.

Graded

Scored against tests or a reference, with no human in the loop and no agent loop either — these measure the model's output directly.

Web development

Human preference between two built web applications. Arena's code arena is web work end to end — every category it publishes is a web application category — so this is the board it actually is, rather than a general coding board with a narrower name.

Arena — Code

Overall as of 2026-08-21

  1. Claude Opus 5 1691
  2. Kimi K3 1674
  3. qwen3.8-max 1669
  4. Grok 4.6 1629
  5. Claude Fable 5 (Opus 4.8 fallback) 1626
  6. GPT-5.6 Sol 1619
  7. GLM-5.3 1599
  8. qwen3.8-27b 1595
  9. Gemini 3.7 Flash 1587
  10. GLM-5.2 1582

all 104 on Arena — Code