OnlySOTA

Coding

Writing and editing code against a specification, measured without an agent loop. Distinct from agentic coding: these benchmarks score the model's output directly rather than a model-plus-harness system.

Ordered by how many of these benchmarks put a model in their top 10, then by its best placing. Each column keeps its own ranking; nothing is averaged.

Model In top 10 Artificial Analysis LLM Leaderboard SciCode Arena — Code Overall
Claude Opus 5.5 anthropic 2/2 #1 66.9% #1 1815
Claude Fable 5.1 anthropic 2/2 #2 63.1% #5 1749
Gemini 4 Argon google 2/2 #3 61.8% #8 1680
Claude Sonnet 5.5 anthropic 2/2 #5 61.0% #3 1786
Kimi K3 kimi 2/2 #9 59.5% #10 1658
GPT 6 Astra openai 1/2 #19 56.5% #2 1788
Claude Fable 5 anthropic 1/2 #4 61.0% #15 1625
GPT 6.1 Sol openai 1/2 #23 55.8% #4 1758
MiMo V2.6 Pro xiaomi 1/2 #6 60.9% #20 1618
Claude Opus 5 anthropic 1/2 #21 56.4% #6 1695
Gemini 3.7 Flash google 1/2 #7 59.8% #23 1592
GPT 6 Sol openai 1/2 #16 57.6% #7 1689
Muse Spark 1.3 meta 1/2 #8 59.7% #11 1657
Qwen 3.8 Max alibaba 1/2 #29 54.1% #9 1671
GLM 5.3 zai 1/2 #10 59.0% #16 1623

A dash means the board does not list that model — not a score of zero. Entries not yet matched to a model are left out here; they still appear on their own board.