OnlySOTA

Coding

Writing and editing code against a specification, measured without an agent loop. Distinct from agentic coding: these benchmarks score the model's output directly rather than a model-plus-harness system.

Ordered by how many of these benchmarks place a model in their top 10, then by its best placing. Each column keeps its own order; nothing here is averaged.

Model In top 10 Artificial Analysis LLM Leaderboard SciCode Artificial Analysis LLM Leaderboard Coding Index Arena — Code Overall
Claude Fable 5 (Opus 4.8 fallback) anthropic 3/3 #1 60.2% #5 76.5 #4 1626
GPT-5.6 Sol openai 3/3 #5 56.9% #1 78.3 #5 1619
Claude Opus 5 anthropic 3/3 #8 55.7% #2 78.0 #1 1691
Kimi K3 kimi 3/3 #3 58.7% #6 76.2 #2 1674
Grok 4.6 xai 3/3 #9 54.6% #3 76.8 #3 1629
Gemini 3.7 Flash google 3/3 #4 57.9% #7 76.1 #7 1587
GLM-5.3 zai 3/3 #6 56.5% #8 74.8 #6 1599
Muse Spark 1.2 meta 2/3 #7 56.4% #10 72.2 #16 1534
Grok 4.5 xai 2/3 #10 54.1% #9 72.4 #11 1556
Gemini 3.1 Pro Preview google 1/3 #2 58.9% #18 68.8 #31 1446
GPT-5.6 Terra openai 1/3 #11 53.9% #4 76.7 #18 1520
GLM-5.2 zai 1/3 #19 50.5% #19 68.8 #8 1582
Claude Opus 4.8 anthropic 1/3 #9 1563
Claude Opus 4.7 anthropic 1/3 #10 1558

A dash means the source does not cover that model, not that it scored zero. Models with no registry entry are excluded here because they cannot be joined across sources — they are still shown, flagged, on their own source's board.