OnlySOTA

Agentic coding

Resolving real repository and terminal tasks through a tool-using loop. Every score here measures a model together with a scaffold, so entries carry the scaffold as a variant and two rows for one model are expected rather than a duplicate.

Ordered by how many of these benchmarks place a model in their top 10, then by its best placing. Each column keeps its own order; nothing here is averaged.

Model In top 10 Artificial Analysis LLM Leaderboard Terminal-Bench 2.1 Artificial Analysis LLM Leaderboard Agentic Index Artificial Analysis Coding Agents Artificial Analysis Coding Agent Index Artificial Analysis Coding Agents Terminal-Bench v2
GPT-5.6 Sol openai 4/4 #1 89.5% #5 57.8 #2 0.666 #1 87.7%
Claude Opus 5 anthropic 4/4 #2 89.1% #1 59.2 #1 0.667 #5 84.9%
GPT-5.6 Terra openai 4/4 #4 88.0% #10 50.2 #4 0.623 #7 84.1%
Kimi K3 kimi 4/4 #6 85.0% #8 54.3 #6 0.613 #8 83.7%
Grok 4.5 xai 3/4 #10 81.6% #14 48.9 #3 0.644 #3 85.3%
Gemini 3.7 Flash google 2/4 #5 85.8% #18 45.1 #11 0.571 #2 86.1%
GLM-5.3 zai 2/4 #8 83.9% #2 59.1
Grok 4.6 xai 2/4 #3 88.4% #3 58.7
Qwen3.8 Max alibaba 2/4 #11 81.3% #4 58.4 #9 0.598 #11 79.4%
GPT-5.5 openai 2/4 #5 0.615 #6 84.1%
Qwen3.8 2.4T A95B alibaba 2/4 #9 82.0% #6 57.1
Claude Fable 5 (Opus 4.8 fallback) anthropic 2/4 #7 84.6% #7 56.6
Claude Opus 4.8 anthropic 2/4 #7 0.606 #9 81.3%
GPT-5.6 Luna openai 2/4 #12 80.9% #16 46.9 #10 0.587 #10 79.8%
DeepSeek V4 Flash deepseek 1/4 #12 0.555 #4 84.9%
Muse Spark 1.2 meta 1/4 #14 80.1% #13 49.3 #8 0.605 #12 78.6%
Qwen3.8 27B alibaba 1/4 #15 79.8% #9 50.9

A dash means the source does not cover that model, not that it scored zero. Models with no registry entry are excluded here because they cannot be joined across sources — they are still shown, flagged, on their own source's board.