OnlySOTA

Agentic work

Multi-step tasks drawn from real occupations rather than from software engineering — spreadsheets, documents, customer-facing procedures — run through a tool-using loop.

Ordered by how many of these benchmarks put a model in their top 10, then by its best placing. Each column keeps its own ranking; nothing is averaged.

Model In top 10 Artificial Analysis LLM Leaderboard τ³-Banking Artificial Analysis LLM Leaderboard GDPval-AA Arena — Agent Overall
Claude Fable 5.1 anthropic 3/3 #6 47.2% #3 61.7% #1 0.143
Qwen 3.8 Max alibaba 2/3 #1 51.3% #8 58.6% #22 0.025
Claude Opus 5.5 anthropic 2/3 — #1 67.3% #2 0.138
Claude Sonnet 5.5 anthropic 2/3 — #2 67.0% #3 0.125
Muse Spark 1.3 meta 2/3 #3 50.5% #7 58.6% #17 0.040
GLM 5.3 zai 2/3 #4 50.3% #9 57.8% #23 0.024
Claude Opus 5 anthropic 2/3 #11 44.7% #5 60.4% #7 0.087
GLM 5.3 Flash zai 2/3 #7 47.2% #10 57.3% #31 -0.004
Grok 4.6 xai 1/3 #2 50.7% #11 56.6% #25 0.013
GPT 6 Astra openai 1/3 #13 43.1% #20 52.1% #4 0.123
Grok 4.7 xai 1/3 — #4 60.5% #15 0.040
Qwen 3.8 alibaba 1/3 #5 48.0% #32 46.2% #36 -0.017
GPT 6.1 Sol openai 1/3 — #18 53.8% #5 0.112
MiMo V2.6 Pro xiaomi 1/3 — #6 59.3% #20 0.033
GPT 6 Sol openai 1/3 — #23 50.3% #6 0.097
Kimi K3 kimi 1/3 #8 46.0% #21 52.0% #13 0.042
Claude Fable 5 anthropic 1/3 #22 38.1% #16 54.8% #8 0.082
Gemini 3.8 Flash google 1/3 #9 45.8% #33 45.6% #21 0.030
Gemini 4 Argon google 1/3 — #13 55.6% #9 0.076
Qwen 3.8 Flash Next alibaba 1/3 #10 45.4% #12 55.6% #33 -0.007
Claude Opus 4.8 anthropic 1/3 #31 34.2% #30 46.9% #10 0.066

A dash means the board does not list that model — not a score of zero. Entries not yet matched to a model are left out here; they still appear on their own board.