OnlySOTA

Agentic work

Multi-step tasks drawn from real occupations rather than from software engineering — spreadsheets, documents, customer-facing procedures — run through a tool-using loop.

Ordered by how many of these benchmarks place a model in their top 10, then by its best placing. Each column keeps its own order; nothing here is averaged.

Model In top 10 Artificial Analysis LLM Leaderboard τ³-Banking Artificial Analysis LLM Leaderboard GDPval-AA Arena — Agent Overall
Qwen3.8 Max alibaba 3/3 #1 51.3% #5 61.8% #9 0.062
Claude Opus 5 anthropic 3/3 #7 44.7% #1 67.2% #1 0.125
Kimi K3 kimi 3/3 #6 46.0% #8 58.9% #2 0.104
GPT-5.6 Sol openai 3/3 #8 44.3% #6 61.1% #3 0.097
Grok 4.6 xai 2/3 #2 50.7% #3 63.3%
GLM-5.3 zai 2/3 #3 50.3% #2 63.4%
Qwen3.8 2.4T A95B alibaba 2/3 #4 49.1% #7 61.0%
Claude Sonnet 5 anthropic 2/3 #14 37.3% #10 54.8% #7 0.066
Grok 4.5 xai 2/3 #9 42.1% #17 51.2% #10 0.062
Claude Fable 5 (Opus 4.8 fallback) anthropic 1/3 #13 38.1% #4 61.9%
Claude Opus 4.8 anthropic 1/3 #4 0.095
Qwen3.8 27B alibaba 1/3 #5 48.0% #15 52.3%
GPT-5.5 openai 1/3 #5 0.085
Claude Opus 4.7 anthropic 1/3 #6 0.081
Claude Opus 4.6 anthropic 1/3 #8 0.066
Muse Spark 1.2 meta 1/3 #17 34.8% #9 56.4% #17 0.021
GPT-5.6 Terra openai 1/3 #10 40.2% #13 53.8% #15 0.032

A dash means the source does not cover that model, not that it scored zero. Models with no registry entry are excluded here because they cannot be joined across sources — they are still shown, flagged, on their own source's board.