OnlySOTA

Reasoning and knowledge

Hard questions with verifiable answers, where the difficulty is in the reasoning rather than in recall alone. Scores here are low in absolute terms by design — these are the benchmarks chosen to still have headroom.

Ordered by how many of these benchmarks place a model in their top 10, then by its best placing. Each column keeps its own order; nothing here is averaged.

Model In top 10 Artificial Analysis LLM Leaderboard GPQA Diamond Artificial Analysis LLM Leaderboard Humanity's Last Exam Artificial Analysis LLM Leaderboard CritPt
GPT-5.6 Sol openai 3/3 #3 94.1% #3 49.5% #1 32.3%
Claude Opus 5 anthropic 3/3 #5 93.7% #2 54.9% #4 29.1%
Kimi K3 kimi 3/3 #6 93.5% #6 46.9% #7 23.4%
Grok 4.6 xai 2/3 #1 94.9% #8 44.1% #12 19.7%
Claude Fable 5 (Opus 4.8 fallback) anthropic 2/3 #13 92.6% #1 55.5% #5 28.6%
Gemini 3.7 Flash google 2/3 #2 94.5% #4 47.9% #21 14.3%
GPT-5.6 Terra openai 2/3 #14 92.5% #10 42.9% #3 30.0%
Gemini 3.1 Pro Preview google 2/3 #4 94.1% #5 47.0% #15 17.7%
Qwen3.8 2.4T A95B alibaba 2/3 #7 93.5% #13 42.4% #10 20.0%
GPT-5.5 Pro openai 1/3 #2 30.6%
Gemini 3 Deep Think google 1/3 #6 25.7%
Muse Spark 1.2 meta 1/3 #23 90.4% #7 45.5% #16 17.7%
Grok 4.5 xai 1/3 #8 93.1% #11 42.7% #20 15.4%
GLM-5.2 zai 1/3 #27 89.5% #17 41.1% #8 20.9%
MiniMax-M3 minimax 1/3 #9 92.9% #22 39.0% #41 3.7%
Qwen3.8 Max alibaba 1/3 #12 92.7% #9 43.0% #11 20.0%
GPT-5.6 Luna openai 1/3 #19 91.1% #21 39.5% #9 20.6%
DeepSeek V4 Pro 0813 deepseek 1/3 #10 92.8% #18 41.0% #14 18.0%

A dash means the source does not cover that model, not that it scored zero. Models with no registry entry are excluded here because they cannot be joined across sources — they are still shown, flagged, on their own source's board.