OnlySOTA

Reasoning and knowledge

Hard questions with verifiable answers, where the difficulty is in the reasoning rather than in recall alone. Scores here are low in absolute terms by design — these are the benchmarks chosen to still have headroom.

Ordered by how many of these benchmarks put a model in their top 10, then by its best placing. Each column keeps its own ranking; nothing is averaged.

Model In top 10 Artificial Analysis LLM Leaderboard GPQA Diamond Artificial Analysis LLM Leaderboard Humanity's Last Exam Artificial Analysis LLM Leaderboard CritPt
GPT 6 Astra openai 3/3 #1 96.3% #7 54.7% #3 31.7%
GPT 5.6 Sol openai 3/3 #5 94.1% #9 49.5% #1 32.3%
Claude Fable 5.1 anthropic 3/3 #8 93.7% #2 59.1% #6 31.1%
Claude Opus 5.5 anthropic 2/3 — #1 61.4% #2 31.7%
GPT 6.1 Sol openai 2/3 — #8 52.9% #4 31.7%
Claude Sonnet 5.5 anthropic 2/3 — #5 55.0% #5 31.4%
Claude Opus 5 anthropic 2/3 #9 93.7% #6 54.9% #11 29.1%
Gemini 3.8 Flash google 1/3 #2 95.3% #15 47.8% #28 18.3%
Grok 4.6 xai 1/3 #3 94.9% #22 44.1% #25 19.7%
Gemini 4 Argon google 1/3 — #3 57.1% #14 27.1%
Gemini 3.7 Flash google 1/3 #4 94.5% #14 47.9% #40 14.3%
Claude Fable 5 anthropic 1/3 #17 92.6% #4 55.5% #12 28.6%
Gemini 3.1 Pro google 1/3 #6 94.1% #16 47.0% #31 17.7%
Muse Spark 1.3 meta 1/3 #7 94.1% #11 48.7% #16 26.0%
GPT 6 Sol openai 1/3 — #13 47.9% #7 30.9%
GPT 5.5 Pro openai 1/3 — — #8 30.6%
GPT 5.4 Pro openai 1/3 — — #9 30.0%
GPT 5.5 openai 1/3 #10 93.5% #20 45.8% #13 27.1%
GPT 5.6 Terra openai 1/3 #18 92.5% #27 42.9% #10 30.0%
MiMo V2.6 Pro xiaomi 1/3 — #10 49.4% #15 26.6%

A dash means the board does not list that model — not a score of zero. Entries not yet matched to a model are left out here; they still appear on their own board.