Reasoning and knowledge
Hard questions with verifiable answers, where the difficulty is in the reasoning rather than in recall alone. Scores here are low in absolute terms by design — these are the benchmarks chosen to still have headroom.
Ordered by how many of these benchmarks place a model in their top 10, then by its best placing. Each column keeps its own order; nothing here is averaged.
| Model | In top 10 | Artificial Analysis LLM Leaderboard GPQA Diamond | Artificial Analysis LLM Leaderboard Humanity's Last Exam | Artificial Analysis LLM Leaderboard CritPt |
|---|---|---|---|---|
| | 3/3 | #3 94.1% | #3 49.5% | #1 32.3% |
| | 3/3 | #5 93.7% | #2 54.9% | #4 29.1% |
| | 3/3 | #6 93.5% | #6 46.9% | #7 23.4% |
| | 2/3 | #1 94.9% | #8 44.1% | #12 19.7% |
| | 2/3 | #13 92.6% | #1 55.5% | #5 28.6% |
| | 2/3 | #2 94.5% | #4 47.9% | #21 14.3% |
| | 2/3 | #14 92.5% | #10 42.9% | #3 30.0% |
| | 2/3 | #4 94.1% | #5 47.0% | #15 17.7% |
| | 2/3 | #7 93.5% | #13 42.4% | #10 20.0% |
| | 1/3 | — | — | #2 30.6% |
| | 1/3 | — | — | #6 25.7% |
| | 1/3 | #23 90.4% | #7 45.5% | #16 17.7% |
| | 1/3 | #8 93.1% | #11 42.7% | #20 15.4% |
| | 1/3 | #27 89.5% | #17 41.1% | #8 20.9% |
| | 1/3 | #9 92.9% | #22 39.0% | #41 3.7% |
| | 1/3 | #12 92.7% | #9 43.0% | #11 20.0% |
| | 1/3 | #19 91.1% | #21 39.5% | #9 20.6% |
| | 1/3 | #10 92.8% | #18 41.0% | #14 18.0% |
A dash means the source does not cover that model, not that it scored zero. Models with no registry entry are excluded here because they cannot be joined across sources — they are still shown, flagged, on their own source's board.