Reasoning and knowledge
Hard questions with verifiable answers, where the difficulty is in the reasoning rather than in recall alone. Scores here are low in absolute terms by design — these are the benchmarks chosen to still have headroom.
Ordered by how many of these benchmarks put a model in their top 10, then by its best placing. Each column keeps its own ranking; nothing is averaged.
| Model | In top 10 | Artificial Analysis LLM Leaderboard GPQA Diamond | Artificial Analysis LLM Leaderboard Humanity's Last Exam | Artificial Analysis LLM Leaderboard CritPt |
|---|---|---|---|---|
| | 3/3 | #1 96.3% | #7 54.7% | #3 31.7% |
| | 3/3 | #5 94.1% | #9 49.5% | #1 32.3% |
| | 3/3 | #8 93.7% | #2 59.1% | #6 31.1% |
| | 2/3 | — | #1 61.4% | #2 31.7% |
| | 2/3 | — | #8 52.9% | #4 31.7% |
| | 2/3 | — | #5 55.0% | #5 31.4% |
| | 2/3 | #9 93.7% | #6 54.9% | #11 29.1% |
| | 1/3 | #2 95.3% | #15 47.8% | #28 18.3% |
| | 1/3 | #3 94.9% | #22 44.1% | #25 19.7% |
| | 1/3 | — | #3 57.1% | #14 27.1% |
| | 1/3 | #4 94.5% | #14 47.9% | #40 14.3% |
| | 1/3 | #17 92.6% | #4 55.5% | #12 28.6% |
| | 1/3 | #6 94.1% | #16 47.0% | #31 17.7% |
| | 1/3 | #7 94.1% | #11 48.7% | #16 26.0% |
| | 1/3 | — | #13 47.9% | #7 30.9% |
| | 1/3 | — | — | #8 30.6% |
| | 1/3 | — | — | #9 30.0% |
| | 1/3 | #10 93.5% | #20 45.8% | #13 27.1% |
| | 1/3 | #18 92.5% | #27 42.9% | #10 30.0% |
| | 1/3 | — | #10 49.4% | #15 26.6% |
A dash means the board does not list that model — not a score of zero. Entries not yet matched to a model are left out here; they still appear on their own board.