Factuality
Whether the model states things that are not so, and whether it declines when it does not know. Measured both as accuracy and as the rate of confident wrong answers, which move independently.
Ordered by how many of these benchmarks place a model in their top 10, then by its best placing. Each column keeps its own order; nothing here is averaged.
| Model | In top 10 | Artificial Analysis LLM Leaderboard AA-Omniscience Index | Artificial Analysis LLM Leaderboard AA-Omniscience Non-Hallucination Rate |
|---|---|---|---|
| | 2/2 | #4 30.5 | #8 76.0% |
| | 1/2 | #1 43.3 | #56 36.4% |
| | 1/2 | #30 -0.800 | #1 99.1% |
| | 1/2 | #2 37.1 | #48 40.5% |
| | 1/2 | #36 -4.367 | #2 88.3% |
| | 1/2 | #3 31.9 | #41 49.1% |
| | 1/2 | #20 3.767 | #3 87.0% |
| | 1/2 | #34 -4.017 | #4 85.8% |
| | 1/2 | #5 27.2 | #25 66.7% |
| | 1/2 | #12 16.7 | #5 83.1% |
| | 1/2 | #6 26.5 | #58 35.5% |
| | 1/2 | #25 1.350 | #6 81.6% |
| | 1/2 | #7 25.3 | #44 45.9% |
| | 1/2 | #37 -6.600 | #7 77.4% |
| | 1/2 | #8 22.1 | #45 44.4% |
| | 1/2 | #9 22.0 | #141 10.6% |
| | 1/2 | #31 -0.850 | #9 75.6% |
| | 1/2 | #10 20.8 | #53 38.2% |
| | 1/2 | #23 3.250 | #10 75.3% |
A dash means the source does not cover that model, not that it scored zero. Models with no registry entry are excluded here because they cannot be joined across sources — they are still shown, flagged, on their own source's board.