OnlySOTA

Factuality

Whether the model states things that are not so, and whether it declines when it does not know. Measured both as accuracy and as the rate of confident wrong answers, which move independently.

Ordered by how many of these benchmarks place a model in their top 10, then by its best placing. Each column keeps its own order; nothing here is averaged.

Model In top 10 Artificial Analysis LLM Leaderboard AA-Omniscience Index Artificial Analysis LLM Leaderboard AA-Omniscience Non-Hallucination Rate
Grok 4.6 xai 2/2 #4 30.5 #8 76.0%
Claude Fable 5 (Opus 4.8 fallback) anthropic 1/2 #1 43.3 #56 36.4%
MiniCPM5-1B openbmb 1/2 #30 -0.800 #1 99.1%
Claude Opus 5 anthropic 1/2 #2 37.1 #48 40.5%
G9v3-3B ai9star 1/2 #36 -4.367 #2 88.3%
Gemini 3.1 Pro Preview google 1/2 #3 31.9 #41 49.1%
G9v3-39A5B ai9star 1/2 #20 3.767 #3 87.0%
Command A+ cohere 1/2 #34 -4.017 #4 85.8%
Muse Spark 1.2 meta 1/2 #5 27.2 #25 66.7%
Grok 4.3 xai 1/2 #12 16.7 #5 83.1%
Gemini 3.7 Flash google 1/2 #6 26.5 #58 35.5%
MiniMax-M3 minimax 1/2 #25 1.350 #6 81.6%
Grok 4.5 xai 1/2 #7 25.3 #44 45.9%
K-EXAONE 2.0 0803 lg 1/2 #37 -6.600 #7 77.4%
Gemini 3.6 Flash google 1/2 #8 22.1 #45 44.4%
GPT-5.6 Sol openai 1/2 #9 22.0 #141 10.6%
Solar Pro 4 upstage 1/2 #31 -0.850 #9 75.6%
Gemini 3.5 Flash google 1/2 #10 20.8 #53 38.2%
MiMo-V2.5-Pro xiaomi 1/2 #23 3.250 #10 75.3%

A dash means the source does not cover that model, not that it scored zero. Models with no registry entry are excluded here because they cannot be joined across sources — they are still shown, flagged, on their own source's board.