OnlySOTA

Factuality

Whether the model states things that are not so, and whether it declines when it does not know. Measured both as accuracy and as the rate of confident wrong answers, which move independently.

Ordered by how many of these benchmarks put a model in their top 10, then by its best placing. Each column keeps its own ranking; nothing is averaged.

Model In top 10 Artificial Analysis LLM Leaderboard AA-Omniscience Index Artificial Analysis LLM Leaderboard AA-Omniscience Non-Hallucination Rate
Gemini 4 Argon google 2/2 #5 42.4 #5 84.9%
Claude Opus 5.5 anthropic 1/2 #1 46.4 #93 41.4%
MiniCPM5 openbmb 1/2 #71 -0.800 #1 99.1%
GPT 6 Astra openai 1/2 #2 43.7 #65 55.2%
G9v3 ai9star 1/2 #79 -4.367 #2 88.3%
Claude Fable 5.1 anthropic 1/2 #3 43.5 #113 34.4%
G9v3 39A5B ai9star 1/2 #53 3.767 #3 87.0%
Claude Fable 5 anthropic 1/2 #4 43.3 #107 36.4%
Command A+ cohere 1/2 #77 -4.017 #4 85.8%
GPT 6.1 Sol openai 1/2 #6 41.5 #74 50.6%
LFM2.5 liquidai 1/2 #98 -10.9 #6 84.0%
Claude Opus 5 anthropic 1/2 #7 37.1 #95 40.5%
Grok 4.3 xai 1/2 #27 18.0 #7 83.1%
Claude Sonnet 5.5 anthropic 1/2 #8 32.3 #69 53.0%
Grok 4.20 xai 1/2 #31 14.8 #8 82.6%
Grok 4.7 xai 1/2 #9 32.0 #29 70.7%
Qwen 3.8 alibaba 1/2 #87 -7.950 #9 81.9%
Gemini 3.1 Pro google 1/2 #10 31.9 #78 49.1%
MiniMax M3 minimax 1/2 #57 1.350 #10 81.6%

A dash means the board does not list that model — not a score of zero. Entries not yet matched to a model are left out here; they still appear on their own board.