OnlySOTA

Multimodal

Questions that cannot be answered from the text alone. Restricted to visual reasoning for now, because that is what the sources measure.

Ordered by how many of these benchmarks place a model in their top 10, then by its best placing. Each column keeps its own order; nothing here is averaged.

Model In top 10 Artificial Analysis LLM Leaderboard MMMU-Pro Arena — Vision Overall (style control)
Claude Opus 5 anthropic 2/2 #2 84.7% #6 1292
Gemini 3.5 Flash google 2/2 #3 83.9% #8 1286
Gemini 3.6 Flash google 2/2 #5 83.2% #10 1285
Gemini 3.7 Flash google 1/2 #1 85.5%
Claude Fable 5 (Opus 4.8 fallback) anthropic 1/2 #1 1312
Claude Opus 4.7 anthropic 1/2 #2 1301
Claude Opus 4.6 anthropic 1/2 #3 1299
GPT-5.6 Sol openai 1/2 #4 83.4% #14 1281
Muse Spark meta 1/2 #4 1294
Muse Spark 1.2 meta 1/2 #5 1292
Gemini 3.1 Pro Preview google 1/2 #6 82.4% #16 1277
Qwen3.8 Max alibaba 1/2 #7 82.3%
Gemini 3 Pro Preview google 1/2 #7 1289
GPT-5.6 Terra openai 1/2 #8 80.7% #21 1266
Kimi K3 kimi 1/2 #9 80.5%
Claude Opus 4.8 anthropic 1/2 #9 1285
Qwen3.7 Plus alibaba 1/2 #10 80.5% #22 1265

A dash means the source does not cover that model, not that it scored zero. Models with no registry entry are excluded here because they cannot be joined across sources — they are still shown, flagged, on their own source's board.