OnlySOTA

Multimodal

Questions that cannot be answered from the text alone. Restricted to visual reasoning for now, because that is what the sources measure.

Ordered by how many of these benchmarks put a model in their top 10, then by its best placing. Each column keeps its own ranking; nothing is averaged.

Model In top 10 Artificial Analysis LLM Leaderboard MMMU-Pro Arena — Vision Overall (style control)
GPT 6.1 Sol openai 2/2 #3 86.0% #8 1291
Gemini 3.8 Flash google 2/2 #4 85.6% #9 1290
Gemini 3.7 Flash google 2/2 #5 85.5% #5 1296
Claude Opus 5.5 anthropic 1/2 #1 87.7% —
Claude Fable 5 anthropic 1/2 — #1 1309
GPT 6 Astra openai 1/2 #2 86.9% #18 1284
Qwen 3.8 Max alibaba 1/2 #11 82.8% #2 1301
Claude Opus 4.6 anthropic 1/2 #44 75.4% #3 1299
Claude Opus 4.7 anthropic 1/2 #27 78.8% #4 1298
Claude Opus 5 anthropic 1/2 #6 84.7% #12 1289
Muse Spark meta 1/2 #17 80.5% #6 1294
Gemini 3.5 Flash google 1/2 #7 84.3% #17 1284
Muse Spark 1.2 meta 1/2 — #7 1293
GPT 5.6 Sol openai 1/2 #8 83.4% #16 1285
Gemini 3.6 Flash google 1/2 #9 83.2% #21 1280
GPT 6 Sol openai 1/2 #10 82.9% —
Muse Spark 1.3 meta 1/2 #13 82.0% #10 1290

A dash means the board does not list that model — not a score of zero. Entries not yet matched to a model are left out here; they still appear on their own board.