Multimodal
Questions that cannot be answered from the text alone. Restricted to visual reasoning for now, because that is what the sources measure.
Ordered by how many of these benchmarks place a model in their top 10, then by its best placing. Each column keeps its own order; nothing here is averaged.
| Model | In top 10 | Artificial Analysis LLM Leaderboard MMMU-Pro | Arena — Vision Overall (style control) |
|---|---|---|---|
| | 2/2 | #2 84.7% | #6 1292 |
| | 2/2 | #3 83.9% | #8 1286 |
| | 2/2 | #5 83.2% | #10 1285 |
| | 1/2 | #1 85.5% | — |
| | 1/2 | — | #1 1312 |
| | 1/2 | — | #2 1301 |
| | 1/2 | — | #3 1299 |
| | 1/2 | #4 83.4% | #14 1281 |
| | 1/2 | — | #4 1294 |
| | 1/2 | — | #5 1292 |
| | 1/2 | #6 82.4% | #16 1277 |
| | 1/2 | #7 82.3% | — |
| | 1/2 | — | #7 1289 |
| | 1/2 | #8 80.7% | #21 1266 |
| | 1/2 | #9 80.5% | — |
| | 1/2 | — | #9 1285 |
| | 1/2 | #10 80.5% | #22 1265 |
A dash means the source does not cover that model, not that it scored zero. Models with no registry entry are excluded here because they cannot be joined across sources — they are still shown, flagged, on their own source's board.