Multimodal
Questions that cannot be answered from the text alone. Restricted to visual reasoning for now, because that is what the sources measure.
Ordered by how many of these benchmarks put a model in their top 10, then by its best placing. Each column keeps its own ranking; nothing is averaged.
| Model | In top 10 | Artificial Analysis LLM Leaderboard MMMU-Pro | Arena — Vision Overall (style control) |
|---|---|---|---|
| | 2/2 | #3 86.0% | #8 1291 |
| | 2/2 | #4 85.6% | #9 1290 |
| | 2/2 | #5 85.5% | #5 1296 |
| | 1/2 | #1 87.7% | — |
| | 1/2 | — | #1 1309 |
| | 1/2 | #2 86.9% | #18 1284 |
| | 1/2 | #11 82.8% | #2 1301 |
| | 1/2 | #44 75.4% | #3 1299 |
| | 1/2 | #27 78.8% | #4 1298 |
| | 1/2 | #6 84.7% | #12 1289 |
| | 1/2 | #17 80.5% | #6 1294 |
| | 1/2 | #7 84.3% | #17 1284 |
| | 1/2 | — | #7 1293 |
| | 1/2 | #8 83.4% | #16 1285 |
| | 1/2 | #9 83.2% | #21 1280 |
| | 1/2 | #10 82.9% | — |
| | 1/2 | #13 82.0% | #10 1290 |
A dash means the board does not list that model — not a score of zero. Entries not yet matched to a model are left out here; they still appear on their own board.