OnlySOTA

Long context

Reasoning that requires holding a large input in view at once. A large advertised context window is not evidence of this; these benchmarks measure whether the window is usable.

Ordered by how many of these benchmarks place a model in their top 10, then by its best placing. Each column keeps its own order; nothing here is averaged.

Model In top 10 Artificial Analysis LLM Leaderboard AA-LCR
Muse Spark 1.2 meta 1/1 #1 83.3%
Kimi K3 kimi 1/1 #2 82.7%
Gemini 3.7 Flash google 1/1 #3 81.0%
MiniMax-M3 minimax 1/1 #4 80.3%
Muse Glimmer meta 1/1 #5 80.0%
GPT-5.6 Terra openai 1/1 #6 79.7%
Gemini 3.5 Flash google 1/1 #7 79.7%
Gemini 3.1 Pro Preview google 1/1 #8 79.0%
Gemini 3.6 Flash google 1/1 #9 79.0%
Claude Opus 5 anthropic 1/1 #10 78.7%

A dash means the source does not cover that model, not that it scored zero. Models with no registry entry are excluded here because they cannot be joined across sources — they are still shown, flagged, on their own source's board.