OnlySOTA

Long context

Reasoning that requires holding a large input in view at once. A large advertised context window is not evidence of this; these benchmarks measure whether the window is usable.

Ordered by how many of these benchmarks put a model in their top 10, then by its best placing. Each column keeps its own ranking; nothing is averaged.

Model In top 10 Artificial Analysis LLM Leaderboard AA-LCR
Kimi K3 kimi 1/1 #1 88.7%
Step 5 stepfun 1/1 #2 88.3%
MiMo V2.6 Pro xiaomi 1/1 #3 86.3%
Claude Fable 5.1 anthropic 1/1 #4 85.3%
Claude Opus 5.5 anthropic 1/1 #5 84.7%
GPT 5.5 openai 1/1 #6 84.3%
DeepSeek V4.1 Flash deepseek 1/1 #7 84.0%
GPT 5.6 Sol openai 1/1 #8 84.0%
GPT 6.1 Sol openai 1/1 #9 84.0%
Gemini 3.8 Flash google 1/1 #10 84.0%

A dash means the board does not list that model — not a score of zero. Entries not yet matched to a model are left out here; they still appear on their own board.