OnlySOTA

Instruction following

Whether the model obeys explicit constraints on its output — format, length, inclusions, exclusions — independently of whether the content is any good.

Ordered by how many of these benchmarks place a model in their top 10, then by its best placing. Each column keeps its own order; nothing here is averaged.

Model In top 10 Artificial Analysis LLM Leaderboard IFBench
Grok 4.3 xai 1/1 #1 83.3%
MiniMax-M3 minimax 1/1 #2 82.9%
Nemotron 3 Ultra 550B A55B nvidia 1/1 #3 81.4%
Nemotron Cascade 2 30B A3B nvidia 1/1 #4 80.4%
MiMo-V2.5-Pro xiaomi 1/1 #5 79.9%
Nova 2.0 Pro Preview aws 1/1 #6 79.6%
Qwen3.5 397B A17B alibaba 1/1 #7 78.8%
Qwen3.7 Plus alibaba 1/1 #8 78.0%
Gemini 3.1 Pro Preview google 1/1 #9 77.1%
DeepSeek V4 Pro deepseek 1/1 #10 76.5%

A dash means the source does not cover that model, not that it scored zero. Models with no registry entry are excluded here because they cannot be joined across sources — they are still shown, flagged, on their own source's board.