OnlySOTA

Instruction following

Whether the model obeys explicit constraints on its output — format, length, inclusions, exclusions — independently of whether the content is any good.

Ordered by how many of these benchmarks put a model in their top 10, then by its best placing. Each column keeps its own ranking; nothing is averaged.

Model In top 10 Artificial Analysis LLM Leaderboard IFBench
Grok 4.3 xai 1/1 #1 83.3%
Grok 4.20 xai 1/1 #2 82.9%
MiniMax M3 minimax 1/1 #3 82.9%
Nemotron 3 Ultra nvidia 1/1 #4 81.4%
Qwen 3.7 Max alibaba 1/1 #5 80.5%
Nemotron Cascade 2 nvidia 1/1 #6 80.4%
MiMo V2.5 Pro xiaomi 1/1 #7 79.9%
Nova 2.0 Pro aws 1/1 #8 79.6%
DeepSeek V4 Flash deepseek 1/1 #9 79.2%
Qwen 3.5 alibaba 1/1 #10 78.8%

A dash means the board does not list that model — not a score of zero. Entries not yet matched to a model are left out here; they still appear on their own board.