Instruction following
Whether the model obeys explicit constraints on its output — format, length, inclusions, exclusions — independently of whether the content is any good.
Ordered by how many of these benchmarks place a model in their top 10, then by its best placing. Each column keeps its own order; nothing here is averaged.
| Model | In top 10 | Artificial Analysis LLM Leaderboard IFBench |
|---|---|---|
| | 1/1 | #1 83.3% |
| | 1/1 | #2 82.9% |
| | 1/1 | #3 81.4% |
| | 1/1 | #4 80.4% |
| | 1/1 | #5 79.9% |
| | 1/1 | #6 79.6% |
| | 1/1 | #7 78.8% |
| | 1/1 | #8 78.0% |
| | 1/1 | #9 77.1% |
| | 1/1 | #10 76.5% |
A dash means the source does not cover that model, not that it scored zero. Models with no registry entry are excluded here because they cannot be joined across sources — they are still shown, flagged, on their own source's board.