OnlySOTA

Speed

How fast output arrives — sustained throughput and time to the first token, which trade off against each other and against cost. Measured on a provider's serving of the model, so it describes an endpoint rather than the weights.

Ordered by how many of these benchmarks place a model in their top 10, then by its best placing. Each column keeps its own order; nothing here is averaged.

Model In top 10 Artificial Analysis LLM Leaderboard Output Speed Artificial Analysis LLM Leaderboard Time To First Token
Celeris-1 celeris 2/2 #1 1573/s #3 0.58s
Command A+ cohere 1/2 #19 196/s #1 0.39s
Mercury 2 inception 1/2 #2 924/s #94 3.54s
North Mini Code cohere 1/2 #105 27/s #2 0.53s
Ling 3.0 Flash inclusionai 1/2 #3 407/s #85 2.70s
Gemini 3.5 Flash-Lite google 1/2 #4 407/s #102 8.96s
Ministral 3 3B mistral 1/2 #16 204/s #4 0.63s
Gemini 3.7 Flash google 1/2 #5 394/s #27 1.04s
Gemma 4 E4B google 1/2 #86 58/s #5 0.65s
LFM2.5-VL-1.6B liquidai 1/2 #6 346/s #33 1.13s
Ministral 3 8B mistral 1/2 #52 114/s #6 0.74s
LFM2.5-8B-A1B liquidai 1/2 #7 339/s #44 1.46s
Qwen3.5 9B alibaba 1/2 #66 86/s #7 0.77s
Nemotron 3 Nano Omni 30B A3B Reasoning nvidia 1/2 #8 326/s #25 1.00s
Granite 4.1 8B ibm 1/2 #53 112/s #8 0.77s
Nemotron 3.5 Lightning nvidia 1/2 #9 314/s #24 0.96s
Llama 4 Scout meta 1/2 #41 133/s #9 0.79s
Nova Micro aws 1/2 #10 273/s #21 0.90s
Muse Glimmer meta 1/2 #59 104/s #10 0.79s

A dash means the source does not cover that model, not that it scored zero. Models with no registry entry are excluded here because they cannot be joined across sources — they are still shown, flagged, on their own source's board.