OnlySOTA

Speed

How fast output arrives — sustained throughput and time to the first token, which trade off against each other and against cost. Measured on a provider's serving of the model, so it describes an endpoint rather than the weights.

Ordered by how many of these benchmarks put a model in their top 10, then by its best placing. Each column keeps its own ranking; nothing is averaged.

Model In top 10 Artificial Analysis LLM Leaderboard Output Speed Artificial Analysis LLM Leaderboard Time To First Token
Celeris 1 celeris 2/2 #1 1523/s #9 0.58s
Gemini 2.5 Flash Lite google 2/2 #5 360/s #1 0.30s
Nemotron 3.5 Lightning nvidia 2/2 #9 305/s #8 0.56s
Mercury 2.5 inception 1/2 #2 691/s #134 3.44s
Nemotron 3 Nano Omni nvidia 1/2 #15 271/s #2 0.35s
Mercury 2 inception 1/2 #3 632/s #148 7.12s
North Mini Code cohere 1/2 #111 83/s #3 0.45s
Trinity Large Thinking arcee 1/2 #4 362/s #61 1.22s
Gemini 2.5 Flash google 1/2 #27 210/s #4 0.46s
Command A+ cohere 1/2 #57 144/s #5 0.48s
Gemini 3.5 Flash Lite google 1/2 #6 347/s #151 9.42s
Granite 4.2 ibm 1/2 #25 216/s #6 0.53s
Ling 3.0 Flash Fin inclusionai 1/2 #7 335/s #83 1.75s
Grok Build 0.1 xai 1/2 #116 80/s #7 0.54s
Ling 3.0 Flash inclusionai 1/2 #8 323/s #111 2.52s
Muse Spark 1.2 meta 1/2 #10 301/s #154 15.17s
GPT 5.4 mini openai 1/2 #14 273/s #10 0.61s

A dash means the board does not list that model — not a score of zero. Entries not yet matched to a model are left out here; they still appear on their own board.