Speed
How fast output arrives — sustained throughput and time to the first token, which trade off against each other and against cost. Measured on a provider's serving of the model, so it describes an endpoint rather than the weights.
Ordered by how many of these benchmarks put a model in their top 10, then by its best placing. Each column keeps its own ranking; nothing is averaged.
| Model | In top 10 | Artificial Analysis LLM Leaderboard Output Speed | Artificial Analysis LLM Leaderboard Time To First Token |
|---|---|---|---|
| | 2/2 | #1 1523/s | #9 0.58s |
| | 2/2 | #5 360/s | #1 0.30s |
| | 2/2 | #9 305/s | #8 0.56s |
| | 1/2 | #2 691/s | #134 3.44s |
| | 1/2 | #15 271/s | #2 0.35s |
| | 1/2 | #3 632/s | #148 7.12s |
| | 1/2 | #111 83/s | #3 0.45s |
| | 1/2 | #4 362/s | #61 1.22s |
| | 1/2 | #27 210/s | #4 0.46s |
| | 1/2 | #57 144/s | #5 0.48s |
| | 1/2 | #6 347/s | #151 9.42s |
| | 1/2 | #25 216/s | #6 0.53s |
| | 1/2 | #7 335/s | #83 1.75s |
| | 1/2 | #116 80/s | #7 0.54s |
| | 1/2 | #8 323/s | #111 2.52s |
| | 1/2 | #10 301/s | #154 15.17s |
| | 1/2 | #14 273/s | #10 0.61s |
A dash means the board does not list that model — not a score of zero. Entries not yet matched to a model are left out here; they still appear on their own board.