Claude Sonnet 5.5 vs
Step 5
Both are SOTA right now. Each row below is one board’s call on the two, and the marked figure is the one that board places higher. No row is added to another, so there is no overall winner: the boards measure different things, and where they split, that is the finding.
| Board | Claude Sonnet 5.5 | Step 5 |
|---|---|---|
| Artificial Analysis LLM Leaderboard Artificial Analysis Intelligence Index | 56.0 No. 2 of 372 | 43.7 No. 17 of 372 |
| Arena — Agent Overall | 0.125 No. 3 of 49 | 0.005 No. 31 of 49 |
| Artificial Analysis Coding Agents Artificial Analysis Coding Agent Index | 0.684 No. 1 of 22 | — |
| Arena — Code Overall | 1786 No. 3 of 109 | 1570 No. 33 of 109 |
| Arena — Text Overall (style control) | 1471 No. 45 of 306 | 1454 No. 74 of 306 |
| Arena — Vision Overall (style control) | 1268 No. 35 of 128 | 1260 No. 45 of 128 |
Each rank is the board’s own: the rank it publishes, or its place in its own order where it publishes none. The bar is how much of that board the model is ahead of. Scores from different boards cannot be compared.
At a glance
| Developer | Anthropic | StepFun |
|---|---|---|
| Weights | closed | closed |
| Released | — | — |
| Price Blended Price, USD per 1M tokens (3:1 input:output) | $4.00 | $1.43 |
| Value: where it sits on the kill line | SOTA | SOTA |
Claude Sonnet 5.5 →Step 5 → How the kill line is drawn →
Claude Sonnet 5.5 against the other SOTA models
- Qwen 3.8 Flash Next
- Qwen 3.8 Max
- Claude Fable 5
- Claude Fable 5.1
- Claude Opus 5.5
- DeepSeek V4.1 Flash
- Gemini 3.8 Flash
- Gemini 4 Argon
- Kimi K3
- Muse Spark 1.3
- GPT 5.6 Terra
- GPT 6.1 Sol
- GPT 6 Astra
- GPT 6 Luna
- Grok 4.7
- MiMo V2.6 Flash
- MiMo V2.6 Pro
- GLM 5.3
- GLM 5.3 Flash