Harnesses
A harness is the program that drives a model through a coding task: Claude Code, Codex, Pi. Each is ranked by its best run on the source chosen above, with the model behind it named, because part of a harness’s score is earned by the model it drives.
Harness Best pairing HTAX-SWE Cost per task
-
1 Claude Code Claude Fable 5 97.8% $1.33
All 7 runs
- Claude Fable 5 SOTA 97.8% $1.33
- Claude Opus 4.8 KILLED 86.7% $0.98
- GPT 5.6 Sol KILLED 77.8% $1.54
- Kimi K3 KILLED 76.7% $0.78
- Claude Sonnet 4.6 KILLED 66.7% $0.67
- GPT 5.6 Luna LOW-COST 55.6% $0.15
- Claude Haiku 4.5 LOW-COST 52.2% $0.43
-
2 Codex Claude Fable 5 96.7% $0.89
All 7 runs
- Claude Fable 5 KILLED 96.7% $0.89
- Claude Opus 4.8 KILLED 88.9% $0.69
- Kimi K3 KILLED 74.4% $0.85
- GPT 5.6 Sol LOW-COST 73.3% $0.56
- Claude Sonnet 4.6 KILLED 68.9% $0.74
- Claude Haiku 4.5 LOW-COST 57.8% $0.39
- GPT 5.6 Luna LOW-COST 55.6% $0.04
-
3 Pi Claude Fable 5 96.7% $0.67
All 7 runs
- Claude Fable 5 Executioner 96.7% $0.67
- Claude Opus 4.8 LOW-COST 82.2% $0.47
- GPT 5.6 Sol LOW-COST 74.4% $0.44
- Kimi K3 LOW-COST 72.2% $0.46
- Claude Sonnet 4.6 KILLED 64.4% $0.68
- Claude Haiku 4.5 LOW-COST 60.0% $0.37
- GPT 5.6 Luna LOW-COST 53.3% $0.03
- Published by
- HarnessTax (Arena)
- Licence
- not stated
- Updates
- frozen
- Fetched
- 2026-10-05
- Entries
- 21
The kill line, over pairings
Every harness-and-model pair the source priced, placed by cost per task and the source’s score. A harness charges nothing itself; the model it drives runs up most of the bill.