OnlySOTA

Harnesses

A harness is the program that drives a model through a coding task: Claude Code, Codex. Each is ranked by its best run on one source, with the model behind it named, because part of a harness’s score is earned by the model it drives.

Harness Best pairing AA Cost per task
  1. 1 Claude Code Claude Sonnet 5.5 · max effort 0.684 $14.19

    All 9 runs

  2. 2 Antigravity CLI Gemini 4 Argon 0.638 $5.84

    All 1 runs

  3. 3 Codex GPT 6.1 Sol · xhigh effort 0.629 $1.04

    All 12 runs

  4. 4 Devin Fusion CLI Claude Fable 5.1 XHigh + SWE-2 Medium 0.617 $7.90

    All 2 runs

    • Claude Fable 5.1 XHigh + SWE-2 Medium KILLED 0.617 $7.90
    • GPT-6 Astra XHigh + SWE-2 Medium KILLED 0.589 $4.54
  5. 5 Grok Build Grok 4.7 · xhigh effort 0.563 $8.82

    All 2 runs

    • Grok 4.7 · xhigh effort KILLED 0.563 $8.82
    • Grok 4.6 · xhigh effort KILLED 0.470 $3.57
  6. 6 Muse Code Muse Spark 1.3 · max effort 0.543 $3.98

    All 1 runs

  7. 7 Opencode GLM 5.3 0.536 $4.24

    All 1 runs

  8. 8 Kimi Code CLI Kimi K3 0.519 $5.05

    All 1 runs

  9. 9 Muse Code 1.0.2 RC Muse Spark 1.3 · xhigh effort 0.483 $3.47

    All 1 runs

  10. 10 Antigravity SDK v0.1.12 Gemini 3.8 Flash · high effort 0.419 $2.47

    All 1 runs

Published by
Artificial Analysis
Licence
not stated
Updates
continuous
Fetched
2026-10-04
Entries
31 · 2 unmatched

The kill line, over pairings

Every harness-and-model pair the source priced, placed by cost per task and agent index. A harness charges nothing itself; the model it drives runs up most of the bill.

Artificial Analysis Coding Agent Index · starts at 00.20.40.60$0.25$1.00$5.00$14.19ExecutionerCodex · GPT 6.1 Sol (xhigh effort)0.629 · $1.04SOTALOW-COSTKILLEDCost · log10
starts at 00.20.40.60$0.50$5.00$14.19ExecutionerCodex · GPT 6.1 Sol (xhigh effort)SOTALOW-COSTKILLEDCost · log10