OnlySOTA

HarnessTax — Terminal-Bench 2.0

Shown as HarnessTax (Arena) publishes it — their metrics, their models. A model name opens its page here; the arrow opens the original board.

Published by
HarnessTax (Arena)
Licence
not stated
Updates
frozen
Fetched
2026-10-05
Entries
21
Top 7 by Terminal-Bench 2.0
  1. GPT 5.6 Sol 83.3%
  2. GPT 5.6 Luna 76.7%
  3. Claude Fable 5 75.6%
  4. Kimi K3 73.3%
  5. Claude Opus 4.8 72.2%
  6. Claude Sonnet 4.6 65.6%
  7. Claude Haiku 4.5 47.8%

21 scored

#Model
1GPT 5.6 SolPiOpenAI83.3%71.1–94.4$0.42
2GPT 5.6 SolCodexOpenAI78.9%65.6–91.1$0.76
3GPT 5.6 LunaPiOpenAI76.7%62.2–90.0$0.05
4Claude Fable 5Claude CodeAnthropic75.6%61.1–88.9$1.55
5Kimi K3PiKimi73.3%60.0–85.6$0.38
6Claude Fable 5CodexAnthropic72.2%56.7–86.7$0.98
7Claude Opus 4.8CodexAnthropic72.2%56.7–86.7$0.85
8Claude Opus 4.8PiAnthropic72.2%58.9–84.4$0.76
9GPT 5.6 LunaCodexOpenAI72.2%57.8–85.6$0.06
10Claude Fable 5PiAnthropic71.1%54.4–85.6$1.08
11GPT 5.6 SolClaude CodeOpenAI71.1%55.6–85.6$1.35
12GPT 5.6 LunaClaude CodeOpenAI70.0%55.6–83.3$0.10
13Kimi K3CodexKimi70.0%55.6–83.3$0.45
14Claude Opus 4.8Claude CodeAnthropic68.9%52.2–84.4$0.90
15Kimi K3Claude CodeKimi66.7%52.2–80.0$0.52
16Claude Sonnet 4.6PiAnthropic65.6%48.9–82.2$0.61
17Claude Sonnet 4.6CodexAnthropic63.3%46.7–78.9$0.55
18Claude Sonnet 4.6Claude CodeAnthropic62.2%46.7–76.7$0.67
19Claude Haiku 4.5PiAnthropic47.8%32.2–63.3$0.25
20Claude Haiku 4.5Claude CodeAnthropic41.1%27.8–55.6$0.26
21Claude Haiku 4.5CodexAnthropic31.1%15.6–47.8$0.21

What each column measures

Terminal-Bench 2.0
Higher is better · Agentic coding
Cost per rollout
Lower is better