Agentic work
Multi-step tasks drawn from real occupations rather than from software engineering — spreadsheets, documents, customer-facing procedures — run through a tool-using loop.
Ordered by how many of these benchmarks place a model in their top 10, then by its best placing. Each column keeps its own order; nothing here is averaged.
| Model | In top 10 | Artificial Analysis LLM Leaderboard τ³-Banking | Artificial Analysis LLM Leaderboard GDPval-AA | Arena — Agent Overall |
|---|---|---|---|---|
| | 3/3 | #1 51.3% | #5 61.8% | #9 0.062 |
| | 3/3 | #7 44.7% | #1 67.2% | #1 0.125 |
| | 3/3 | #6 46.0% | #8 58.9% | #2 0.104 |
| | 3/3 | #8 44.3% | #6 61.1% | #3 0.097 |
| | 2/3 | #2 50.7% | #3 63.3% | — |
| | 2/3 | #3 50.3% | #2 63.4% | — |
| | 2/3 | #4 49.1% | #7 61.0% | — |
| | 2/3 | #14 37.3% | #10 54.8% | #7 0.066 |
| | 2/3 | #9 42.1% | #17 51.2% | #10 0.062 |
| | 1/3 | #13 38.1% | #4 61.9% | — |
| | 1/3 | — | — | #4 0.095 |
| | 1/3 | #5 48.0% | #15 52.3% | — |
| | 1/3 | — | — | #5 0.085 |
| | 1/3 | — | — | #6 0.081 |
| | 1/3 | — | — | #8 0.066 |
| | 1/3 | #17 34.8% | #9 56.4% | #17 0.021 |
| | 1/3 | #10 40.2% | #13 53.8% | #15 0.032 |
A dash means the source does not cover that model, not that it scored zero. Models with no registry entry are excluded here because they cannot be joined across sources — they are still shown, flagged, on their own source's board.