Agentic work
Multi-step tasks drawn from real occupations rather than from software engineering — spreadsheets, documents, customer-facing procedures — run through a tool-using loop.
Ordered by how many of these benchmarks put a model in their top 10, then by its best placing. Each column keeps its own ranking; nothing is averaged.
| Model | In top 10 | Artificial Analysis LLM Leaderboard τ³-Banking | Artificial Analysis LLM Leaderboard GDPval-AA | Arena — Agent Overall |
|---|---|---|---|---|
| | 3/3 | #6 47.2% | #3 61.7% | #1 0.143 |
| | 2/3 | #1 51.3% | #8 58.6% | #22 0.025 |
| | 2/3 | — | #1 67.3% | #2 0.138 |
| | 2/3 | — | #2 67.0% | #3 0.125 |
| | 2/3 | #3 50.5% | #7 58.6% | #17 0.040 |
| | 2/3 | #4 50.3% | #9 57.8% | #23 0.024 |
| | 2/3 | #11 44.7% | #5 60.4% | #7 0.087 |
| | 2/3 | #7 47.2% | #10 57.3% | #31 -0.004 |
| | 1/3 | #2 50.7% | #11 56.6% | #25 0.013 |
| | 1/3 | #13 43.1% | #20 52.1% | #4 0.123 |
| | 1/3 | — | #4 60.5% | #15 0.040 |
| | 1/3 | #5 48.0% | #32 46.2% | #36 -0.017 |
| | 1/3 | — | #18 53.8% | #5 0.112 |
| | 1/3 | — | #6 59.3% | #20 0.033 |
| | 1/3 | — | #23 50.3% | #6 0.097 |
| | 1/3 | #8 46.0% | #21 52.0% | #13 0.042 |
| | 1/3 | #22 38.1% | #16 54.8% | #8 0.082 |
| | 1/3 | #9 45.8% | #33 45.6% | #21 0.030 |
| | 1/3 | — | #13 55.6% | #9 0.076 |
| | 1/3 | #10 45.4% | #12 55.6% | #33 -0.007 |
| | 1/3 | #31 34.2% | #30 46.9% | #10 0.066 |
A dash means the board does not list that model — not a score of zero. Entries not yet matched to a model are left out here; they still appear on their own board.