OnlySOTA does not evaluate models. Every number on this site was published by somebody else, and every one of them is shown beside the name of whoever published it. What this site adds is the join: the same model, under whatever name each source spells it, placed next to itself.
That sounds smaller than it is. A leaderboard is easy to read and hard to compare, because the hard part is never the ranking — it is knowing whether the row you are looking at on one board is the same thing as the row you are looking at on another.
Where the numbers come from
Each source is fetched on its own schedule and stored as a committed snapshot. Fetching and building are separate jobs: the site is rebuilt from files that are already on disk, never from a live call to anyone’s API. Two consequences follow, and both are deliberate. A deploy cannot be broken by an upstream outage. And the date beside a figure is the date it was fetched, not the date the page was rebuilt — so a page that has not changed says so honestly.
Every source’s own board is reproduced at /leaderboards, close to the shape
its authors publish it in, with a provenance bar naming the publisher, the
licence where one is stated, and how often they refresh.
What normalizing means here
Three things, and nothing beyond them.
Names are resolved to models. claude-opus-4-5, Claude Opus 4.5 and
anthropic/claude-opus-4.5 are one model. The match is exact against a
hand-curated alias list — there is no fuzzy fallback, because a fuzzy match that
is wrong is invisible, and an aggregator that quietly merges two different
models is worse than one that admits it cannot tell.
Run configuration is separated from identity. Claude Opus 5 (Max Effort)
is one model and one setting, not a second model. The same goes for a size, a
release stage, and a dated build: eleven scales of one release are one model at
eleven rungs.
A name that matches nothing is still shown. It is flagged as unmatched and left on its source’s board rather than dropped. Silent dropping is the failure mode that rots an aggregator invisibly — the board stays plausible while it quietly stops being complete.
What is never done: there is no blended score
Nothing on this site averages, weights or otherwise fuses scores across sources into a single ranking. Not on the home page, not in the structured data, not anywhere.
This is the one refusal worth stating plainly, because it is the feature every aggregator is asked for. Two boards averaged read as one board that agrees — and the agreement is manufactured. Benchmarks measure different things under different conditions: a graded examination and a count of which answer people preferred do not become more true by being added together. Where two sources disagree, the disagreement is the reading, and both are right about what they measured.
So panels sit beside each other, each keeping its own order, and the columns never merge.
The kill line
The one figure here that is derived rather than reproduced, and it is derived from a single source’s own two axes — capability and price — never from two sources borrowed from each other.
The executioner is the model that kills the most others, where killing means being at least as cheap and at least as capable. Its cross splits the board in three: everything smarter is SOTA, everything cheaper but weaker is low-cost, and everything dearer and weaker is killed. It changes hands on a repricing, not only on a release, which is why it is recomputed on every refresh.
A withdrawn model is placed against that cross but never draws it: an executioner nobody can buy makes nothing not worth buying.
What NEW means
Nothing upstream publishes an arrival date, and this site does not invent one. A NEW badge means exactly one thing: a row turned up on a board we were already watching, within the last seven days.
A source’s first snapshot is a baseline and badges nothing — otherwise the day a board is added, every model on it looks new. A snapshot that brings a fifth of a board at once is treated as a different population rather than as news. And the ledger is keyed per source by the spelling that source uses, so an adapter changing how it writes a name is not an arrival.
What RETIRED means
Sources withdraw models — one of them hides more than half its rows behind a “current” filter. Those rows are kept and flagged rather than hidden, because “what should I use” wants them gone and “how good did this get” does not.
A retired model keeps its place on a board, reads the best configuration you can still buy, and is off the price chart, where there is no price to pay.
What is not here
No re-run evaluations. No blended score. No editorial ranking of the sources against each other. No number without the name of whoever published it.