The vocabulary on these boards comes from three places: the field, the sources, and this site. Which is which matters, because only the third kind is ours to define.
SOTA
State of the art. On this site it is not a compliment but a position: a model is SOTA when it sits above the executioner’s cross — at least as capable as the model that kills the most others, whatever it costs.
Benchmark
One published measurement, named as its publisher names it. A benchmark belongs to whoever ran it; the same name run by two organisations is two benchmarks here, because the harness, the prompts and the grading are part of the number.
Capability
A dimension several benchmarks speak to — coding, agentic work, cost. A capability page puts the contributing benchmarks side by side without combining them, so the question it answers is “where do the sources disagree about this”, not “what is the score”.
Board
A question a visitor arrives with — what should I use for this kind of work — answered by every source that speaks to it. A board adds sections, which are the sub-questions a domain actually has: models against harnesses, graded coding against coding judged by preference.
Elo
A rating from pairwise comparisons: people are shown two answers and pick one. An Elo has no zero. It is an arbitrary anchor, not “no skill”, which is why an Elo is never drawn as a bar growing from the left edge on this site — the distance between two rows is the reading, the distance from the edge is not.
Kill line and the executioner
The executioner is the model that kills the most others, where to kill is to be at least as cheap and at least as capable. The kill line is its cross, and it splits the board into SOTA, low-cost and killed. Both axes must come from one source; borrowing one source’s prices for another’s scores would draw a line nobody measured.
The fainter line joining the models nothing beats on both axes at once is the frontier. It runs point to point, so a score read off a slope falls between two models rather than on one you can buy.
Variant
A run configuration, as distinct from a model. Reasoning effort, a harness, a
context window, a provider, a size, a release stage, a dated build — all of
these are settings on a model, not separate models. Which words count as
settings is a closed, hand-checked list, because Max is a setting on one
developer’s model and part of the name on another’s.
Checkpoint
A dated build — 20260813. A checkpoint is a rung, not a model: it never gets
an id of its own.
Line, or family
The tier a model belongs to, declared rather than parsed from its name. Claude Opus is a line and holds Opus 3 through Opus 5; GPT 5.6 carries a version and
is therefore not a line but a group over the models inside it. Deriving this
from names does not work — the longest shared prefix of two ids splits one line
by version, and a rename predates itself.
Blended price
A single price per million tokens, mixing input and output at a fixed ratio, published by the source that publishes it. It is a comparison rate, not a quote: nothing on this site is an offer, and a real bill depends on your own traffic shape.
Unmatched
An upstream row whose name matches no model in the registry yet. It is shown, flagged, and kept on its source’s own board. It is absent from cross-source views for one reason only: it cannot be joined to anything until somebody confirms what it is.
Retired
Withdrawn by the source that published it. Kept on the board and flagged, read at the best configuration you can still buy, and absent from the price chart — where there is no longer a price to pay. Retirement is per configuration: a model can have a withdrawn high-effort rung and a live one.
Estimated
Marked by the source as projected rather than measured. It is carried through and labelled, including on the kill line’s most load-bearing point.