OnlySOTA

The kill line: which AI models are worth what they cost

One model on each board beats more rivals on price and score than any other. Its cross sorts the rest into SOTA, low-cost and killed. How the line is drawn.

Every price-against-score chart of AI models asks you to do the same squint: find the dots that sit high and to the left, and quietly ignore the rest. The kill line does the squint once, the same way every time, and writes the answer on the chart.

The executioner

Start with one rule. A model kills another when it is at least as cheap and at least as capable — same price for a better score, or the same score for less. Count the kills for every model on the board, and the one with the most is the executioner. Ties go to the cheaper model, then to the stronger one.

The executioner is not the best model and not the cheapest. It is the price at which the board stops rewarding you for paying more, and on most boards it is a model few people would have named before looking.

Three zones, one cross

Draw a vertical line at the executioner’s price and a horizontal one at its score. That cross is the kill line, and every model on the board lands in exactly one of three places:

The fourth corner is empty by construction. A model cheaper and stronger than the executioner would kill everything the executioner kills, and the executioner too, so it would have been the executioner.

The home page draws this for the overall board. The coding and agent boards each draw their own, on the benchmark that use case leans on, so the model that is the best value for writing code need not be the one that is the best value overall.

What the line refuses to do

Both axes come from one source. The price and the score are published by the same organisation, measured the same way. Pairing one board’s prices with another’s scores would draw a line through points nobody measured, so a board that publishes no prices ranks models but draws no line.

It is recomputed on every refresh. The executioner can change hands on a repricing as easily as on a release: when one model’s price per million tokens quadruples, the cross moves and a different model takes the role. The line is not smoothed to look steady, and every board names the source and the fetch time it was drawn from, so a move can always be traced.

A withdrawn model never draws it. A model its source no longer lists as current is placed against the cross but cannot be the executioner, counted among its kills or tallied in a zone. An executioner nobody can buy would make nothing not worth buying.

What it does not tell you

The line reads price and one measured capability, nothing else. A killed model can still be the right choice: open weights you can run yourself, a longer context window, a licence your lawyers accept, a latency budget. What the line does say is that on this source’s measurements, for this one capability, you are paying for something other than the score.

The fainter line joining the models nothing beats on both axes at once is the frontier. It answers a different question, “at my budget, what is the best I can get”, and it runs point to point, so a score read off a slope falls between two models rather than on one you can buy.

Where to go from here