Document summarisation
Three long business documents, each summarised under a 150-word limit. A single-turn task with an unambiguous brief, so differences in the result are differences in the model rather than in the scaffolding around it.
Choosing a model is a trade between three things at once, and a flat chart can only ever show you two of them. Here each measured run is a point placed by what it cost, how long it took and how well it was judged — so the shape of the trade-off is the thing you are looking at.
Drag to rotate, scroll to zoom, and click any point to read exactly what that model wrote and how each judge scored it. The cube’s edges are coloured toward the two corners that matter — cheap, fast and accurate at one, dear, slow and poor at the other — and the fixed views collapse it to any two axes at a time, on a parallel projection so a face-on reading actually lines up.
Warning: Judges disagree materially on some models
Where the spread exceeds 0.8 on a 1-5 scale, treat the ordering as noise.
Loading 90 runs…
Click a point
Every point is one run. Opening it shows what that model actually wrote, what it cost, how long it took, and how each judge scored it.
Why three axes
Cost and quality alone will tell you a cheap model is a bargain, right up until its p95 latency makes the product unusable. Time is not a footnote to the decision; it is one of its dimensions.
Why individual runs
The leaderboard shows each model as a single averaged dot. Here the same model is a cluster, and how tightly it holds together is itself a result — a wide spread means the average is hiding something.
Both scales are logarithmic
Cost and time each span orders of magnitude across these models. On a linear axis every cheap, fast run would collapse into one corner of the cube.
Prefer the numbers? The full report has every figure as a sortable table, the cost/quality frontier, judge coverage and the same runs in a searchable list.