Document summarisation
Three long business documents, each summarised under a 150-word limit. A single-turn task with an unambiguous brief, so differences in the result are differences in the model rather than in the scaffolding around it.
Warning: Judges disagree materially on some models
Where the spread exceeds 0.8 on a 1-5 scale, treat the ordering as noise.
Best value
Mistral Small 3.2
4.50/5 at $0.00011 per run — $0.11 per thousand
Highest quality
GPT-4.1
4.78/5 at $0.00346 per run
Cost spread
126×
Between the dearest and cheapest model measured, for the same work
The frontier
What each point of quality costs
Every model measured on the same work. A model is on the frontier when nothing else is both better and cheaper — those are the only ones worth choosing between. The gap between the cheapest and dearest here is 126×.
This chart holds two of the three dimensions. The 3D explorer shows cost, time and quality at once, one point per run rather than one per model.
Full results
Every model, every measure
Sort by any column. Cost is what OpenRouter actually billed for these calls, retrieved per generation after the fact — not a figure computed from the published price table.
| Gemini 2.5 Progemini-2.5-pro · Google | 4.78 | $0.01 | $14.47 | 1.85 s | 2.00 s | 12.8 s | 141 tok/s | 1,332 |
| GPT-4.1★on the frontiergpt-4.1 · OpenAI | 4.78 | $0.00346 | $3.46 | 640 ms | 803 ms | 2.58 s | 150 tok/s | 215 |
| Claude Haiku 4.5claude-haiku-4.5 · Amazon Bedrock | 4.67 | $0.00220 | $2.20 | 1.18 s | 1.82 s | 4.72 s | 102 tok/s | 248 |
| Gemini 2.5 Flash★on the frontiergemini-2.5-flash · Google | 4.67 | $0.00088 | $0.88 | 361 ms | 1.16 s | 2.54 s | 218 tok/s | 243 |
| DeepSeek V3★on the frontierdeepseek-chat-v3-0324 · Crusoe, SiliconFlow | 4.56 | $0.00065 | $0.65 | 657 ms | 1.49 s | 12.5 s | 28 tok/s | 248 |
| Claude Sonnet 4.5claude-sonnet-4.5 · Amazon Bedrock | 4.50 | $0.00644 | $6.44 | 1.96 s | 2.72 s | 6.88 s | 58 tok/s | 238 |
| GPT-4.1 minigpt-4.1-mini · OpenAI | 4.50 | $0.00075 | $0.75 | 740 ms | 973 ms | 4.23 s | 87 tok/s | 253 |
| Mistral Small 3.2★on the frontiermistral-small-3.2-24b-instruct · DeepInfra | 4.50 | $0.00011 | $0.11 | 404 ms | 1.39 s | 6.51 s | 53 tok/s | 225 |
| Llama 3.3 70Bllama-3.3-70b-instruct · AkashML, DeepInfra, Groq, Novita | 4.44 | $0.00023 | $0.23 | 922 ms | 5.11 s | 12.8 s | 30 tok/s | 144 |
| Qwen3 235Bqwen3-235b-a22b · Alibaba | 4.17 | $0.00201 | $2.01 | 690 ms | 753 ms | 28.7 s | 45 tok/s | 871 |
★ marks the cost/quality frontier — models with nothing both better and cheaper. Cost is the amount OpenRouter billed, not a price-table estimate.
Speed
How long before something appears
Latency is skewed, so these are percentiles rather than averages. The distance between p50 and p95 is the tail your users will actually notice.
- Gemini 2.5 Flash361 ms / 1.16 s
- Mistral Small 3.2404 ms / 1.39 s
- GPT-4.1640 ms / 803 ms
- DeepSeek V3657 ms / 1.49 s
- Qwen3 235B690 ms / 753 ms
- GPT-4.1 mini740 ms / 973 ms
- Llama 3.3 70B922 ms / 5.11 s
- Claude Haiku 4.51.18 s / 1.82 s
- Gemini 2.5 Pro1.85 s / 2.00 s
- Claude Sonnet 4.51.96 s / 2.72 s
- Gemini 2.5 Flash218
- GPT-4.1150
- Gemini 2.5 Pro141
- Claude Haiku 4.5102
- GPT-4.1 mini87
- Claude Sonnet 4.558
- Mistral Small 3.253
- Qwen3 235B45
- Llama 3.3 70B30
- DeepSeek V328
Quality, decomposed
Where the score comes from
Each judged output is scored on 4 dimensions before an overall mark. A model can be strong on faithfulness and weak on following the brief — the overall number hides that.
| Model | concision | coverage | faithfulness | instruction adherence |
|---|---|---|---|---|
| GPT-4.1gpt-4.1 | 4.89 | 4.44 | 4.94 | 4.72 |
| Gemini 2.5 Progemini-2.5-pro | 5.00 | 4.17 | 4.83 | 5.00 |
| Gemini 2.5 Flashgemini-2.5-flash | 4.78 | 4.67 | 5.00 | 4.50 |
| Claude Haiku 4.5claude-haiku-4.5 | 4.67 | 4.39 | 5.00 | 4.50 |
| DeepSeek V3deepseek-chat-v3-0324 | 4.83 | 5.00 | 5.00 | 4.17 |
| Mistral Small 3.2mistral-small-3.2-24b-instruct | 4.50 | 4.94 | 5.00 | 4.11 |
| GPT-4.1 minigpt-4.1-mini | 4.83 | 5.00 | 5.00 | 4.00 |
| Claude Sonnet 4.5claude-sonnet-4.5 | 4.89 | 4.06 | 5.00 | 4.83 |
| Llama 3.3 70Bllama-3.3-70b-instruct | 4.89 | 3.78 | 4.94 | 5.00 |
| Qwen3 235Bqwen3-235b-a22b | 4.11 | 5.00 | 4.83 | 3.72 |
By case
Consistency across the set
The same model does not perform identically on every input. A model that is excellent on one case and poor on another is a different risk from one that is uniformly good.
| Model | Incident postmortem | Procurement policy | Research note |
|---|---|---|---|
| GPT-4.1gpt-4.1 | 5.00 | 4.50 | 4.83 |
| Gemini 2.5 Progemini-2.5-pro | 5.00 | 4.83 | 4.50 |
| Gemini 2.5 Flashgemini-2.5-flash | 5.00 | 4.00 | 5.00 |
| Claude Haiku 4.5claude-haiku-4.5 | 5.00 | 4.00 | 5.00 |
| DeepSeek V3deepseek-chat-v3-0324 | 4.67 | 4.33 | 4.67 |
| Mistral Small 3.2mistral-small-3.2-24b-instruct | 4.50 | 4.00 | 5.00 |
| GPT-4.1 minigpt-4.1-mini | 4.50 | 4.17 | 4.83 |
| Claude Sonnet 4.5claude-sonnet-4.5 | 5.00 | 3.50 | 5.00 |
| Llama 3.3 70Bllama-3.3-70b-instruct | 4.50 | 4.33 | 4.50 |
| Qwen3 235Bqwen3-235b-a22b | 3.67 | 3.83 | 5.00 |
Deterministic checks
What can be measured without a judge
These need no model to evaluate them, so they carry none of a judge's bias. Where they disagree with the judged score, trust these.
| Model | Compression | Empty | Entity recall | Has bullets | Has headings | Number recall | Over limit % | Within limit | Words |
|---|---|---|---|---|---|---|---|---|---|
| GPT-4.1 | 27% | 0% | 17% | 0% | 0% | 66% | 0.1 | 89% | 147 |
| Gemini 2.5 Pro | 25% | 0% | 24% | 0% | 0% | 47% | 0.0 | 100% | 135 |
| Gemini 2.5 Flash | 27% | 0% | 22% | 0% | 0% | 60% | 15.3 | 67% | 151 |
| Claude Haiku 4.5 | 26% | 0% | 11% | 0% | 0% | 66% | 6.0 | 67% | 144 |
| DeepSeek V3 | 28% | 0% | 13% | 0% | 0% | 68% | 4.7 | 44% | 155 |
| Mistral Small 3.2 | 26% | 0% | 9% | 0% | 0% | 43% | 4.0 | 56% | 145 |
| GPT-4.1 mini | 31% | 0% | 11% | 0% | 0% | 76% | 12.9 | 0% | 169 |
| Claude Sonnet 4.5 | 27% | 0% | 19% | 0% | 0% | 67% | 2.2 | 67% | 147 |
| Llama 3.3 70B | 19% | 0% | 2% | 0% | 0% | 36% | 0.0 | 100% | 105 |
| Qwen3 235B | 31% | 0% | 17% | 0% | 0% | 70% | 16.2 | 33% | 169 |
Rates are the share of runs passing; ratios are shown as percentages; everything else is a mean.
The judges
Who scored this, and how well they agreed
A ranking is only as good as what produced it. One judge on its own shows self-preference bias, so this is published rather than summarised away.
claude-sonnet-4.5
✓ 100% coverageScored 90 of 90 runs, averaging 4.81 out of 5 across every model.
gemini-2.5-pro
✓ 100% coverageScored 90 of 90 runs, averaging 4.30 out of 5 across every model.
| Model | claude-sonnet-4.5 | gemini-2.5-pro | Spread |
|---|---|---|---|
| GPT-4.1 | 5.00 | 4.56 | 0.44 |
| Gemini 2.5 Pro | 4.89 | 4.67 | 0.33 |
| Gemini 2.5 Flash | 5.00 | 4.33 | 0.67 |
| Claude Haiku 4.5 | 5.00 | 4.33 | 0.67 |
| DeepSeek V3 | 5.00 | 4.11 | 0.89 |
| Mistral Small 3.2 | 4.67 | 4.33 | 0.33 |
| GPT-4.1 mini | 5.00 | 4.00 | 1.00 |
| Claude Sonnet 4.5 | 4.67 | 4.33 | 0.33 |
| Llama 3.3 70B | 4.11 | 4.78 | 1.11 |
| Qwen3 235B | 4.78 | 3.56 | 1.22 |
Above roughly 0.8 of spread on a 1–5 scale, the judges are not really agreeing and the ordering should be read as noise.
Every run
Read the actual output
Nothing here is summarised away. Every model response and every judge's reasoning is published exactly as recorded, so you can decide whether you agree with the score it was given.
Loading 90 runs…
Reproducibility
How this run was configured
The grid, the parameters and the judging setup, exactly as the harness read them.
Grid
- Task
- summarise
- Runner
- raw (single streamed call)
- Models
- 10
- Cases
- 3
- Repeats
- 3
- Concurrency
- 4
- Output tokens
- 36,157
- Cost reconciled
- all runs
Parameters
- max tokens
- 2400
- temperature
- 0
- Judges
- claude-sonnet-4.5
- gemini-2.5-pro
- judge temperature
- 0
- judge max tokens
- 4000
Note: 3 cases × 3 repeats per model
A small grid. Percentiles are indicative of the tail rather than a precise measure of it, and results are a snapshot of the providers routed to on the day.
The full method, and why cost is never computed from a price table, is on the methodology page.