Skip to content
ManchesterAI Bench
← All benchmarks

Document summarisation

Three long business documents, each summarised under a 150-word limit. A single-turn task with an unambiguous brief, so differences in the result are differences in the model rather than in the scaffolding around it.

Single turn10 models3 cases3 repeats90 runsRun 26 August 2026

Warning: Judges disagree materially on some models

Where the spread exceeds 0.8 on a 1-5 scale, treat the ordering as noise.

Best value

Mistral Small 3.2

4.50/5 at $0.00011 per run — $0.11 per thousand

Highest quality

GPT-4.1

4.78/5 at $0.00346 per run

Cost spread

126×

Between the dearest and cheapest model measured, for the same work

The frontier

What each point of quality costs

Every model measured on the same work. A model is on the frontier when nothing else is both better and cheaper — those are the only ones worth choosing between. The gap between the cheapest and dearest here is 126×.

4.004.254.504.75$0.0001$0.0003$0.001$0.003$0.01Mean cost per run — log scaleJudged quality (1–5)GPT-4.1Gemini 2.5 ProGemini 2.5 FlashClaude Haiku 4.5DeepSeek V3Mistral Small 3.2GPT-4.1 miniClaude Sonnet 4.5Llama 3.3 70BQwen3 235B
On the frontierSomething else is better and cheaperUp and to the left is better. Hover a point for detail; every figure is in the table below.

This chart holds two of the three dimensions. The 3D explorer shows cost, time and quality at once, one point per run rather than one per model.

Full results

Every model, every measure

Sort by any column. Cost is what OpenRouter actually billed for these calls, retrieved per generation after the fact — not a figure computed from the published price table.

Every model measured, with judged quality, billed cost and latency percentiles. Sortable by any column.
Gemini 2.5 Progemini-2.5-pro · Google4.78$0.01$14.471.85 s2.00 s12.8 s141 tok/s1,332
GPT-4.1on the frontiergpt-4.1 · OpenAI4.78$0.00346$3.46640 ms803 ms2.58 s150 tok/s215
Claude Haiku 4.5claude-haiku-4.5 · Amazon Bedrock4.67$0.00220$2.201.18 s1.82 s4.72 s102 tok/s248
Gemini 2.5 Flashon the frontiergemini-2.5-flash · Google4.67$0.00088$0.88361 ms1.16 s2.54 s218 tok/s243
DeepSeek V3on the frontierdeepseek-chat-v3-0324 · Crusoe, SiliconFlow4.56$0.00065$0.65657 ms1.49 s12.5 s28 tok/s248
Claude Sonnet 4.5claude-sonnet-4.5 · Amazon Bedrock4.50$0.00644$6.441.96 s2.72 s6.88 s58 tok/s238
GPT-4.1 minigpt-4.1-mini · OpenAI4.50$0.00075$0.75740 ms973 ms4.23 s87 tok/s253
Mistral Small 3.2on the frontiermistral-small-3.2-24b-instruct · DeepInfra4.50$0.00011$0.11404 ms1.39 s6.51 s53 tok/s225
Llama 3.3 70Bllama-3.3-70b-instruct · AkashML, DeepInfra, Groq, Novita4.44$0.00023$0.23922 ms5.11 s12.8 s30 tok/s144
Qwen3 235Bqwen3-235b-a22b · Alibaba4.17$0.00201$2.01690 ms753 ms28.7 s45 tok/s871

★ marks the cost/quality frontier — models with nothing both better and cheaper. Cost is the amount OpenRouter billed, not a price-table estimate.

Speed

How long before something appears

Latency is skewed, so these are percentiles rather than averages. The distance between p50 and p95 is the tail your users will actually notice.

Time to first token (p50)p95
  • Gemini 2.5 Flash361 ms / 1.16 s
  • Mistral Small 3.2404 ms / 1.39 s
  • GPT-4.1640 ms / 803 ms
  • DeepSeek V3657 ms / 1.49 s
  • Qwen3 235B690 ms / 753 ms
  • GPT-4.1 mini740 ms / 973 ms
  • Llama 3.3 70B922 ms / 5.11 s
  • Claude Haiku 4.51.18 s / 1.82 s
  • Gemini 2.5 Pro1.85 s / 2.00 s
  • Claude Sonnet 4.51.96 s / 2.72 s
Output tokens per second (p50)
  • Gemini 2.5 Flash218
  • GPT-4.1150
  • Gemini 2.5 Pro141
  • Claude Haiku 4.5102
  • GPT-4.1 mini87
  • Claude Sonnet 4.558
  • Mistral Small 3.253
  • Qwen3 235B45
  • Llama 3.3 70B30
  • DeepSeek V328

Quality, decomposed

Where the score comes from

Each judged output is scored on 4 dimensions before an overall mark. A model can be strong on faithfulness and weak on following the brief — the overall number hides that.

score out of 5 for each model across 4 cases.
Modelconcisioncoveragefaithfulnessinstruction adherence
GPT-4.1gpt-4.1
4.89
4.44
4.94
4.72
Gemini 2.5 Progemini-2.5-pro
5.00
4.17
4.83
5.00
Gemini 2.5 Flashgemini-2.5-flash
4.78
4.67
5.00
4.50
Claude Haiku 4.5claude-haiku-4.5
4.67
4.39
5.00
4.50
DeepSeek V3deepseek-chat-v3-0324
4.83
5.00
5.00
4.17
Mistral Small 3.2mistral-small-3.2-24b-instruct
4.50
4.94
5.00
4.11
GPT-4.1 minigpt-4.1-mini
4.83
5.00
5.00
4.00
Claude Sonnet 4.5claude-sonnet-4.5
4.89
4.06
5.00
4.83
Llama 3.3 70Bllama-3.3-70b-instruct
4.89
3.78
4.94
5.00
Qwen3 235Bqwen3-235b-a22b
4.11
5.00
4.83
3.72
LowerHigherscore out of 5Each cell is the mean score that dimension received across every judged run for that model.

By case

Consistency across the set

The same model does not perform identically on every input. A model that is excellent on one case and poor on another is a different risk from one that is uniformly good.

judged quality out of 5 for each model across 3 cases.
ModelIncident postmortemProcurement policyResearch note
GPT-4.1gpt-4.1
5.00
4.50
4.83
Gemini 2.5 Progemini-2.5-pro
5.00
4.83
4.50
Gemini 2.5 Flashgemini-2.5-flash
5.00
4.00
5.00
Claude Haiku 4.5claude-haiku-4.5
5.00
4.00
5.00
DeepSeek V3deepseek-chat-v3-0324
4.67
4.33
4.67
Mistral Small 3.2mistral-small-3.2-24b-instruct
4.50
4.00
5.00
GPT-4.1 minigpt-4.1-mini
4.50
4.17
4.83
Claude Sonnet 4.5claude-sonnet-4.5
5.00
3.50
5.00
Llama 3.3 70Bllama-3.3-70b-instruct
4.50
4.33
4.50
Qwen3 235Bqwen3-235b-a22b
3.67
3.83
5.00
LowerHigherjudged quality out of 5Each cell is the mean judged quality across that model's repeats for one case.

Deterministic checks

What can be measured without a judge

These need no model to evaluate them, so they carry none of a judge's bias. Where they disagree with the judged score, trust these.

Deterministic check results for each model.
ModelCompressionEmptyEntity recallHas bulletsHas headingsNumber recallOver limit %Within limitWords
GPT-4.127%0%17%0%0%66%0.189%147
Gemini 2.5 Pro25%0%24%0%0%47%0.0100%135
Gemini 2.5 Flash27%0%22%0%0%60%15.367%151
Claude Haiku 4.526%0%11%0%0%66%6.067%144
DeepSeek V328%0%13%0%0%68%4.744%155
Mistral Small 3.226%0%9%0%0%43%4.056%145
GPT-4.1 mini31%0%11%0%0%76%12.90%169
Claude Sonnet 4.527%0%19%0%0%67%2.267%147
Llama 3.3 70B19%0%2%0%0%36%0.0100%105
Qwen3 235B31%0%17%0%0%70%16.233%169

Rates are the share of runs passing; ratios are shown as percentages; everything else is a mean.

The judges

Who scored this, and how well they agreed

A ranking is only as good as what produced it. One judge on its own shows self-preference bias, so this is published rather than summarised away.

claude-sonnet-4.5

100% coverage

Scored 90 of 90 runs, averaging 4.81 out of 5 across every model.

gemini-2.5-pro

100% coverage

Scored 90 of 90 runs, averaging 4.30 out of 5 across every model.

Mean score each judge gave each model.
Modelclaude-sonnet-4.5gemini-2.5-proSpread
GPT-4.15.004.560.44
Gemini 2.5 Pro4.894.670.33
Gemini 2.5 Flash5.004.330.67
Claude Haiku 4.55.004.330.67
DeepSeek V35.004.110.89
Mistral Small 3.24.674.330.33
GPT-4.1 mini5.004.001.00
Claude Sonnet 4.54.674.330.33
Llama 3.3 70B4.114.781.11
Qwen3 235B4.783.561.22

Above roughly 0.8 of spread on a 1–5 scale, the judges are not really agreeing and the ordering should be read as noise.

Every run

Read the actual output

Nothing here is summarised away. Every model response and every judge's reasoning is published exactly as recorded, so you can decide whether you agree with the score it was given.

Loading 90 runs…

Loading run detail…

Reproducibility

How this run was configured

The grid, the parameters and the judging setup, exactly as the harness read them.

Grid

Task
summarise
Runner
raw (single streamed call)
Models
10
Cases
3
Repeats
3
Concurrency
4
Output tokens
36,157
Cost reconciled
all runs

Parameters

max tokens
2400
temperature
0
Judges
claude-sonnet-4.5
gemini-2.5-pro
judge temperature
0
judge max tokens
4000

Note: 3 cases × 3 repeats per model

A small grid. Percentiles are indicative of the tail rather than a precise measure of it, and results are a snapshot of the providers routed to on the day.

The full method, and why cost is never computed from a price table, is on the methodology page.