Skip to content
ManchesterAI Bench

ManchesterAI Bench

What a language model actually costs you.

Published prices tell you what a token costs. They do not tell you what the work costs — reasoning tokens, cache discounts, per-request fees and retries all sit between the price table and your invoice. So we measure the invoice.

Every model here runs the same task, on the same inputs, through the same harness. Cost is retrieved from the provider after each call. Speed is measured at the stream, not the wrapper. Quality is scored blind. Then we publish all of it, including the parts that did not work.

On the tasks measured so far, the dearest model costs 126× what the cheapest does — for work a judge scores about the same.

Benchmarks published

1

Each one a different kind of work

Measured runs

90

Every one billed and reconciled

Blind judgements

180

Scored without knowing which model produced what

Benchmarks

The tasks

Each benchmark is a different shape of work, because models do not rank the same way across them. A model that summarises beautifully may be hopeless at holding a tool-calling loop together.

Every benchmark also has a 3D explorer, placing each run by cost, time and quality at once.

The full method is on the methodology page, including what these numbers cannot tell you.

Want this run against your own workload?

These are public tasks. The interesting version is the same measurement on your prompts, your traffic shape and your quality bar — which is the work ManchesterAI does.

Talk to us