ManchesterAI Bench
What a language model actually costs you.
Published prices tell you what a token costs. They do not tell you what the work costs — reasoning tokens, cache discounts, per-request fees and retries all sit between the price table and your invoice. So we measure the invoice.
Every model here runs the same task, on the same inputs, through the same harness. Cost is retrieved from the provider after each call. Speed is measured at the stream, not the wrapper. Quality is scored blind. Then we publish all of it, including the parts that did not work.
On the tasks measured so far, the dearest model costs 126× what the cheapest does — for work a judge scores about the same.
Benchmarks published
1
Each one a different kind of work
Measured runs
90
Every one billed and reconciled
Blind judgements
180
Scored without knowing which model produced what
Benchmarks
The tasks
Each benchmark is a different shape of work, because models do not rank the same way across them. A model that summarises beautifully may be hopeless at holding a tool-calling loop together.
Every benchmark also has a 3D explorer, placing each run by cost, time and quality at once.
The full method is on the methodology page, including what these numbers cannot tell you.
Want this run against your own workload?
These are public tasks. The interesting version is the same measurement on your prompts, your traffic shape and your quality bar — which is the work ManchesterAI does.