Learn

Which AI model is worth paying for?

Price per token is the number on the price page. Cost per successful task is the number on your bill. Here's why they differ, when a smaller model beats a frontier model and when it doesn't, and how to keep costs from drifting up. With research and worked examples.

What happened when researchers measured it

Independent studies that tracked real spend, not list prices, keep finding the same gaps.

32%

of model pairs where the model with the lower list price cost more in total.

Chen, Zhang et al., 2026
28×

how much more the "cheaper" model cost in the worst case.

Chen, Zhang et al., 2026
30×

difference in tokens between runs of the same agent task.

Bai et al., 2026
11–58×

more variation from rewording a prompt than from simply running it again.

Chen, Qian et al., 2026
The price list isn't the bill

A cheaper price per token can mean a bigger invoice

Providers price tokens. You pay for how many tokens a model spends to finish your task: its reasoning, its retries, its tool calls and the context it drags along. Two models of similar quality can spend very different amounts on exactly the same work.

In one pair from a 2026 study, a model listed at 80% less than its rival ended up costing 38% more across the same tasks. Mostly it spent more tokens thinking.

It cuts both ways: most of the time the cheaper model really is cheaper. The trouble is that you can't tell which case you're in from a price page. Only your own results can.

20100List price per token138100Actual cost across the tasks
Model AModel B (= 100)One model pair reported by Chen, Zhang et al., 2026. Chart is ours.
It depends on the task

Sometimes the smaller model wins. Sometimes it doesn't.

The same two models, on two different jobs. What changes is how hard the task is, how much each model thinks, and what a failure costs you. Every failed task is paid for at least twice: once for the wrong answer, again for the retry, the fallback or the person who fixes it.

Simple extraction

Pull fields out of short documents. Both models are reliable; a miss is cheap to retry.

Smaller modelFrontier model
Cost per call$0.00032$0.00900
Success, after retries99.84%99.96%
Cost per successful task$0.00041$0.00920

The smaller model is 23× cheaper per successful task. The frontier model's extra reliability buys almost nothing here.

Complex multi-step reasoning

Hard analytical tasks. The frontier model thinks at length; failures need an engineer.

Smaller modelFrontier model
Cost per call$0.00150$0.102
Success, after retries73%99%
Cost per successful task$0.552$0.142

The frontier model costs 68× more per call, yet is 3.9× cheaper per successful task, because the smaller model's failures are expensive to fix.

Illustrative numbers. Prices are typical of a small and a frontier model, not any specific provider.

The question everyone asks

"Can't I just always use the frontier model?"

You can, and you'll get strong answers. Whether it's the right call depends on the task, and the research is fairly consistent about why.

"Frontier model" here means the newest, most capable model. It's a different idea from the Pareto frontier further down.

No model is the best value everywhere

When researchers priced each correct answer, lightweight models were the most economical for basic quantitative tasks, large models for knowledge-heavy ones, and reasoning models for complex quantitative problems (Erol et al., 2025).

Many requests don't need it

Sending each request to the cheapest model likely to handle it matched the best single model at up to 98% lower cost (Chen et al., 2023). Choosing between a strong and a weak model per query cut costs by more than half in some cases, with no loss in quality (Ong et al., 2024).

Past a point, more spend buys little

In agent tasks, accuracy often peaked at a moderate level of spend and then flattened (Bai et al., 2026). Above the success rate you need, extra capability may not be worth its price.

But on hard tasks it can be the bargain

When a cheaper model fails often, and failures are expensive to fix, the more capable model can cost less per success, even at many times the price per call. That's the second example above.

So the honest answer is: it depends. On how complex the task is. On how much prompt and context you send each time. On how much the model thinks. On what a failure costs you, and how much failure you can tolerate. On the success rate you need, your volume, and how long users can wait.

None of those are on a price page. They're in your eval results, and that's where ParetoOps looks.

ParetoOps measures all of this on your own eval results

Cost per successful task for every model and prompt you test, in one command. Free, no sign-up.

More spend isn't more accuracy

Some options are simply worse, and you may be paying for them

Accuracy often climbs with spend, then flattens. Past that point, extra money buys very little (Bai et al., 2026). And when systems are tuned for accuracy alone, they drift toward complexity and cost that isn't needed (Kapoor et al., 2024).

Plot every configuration by cost per success and accuracy, and the picture gets clear. The ones nothing beats on both form the Pareto frontier. Everything else is dominated: something else is at least as accurate for less.

In this sample eval, one popular model sits off the frontier: it costs five times as much per success as an equally accurate alternative.

50%60%70%80%90%100%$0.002$0.005$0.01$0.02$0.05Cost per successful task (log scale) →Success rate ↑Your accuracy bar: 90%Gemini 2.5 FlashClaude Haiku 4.5Claude Sonnet 5Claude Opus 4.8 (legacy)
On the frontierSweet spot: cheapest frontier option above the barDominated: costs more for the same or worse results
Small edits, big swings

Would you notice if last week's prompt change made things worse?

Rewording a prompt, without changing what it asks, can move results far more than rerunning it ever would (Chen, Qian et al., 2026). And runs of the same task can differ wildly in tokens.

So comparing two averages by eye is a coin flip. The reliable way to compare two runs is task by task, with error bars on the difference (Miller, 2024), against a run you know was good.

The change on the right still passed its 50% accuracy threshold. Compared task by task with the version on main, it lost 15 points of accuracy, and each success cost 42% more.

ParetoOpscommented on this pull request❌ FAILED
🔁 Change vs Baseline (200 shared tasks) — WORSE
MetricBaselineThis changeChange (95% CI)Verdict
Accuracy81.5%66.5%-15.0pp (-22.5pp to -7.5pp)WORSE
Cost / Success$0.0155$0.0219×1.42 (×1.24–×1.64)WORSE

The change still clears the 50% accuracy threshold, so a threshold check alone would have passed it.

Excerpt of the Pro CI gate's pull request comment (full report). Real output from the sample regression dataset.

Find out what a successful task costs you

Run ParetoOps on your own eval results today, free. Get early access to Pro.