Interactive Cost Calculator

Which model is cheaper for your task?

Neither the smaller model nor the frontier model always wins. It depends on how hard the task is, how much context you send, how much the model thinks, what a failure costs downstream, and the success rate you need. Pick a scenario, select model presets, or customize your own numbers.

Start from a scenario (illustrative numbers)

Pull fields out of short documents. Both models are reliable; a miss is cheap to retry.

Your task
Examples:
Smaller model
Frontier model
Smaller model: cost per successful task–Monthly: –
  • Cost per call: –
  • Model succeeds, after retries: –
  • Needs the fallback: –
Frontier model: cost per successful task–Monthly: –
  • Cost per call: –
  • Model succeeds, after retries: –
  • Needs the fallback: –

Live Pareto Plot

Visualizing cost vs. accuracy trade-off between the two models
Quality barSweet spotOn Pareto frontierDominated

Change the scenario, model presets, context size or failure cost, and the answer can flip.

From Estimation to Reality

Stop estimating. Let your evals measure the frontier.

The calculator above uses estimated success rates and token counts. In production, prompt rewrites move results 11–58× more than rerunning them, and token spend varies up to 30× across tasks.

ParetoOps replaces manual guesswork with deterministic CI analysis across all your prompts and models:

  • Reads your real eval exports: Directly ingests results from Promptfoo, LangSmith, Braintrust, DeepEval, CSV, and JSON.
  • Discovers the true Pareto frontier: Evaluates dozens of models and prompt variants at once, identifying dominated options and statistical sweet spots with 95% Wilson confidence intervals.
  • Guards your budget in CI: Exits 1 if a pull request regresses accuracy or increases cost per success, posting an explanatory diff right on the PR.

What moves the answer

Each of these can make the smaller model the right choice, or the frontier model.

Task complexity

On easy tasks, models score alike and the cheaper one wins. On hard ones, success rates drift apart, and failures dominate the cost. Different kinds of models turn out to be the most economical in different domains (Erol et al., 2025).

Prompt and context size

Long instructions, documents and history are sent on every call and multiply the price gap between models. In agent workloads, input tokens drove most of the cost (Bai et al., 2026).

How much the model thinks

Reasoning tokens are billed as output and vary widely: between models, and between runs of the same query. That is how a model with a lower list price can end up costing more (Chen, Zhang et al., 2026).

What a failure costs you

If a miss is dropped or retried cheaply, a less reliable model can be fine. If it goes to a support agent or an engineer, reliability is worth paying for.

The success rate you need

A model below your bar is out, however cheap. Above it, extra accuracy may be worth little: accuracy often plateaus while spend keeps rising (Bai et al., 2026).

Retries, volume and latency

Retries lift success but add calls and wait time, and a hard task that failed once often fails again, so treat retry gains as optimistic. At volume, fractions of a cent become real money.

How it's calculated

With first-try success rate p and up to n attempts, a task fails every attempt with probability (1 − p)n.

  • cost per call = context tokens × input price + output tokens × output price
  • model success = 1 − (1 − p)ⁿ
  • expected calls per task = model success ÷ p
  • Failures completed by a fallback: cost per successful task = cost per call × expected calls + (1 − p)ⁿ × cost per failure
  • Failures accepted: cost per successful task = cost per call ÷ p

It assumes attempts succeed independently, which flatters retries. And it needs a success rate you can trust: measured on many tasks, on your prompts. A small eval can be off by more than the gap between two models.

ParetoOps does this from your real eval results, for every model and prompt you test at once, and shows which ones are worth keeping for each job. See how →

Stop estimating. Measure it on your tasks.

ParetoOps Free computes this for every configuration in your eval results. No sign-up.