Know which AI model is worth paying for.
ParetoOps turns the eval results you already have into a clear answer: the model and prompt with the lowest cost per successful task that still meets your quality bar. Then it guards that answer in CI, so no change quietly makes you pay more for worse results.
Free plan · No sign-up · Runs offline, your data stays with you
$ npx pareto-ops analyze results.json --min-acc 0.9 Model Success $/success Status ---------------------------------------------------------- Gemini 2.5 Flash 67.0% $0.0022 ✔ OPTIMAL Claude Haiku 4.5 92.0% $0.0035 ✔ OPTIMAL Claude Sonnet 5 94.0% $0.0233 ✔ OPTIMAL Claude Opus 4.8 (legacy) 80.0% $0.0327 ✖ DOMINATED 🌟 Sweet spot: Claude Haiku 4.5, $0.0035/success, 85% cheaper than Claude Sonnet 5 for 2 points of accuracy $ npx pareto-ops gate results.json \ --config claude-haiku-4-5 --min-acc 0.9 ✅ CI Gate PASSED.
Three ways AI budgets leak, and how ParetoOps plugs them
Choosing models by price per token
In 32% of model pairs studied, the model with the lower list price cost more overall (Chen, Zhang et al., 2026).
ParetoOps ranks every option by what a successful task really costs on your data.
Prompt changes that ship blind
Rewording a prompt moved results 11 to 58 times more than rerunning it (Chen, Qian et al., 2026). Comparing two averages can't catch that.
ParetoOps compares every pull request with your last good run, task by task.
Eval runs that keep spending after they've failed
In our sample, a run clearly below its accuracy bar was stopped after 86 of 200 tasks, saving 57% of its spend.
ParetoOps stops a failing run as soon as the evidence is clear.
Find the cheapest option that's good enough
Every model and prompt you test, ranked by cost per successful task. No spreadsheets, no guesswork.
- Flags dominated options: beaten on both cost and accuracy, so you can drop them today.
- Recommends a sweet spot: the cheapest option above your accuracy bar.
- Prices tokens correctly: prompt caching and reasoning tokens included.
Catch regressions before they merge
Every prompt or model change is checked in CI before it reaches your users, or your bill.
- Gate on your limits (Free): fail the build when accuracy drops below your bar or cost per success goes over budget, in any CI.
- Compare with your last good run (Pro): paired task by task, with confidence intervals, so real regressions aren't lost in noise.
- See it on the pull request (Pro): one comment on the pull request shows what changed and why it failed, from the CI gate's GitHub Actions step or any CI.
| Metric | Baseline | This change | Change (95% CI) | Verdict |
|---|---|---|---|---|
| Accuracy | 81.5% | 66.5% | -15.0pp (-22.5pp to -7.5pp) | WORSE |
| Cost / Success | $0.0155 | $0.0219 | ×1.42 (×1.24–×1.64) | WORSE |
The change still clears the 50% accuracy threshold, so a threshold check alone would have passed it.
Stop paying for evals that already failed
ParetoOps watches your eval as it runs and ends it once the result is clear, instead of paying for every remaining task.
- Stops on your limits: minimum accuracy, maximum cost per success, a hard spend cap, or doing worse than your baseline.
- Statistically safe: checks after every task while keeping false stops below a rate you choose.
- Wraps any eval runner that writes one result per line.
$ pareto-ops watch run.jsonl --min-acc 0.80 \ --total-tasks 200 -- npm run evals ⏹ STOPPED EARLY — STOP_BELOW_MIN_ACCURACY Success rate is significantly below the 80.0% minimum (65.1% over 86 trials). - Trials evaluated: 86 · success 65.1% - Spend so far: $1.2561 - Estimated spend avoided: $1.6651
Real engine output on the sample dataset that ships with ParetoOps.
The right answer changes with the task. ParetoOps tells you which.
Same two models, two jobs, opposite winners. It depends on task complexity, context size, how much the model thinks and what a failure costs you.
Simple extraction
The smaller model wins.
Complex multi-step reasoning
The frontier model wins, despite a much higher price per call.
Illustrative scenarios.
Built different, on purpose
Reads the results your eval tool already exports. ParetoOps never calls a model and sends no telemetry.
The gate is arithmetic on your data, not another model's judgment. Same results in, same verdict out, every time.
Runs on your laptop or CI runner. License keys are checked offline; nothing phones home.
Every stat on this site traces to a cited study or the sample dataset shipped with the engine. See the research →
One command that exits 0 or 1. GitHub Actions, GitLab, Jenkins, or a cron job — same gate everywhere.
Python SDK
The same analysis from Python: pip install pareto-ops.
Baselines without infrastructure Pro
Last known good results live on a git branch: no artifact storage or database.
Survives price changes Pro
Re-prices old and new runs from token counts, so a provider's price change isn't mistaken for a regression.
See every command in depth on the product page →
Start free today. Add Pro when your team ships every week.
Free
- Cost per successful task for every option
- Dominated options and sweet spot
- CI gate on accuracy and cost limits
- All eval formats, Python SDK
Pro
- Everything in Free
- CI gate with pull request comments
- Regression gate vs your last good run
- Early stopping for failing evals
- Baselines on a git branch, latency as a third axis
Questions
What is cost per successful task?
It is the total you spent on a set of tasks divided by the number of tasks that actually succeeded. A failed call still costs money, so a model that is cheap per call but fails often can cost more per useful result than a pricier, more reliable one. Researchers call the same quantity cost-of-pass.
Should I always use the most capable model?
Not necessarily. Research that priced each correct answer found that different kinds of models are the most economical for different tasks, and that sending each request to the cheapest model able to handle it can match the best single model at a fraction of the cost. On hard tasks where failures are expensive, the most capable model can still be the cheapest per success. The right choice depends on task complexity, prompt and context size, how much the model reasons, what a failure costs you and the success rate you need, so it has to be measured on your own tasks.
Does ParetoOps run my evals or call model APIs?
No. It reads the results your eval tool already produces and does the analysis on top. It never calls a model provider.
Which eval tools does it work with?
Promptfoo, LangSmith, Braintrust and DeepEval exports, plus CSV and a simple JSON format. Each trial needs a configuration name, its cost and whether it passed. Latency is optional.
Does my data leave my machine?
No. ParetoOps runs on your laptop or CI runner, sends no telemetry, and checks license keys offline.
What does Pro add?
The checks a team needs on every pull request: a CI gate that comments on the pull request (a GitHub Actions step, or any CI), a regression gate that compares each change with your last known good run task by task, baselines stored on a git branch, early stopping for eval runs that are already failing, and a 3D frontier that includes P95 latency.
See what your AI really costs per result
Run ParetoOps on your eval results in a minute. Free, no sign-up.