Three commands. One economic model of your AI.
ParetoOps doesn't run evals, and it never calls a model. It reads the results your eval tool already wrote, does the cost-per-successful-task math, and gives you three things: a ranked list of what's worth paying for, a CI gate that enforces it, and a way to stop a failing run before it burns its full budget.
analyze — find the frontier
Every configuration you tested, ranked by what a successful task actually cost, with the ones nothing beats on both cost and accuracy marked out.
$ npx pareto-ops analyze eval_benchmark.json --min-acc 0.9 Config / Model Success % 95% Wilson Avg Cost Cost/Success Status -------------------------------------------------------------------------------- Gemini 2.5 Flash 67.0% 57% min $0.0015 $0.0022 ✔ OPTIMAL Claude Haiku 4.5 92.0% 85% min $0.0032 $0.0035 ✔ OPTIMAL Claude Sonnet 5 94.0% 88% min $0.0219 $0.0233 ✔ OPTIMAL Claude Opus 4.8 (legac 80.0% 71% min $0.0261 $0.0327 ✖ DOMINATED 🌟 Recommended Sweet Spot: Claude Haiku 4.5 at $0.0035/successful task (92.0% accuracy)
Real engine output on the sample dataset that ships with ParetoOps.
- 95% Wilson lower bound: how low the real success rate could plausibly be, so a handful of lucky trials doesn't look like a result.
- DOMINATED means overpaying: here, the legacy Opus 4.8 setup costs about 9× more per success than Haiku 4.5 and is less accurate — money with no case for it.
- Sweet spot: the cheapest frontier configuration that clears your accuracy bar, set with
--min-acc, computed for you, not eyeballed off a chart. - Prices tokens correctly: the built-in catalog accounts for prompt caching and reasoning-token rates.
gate — stop regressions before they ship
A configuration that fails your thresholds fails the build. On Pro, every pull request is also compared with your last known good run, task by task, with the diff posted as a comment.
$ npx pareto-ops gate results.json \ --config claude-haiku-4-5 --min-acc 0.9 \ --max-cost-per-success 0.01 Evaluating CI Gate for Config: Claude Haiku 4.5 - Success Rate: 92.0% (Threshold: 90.0%) - Cost per Success: $0.0035 - Status: PARETO OPTIMAL ✅ CI Gate PASSED. Configuration meets economic & accuracy requirements.
Exit code 0 on pass, 1 on fail, 2 on usage error — any CI can use it.
- Works in any CI (Free): one shell command, no account, in GitHub Actions, GitLab, or a plain script.
- Regression gate (Pro):
pareto-ops cicompares the pull request's run with your last good run on the same tasks, so a drop that still clears your threshold doesn't slip through. - Baselines on a git branch (Pro): no artifact storage or database — the last good run is just a commit.
- Survives price changes (Pro): re-prices old and new runs from token counts, so a provider's price change isn't mistaken for a regression.
| Metric | Baseline | This change | Change (95% CI) | Verdict |
|---|---|---|---|---|
| Accuracy | 81.5% | 66.5% | -15.0pp (-22.5pp to -7.5pp) | WORSE |
| Cost / Success | $0.0155 | $0.0219 | ×1.42 (×1.24–×1.64) | WORSE |
The change still clears the 50% accuracy threshold, so a threshold check alone would have passed it.
watch — stop paying for evals that already failed
Wraps any eval runner that writes one result per line, and ends the run as soon as the result is statistically clear — not after every task has been paid for.
$ pareto-ops watch run.jsonl --min-acc 0.80 \ --total-tasks 200 -- npm run evals ⏹ STOPPED EARLY — STOP_BELOW_MIN_ACCURACY Success rate is significantly below the 80.0% minimum (65.1% over 86 trials). - Trials evaluated: 86 · success 65.1% - Spend so far: $1.2561 - Estimated spend avoided: $1.6651
Real engine output on the sample dataset that ships with ParetoOps.
- Stops on your limits: minimum accuracy, maximum cost per success, a hard
--max-spendcap, or doing worse than a baseline run. - Statistically safe by default: checks after every task while keeping the false-stop rate at or below 5% (
--alpha), and waits for at least 20 trials (--min-tasks) before it will stop at all. - Wraps any runner: point it at the command that writes your eval's results, one JSON line per task, and it tails the file.
The same analysis, from Python
pip install pareto-ops installs the identical native CLI plus a Python API, for teams that want the frontier in a notebook or a pipeline instead of a terminal.
- Free:
ParetoOptimizer, every importer, the pricing catalog. - Pro: the 3D frontier (
enable_3d_optimization) and confidence-bound ranking (use_statistical_confidence), unlocked with alicense_key.
from pareto_ops import ParetoOptimizer, load_trials, \
ParetoAnalysisOptions
trials = load_trials("promptfoo_results.json")
result = ParetoOptimizer(trials).compute(
options=ParetoAnalysisOptions(min_success_rate=0.85)
)
print(f"Sweet spot: {result.sweet_spot.config_name}")
for rec in result.recommendations:
print(f"[{rec.type}] {rec.description}")Run it on your own eval results
Free, no sign-up, no telemetry. Pro adds the regression gate, baselines and early stopping.