Product

Three commands. One economic model of your AI.

ParetoOps doesn't run evals, and it never calls a model. It reads the results your eval tool already wrote, does the cost-per-successful-task math, and gives you three things: a ranked list of what's worth paying for, a CI gate that enforces it, and a way to stop a failing run before it burns its full budget.

Free

analyze — find the frontier

Every configuration you tested, ranked by what a successful task actually cost, with the ones nothing beats on both cost and accuracy marked out.

  • 95% Wilson lower bound: how low the real success rate could plausibly be, so a handful of lucky trials doesn't look like a result.
  • DOMINATED means overpaying: here, the legacy Opus 4.8 setup costs about 9× more per success than Haiku 4.5 and is less accurate — money with no case for it.
  • Sweet spot: the cheapest frontier configuration that clears your accuracy bar, set with --min-acc, computed for you, not eyeballed off a chart.
  • Prices tokens correctly: the built-in catalog accounts for prompt caching and reasoning-token rates.
50%60%70%80%90%100%$0.002$0.005$0.01$0.02$0.05Cost per successful task (log scale) →Success rate ↑Your accuracy bar: 90%Gemini 2.5 FlashClaude Haiku 4.5Claude Sonnet 5Claude Opus 4.8 (legacy)
On the frontierSweet spot: cheapest frontier option above the barDominated: costs more for the same or worse results
Free Pro

gate — stop regressions before they ship

A configuration that fails your thresholds fails the build. On Pro, every pull request is also compared with your last known good run, task by task, with the diff posted as a comment.

  • Works in any CI (Free): one shell command, no account, in GitHub Actions, GitLab, or a plain script.
  • Regression gate (Pro): pareto-ops ci compares the pull request's run with your last good run on the same tasks, so a drop that still clears your threshold doesn't slip through.
  • Baselines on a git branch (Pro): no artifact storage or database — the last good run is just a commit.
  • Survives price changes (Pro): re-prices old and new runs from token counts, so a provider's price change isn't mistaken for a regression.

See a full sample Pro report →

ParetoOpscommented on this pull request❌ FAILED
🔁 Change vs Baseline (200 shared tasks) — WORSE
MetricBaselineThis changeChange (95% CI)Verdict
Accuracy81.5%66.5%-15.0pp (-22.5pp to -7.5pp)WORSE
Cost / Success$0.0155$0.0219×1.42 (×1.24–×1.64)WORSE

The change still clears the 50% accuracy threshold, so a threshold check alone would have passed it.

Excerpt of the Pro CI gate's pull request comment (full report). Real output from the sample regression dataset.
Pro

watch — stop paying for evals that already failed

Wraps any eval runner that writes one result per line, and ends the run as soon as the result is statistically clear — not after every task has been paid for.

  • Stops on your limits: minimum accuracy, maximum cost per success, a hard --max-spend cap, or doing worse than a baseline run.
  • Statistically safe by default: checks after every task while keeping the false-stop rate at or below 5% (--alpha), and waits for at least 20 trials (--min-tasks) before it will stop at all.
  • Wraps any runner: point it at the command that writes your eval's results, one JSON line per task, and it tails the file.

The same analysis, from Python

pip install pareto-ops installs the identical native CLI plus a Python API, for teams that want the frontier in a notebook or a pipeline instead of a terminal.

  • Free: ParetoOptimizer, every importer, the pricing catalog.
  • Pro: the 3D frontier (enable_3d_optimization) and confidence-bound ranking (use_statistical_confidence), unlocked with a license_key.

Run it on your own eval results

Free, no sign-up, no telemetry. Pro adds the regression gate, baselines and early stopping.