Introduction
ParetoOps tells you which model and prompt configuration gives you the lowest cost per successful task while still meeting your accuracy bar. It also stops changes that would quietly raise your inference bill from reaching your main branch.
The idea in one formula
Section titled “The idea in one formula”cost per successful task = total spend ÷ number of tasks that passedCost per call ignores failures, and failures are what make cheap models expensive: every failed task still costs money, plus a retry, a fallback or a person. See Cost per successful task.
What it does
Section titled “What it does”- Find the frontier.
pareto-ops analyzeranks every configuration in your eval results, removes the ones that are worse on every axis, and recommends a sweet spot. - Gate CI.
pareto-ops gatefails a CI job when a configuration falls below your accuracy threshold or goes over your cost-per-success budget. - Catch regressions (Pro, coming soon). Compare each pull request with your last known good
run, task by task, with the CI gate (
pareto-ops ci), which also comments on the pull request. - Stop losing runs early (Pro, coming soon). End an eval run once it is clearly failing your limits.
Where it fits
Section titled “Where it fits”ParetoOps doesn’t run evals. It reads the results your eval tool already produces (Promptfoo, LangSmith, Braintrust, DeepEval, CSV or JSON) and does the economics on top. It runs on your laptop or CI runner, with no telemetry: your data stays with you.