Skip to content

Introduction

ParetoOps tells you which model and prompt configuration gives you the lowest cost per successful task while still meeting your accuracy bar. It also stops changes that would quietly raise your inference bill from reaching your main branch.

cost per successful task = total spend ÷ number of tasks that passed

Cost per call ignores failures, and failures are what make cheap models expensive: every failed task still costs money, plus a retry, a fallback or a person. See Cost per successful task.

  • Find the frontier. pareto-ops analyze ranks every configuration in your eval results, removes the ones that are worse on every axis, and recommends a sweet spot.
  • Gate CI. pareto-ops gate fails a CI job when a configuration falls below your accuracy threshold or goes over your cost-per-success budget.
  • Catch regressions (Pro, coming soon). Compare each pull request with your last known good run, task by task, with the CI gate (pareto-ops ci), which also comments on the pull request.
  • Stop losing runs early (Pro, coming soon). End an eval run once it is clearly failing your limits.

ParetoOps doesn’t run evals. It reads the results your eval tool already produces (Promptfoo, LangSmith, Braintrust, DeepEval, CSV or JSON) and does the economics on top. It runs on your laptop or CI runner, with no telemetry: your data stays with you.