Add a CI gate
pareto-ops gate checks one configuration against your limits. It exits with code 0 when the
configuration passes, 1 when it doesn’t (or the input can’t be read) and 2 for a usage error,
so any CI system can use it. If a configuration has fewer than 10 trials, the gate warns that the
result is preliminary.
npx pareto-ops gate results.json --config claude-haiku-4-5 --min-acc 0.9 --max-cost-per-success 0.01Evaluating CI Gate for Config: Claude Haiku 4.5- Success Rate: 92.0% (Threshold: 90.0%)- Cost per Success: $0.0035- Status: PARETO OPTIMAL
✅ CI Gate PASSED. Configuration meets economic & accuracy requirements.| Flag | Meaning |
|---|---|
--config <id> |
The configuration to check, by configId |
--min-acc <n> |
Minimum success rate, from 0 to 1 (default 0.9) |
--max-cost-per-success <n> |
Maximum cost per successful task, in USD |
--format <format> |
Input format (default: detected automatically) |
GitHub Actions
Section titled “GitHub Actions”Run your evals in an earlier step, then gate on the results file:
- name: Run evals run: npm run evals -- --output results.json
- name: ParetoOps gate run: npx pareto-ops gate results.json --config claude-haiku-4-5 --min-acc 0.9 --max-cost-per-success 0.01GitLab CI and other CI systems
Section titled “GitLab CI and other CI systems”The command is the same. A failing gate exits with 1, which fails the job:
evals: script: - npm run evals -- --output results.json - npx pareto-ops gate results.json --config claude-haiku-4-5 --min-acc 0.9In Pro (coming soon)
Section titled “In Pro (coming soon)”Thresholds catch a configuration that’s bad in absolute terms. Pro adds the checks teams need on every pull request:
- the CI gate (
pareto-ops ci, also available as a GitHub Actions container step), with a pull request comment and job summary showing what changed; - a regression gate against your last known good run, paired task by task, so a small drop is caught even when the result still clears your thresholds;
- baselines stored on a git branch, with no artifact storage to manage;
- early stopping, so a failing eval run ends before it spends the full budget.
Join the interest list to hear when Pro launches.