What ParetoOps Pro posts on a pull request
The Pro CI gate (pareto-ops ci) writes one report to the job summary and keeps one comment up to date on the pull request. Both reports below are its real output on the sample data shipped with ParetoOps, rendered the way GitHub shows them. No invented numbers, nothing redacted.
A pull request that passes
Four model configurations on a 100-task benchmark, with a 90% accuracy bar. Claude Haiku 4.5 clears the bar at 85% less cost per success than Claude Sonnet 5, so the report recommends it; the legacy Opus setup is dominated.
๐ ParetoOps Gate Report โ โ PASSED
๐ Recommended Sweet Spot
Claude Haiku 4.5 (claude-haiku-4-5) delivers optimal cost efficiency at $0.0035 per successful task with 92.0% task accuracy.
๐ Pareto Frontier
๐ Model Efficiency Matrix
| Configuration | Trials | Success Rate | Avg Cost | Cost / Success | Status |
|---|---|---|---|---|---|
| Gemini 2.5 Flash | 100 | 67.0% | $0.0015 | $0.0022 | ๐ข Optimal |
| Claude Haiku 4.5 | 100 | 92.0% | $0.0032 | $0.0035 | ๐ข Optimal |
| Claude Sonnet 5 | 100 | 94.0% | $0.0219 | $0.0233 | ๐ข Optimal |
| Claude Opus 4.8 (legacy) | 100 | 80.0% | $0.0261 | $0.0327 | ๐ด Dominated |
๐ก Optimization Actions
- ๐ Sweet Spot: Claude Haiku 4.5 delivers the best cost efficiency at 92% success rate and $0.0035/successful task โ 85% cheaper than Claude Sonnet 5.
- ๐ Eliminate: Claude Opus 4.8 (legacy) is mathematically dominated by [Claude Haiku 4.5, Claude Sonnet 5]. You are paying more per successful task for lower or identical reliability.
- ๐ Cost Downgrade: Switching to Claude Haiku 4.5 saves ~85% on cost-per-successful-task with only a 2.0% delta in task success rate.
- ๐ Reliability Upgrade: Claude Opus 4.8 (legacy) achieves 80% success (below 90% SLA). Upgrading to Claude Haiku 4.5 gains +12.0% success rate and meets your SLA threshold.
Generated by ParetoOps โ cost per successful task, gated in CI.
A regression caught before merge
The change still clears its 50% accuracy bar, so a threshold check alone would pass it. Compared task by task with the saved baseline, it is significantly less accurate and more expensive per success, so the gate fails the pull request.
๐ ParetoOps Gate Report โ โ FAILED
โ Gate Violations
- Regression vs baseline sonnet-5@v1: Accuracy is significantly lower: -15.0pp (CI -22.5pp to -7.5pp), beyond the 2.0pp margin. Cost per successful task is significantly higher: ร1.42 (CI ร1.24โร1.64), beyond the +10% margin.
๐ Change vs Baseline (200 shared tasks) โ WORSE
| Metric | Baseline | This change | Change (95% CI) | Verdict |
|---|---|---|---|---|
| Accuracy | 81.5% | 66.5% | -15.0pp (-22.5pp to -7.5pp) | WORSE |
| Cost / Success | $0.0155 | $0.0219 | ร1.42 (ร1.24โร1.64) | WORSE |
๐ Model Efficiency Matrix
| Configuration | Trials | Success Rate | Avg Cost | Cost / Success | Status |
|---|---|---|---|---|---|
| Claude Sonnet 5 (prompt v1) | 200 | 66.5% | $0.0146 | $0.0219 | ๐ข Optimal |
Generated by ParetoOps โ cost per successful task, gated in CI.