Research

What researchers found when they measured real AI costs

A growing body of work measures what models actually cost per correct answer, and how to tell whether a change really helped. Here is what it found, in plain language, and what it means if you run AI in production.

These are summaries in our own words, based on each paper's public abstract. The research belongs to its authors, and ParetoOps is not affiliated with them. Follow the links for the full papers and their methods.

The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More

Chen, Zhang, He, Stoica, Zaharia and Zou · 2026 · arXiv
  • Across 8 reasoning models and 12 tasks, the model with the lower list price cost more in total in 32% of model pairs, by as much as 28 times.
  • In one pair, a model listed 80% cheaper than the other ended up 38% more expensive across the tasks.
  • The main driver was how many thinking tokens each model spent. Running the same query again could change its thinking-token count by up to 9.7 times.

What it means in practice: Price per token does not predict what a task will cost you. The only reliable number is measured spend on your own tasks.

How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks

Bai, Huang, Wang, Sun, Mihalcea, Brynjolfsson, Pentland and Pei · 2026 · arXiv
  • Runs of the same agentic coding task varied by up to 30 times in total tokens.
  • Input tokens, not output tokens, drove most of the cost.
  • Accuracy often peaked at a moderate level of spend and then flattened out as spend kept rising.
  • Models were poor at predicting their own token usage, with correlations of 0.39 at best.

What it means in practice: One run tells you little about cost: measure across repeated trials. And past a certain point, paying more stops buying accuracy.

Cost-of-Pass: An Economic Framework for Evaluating Language Models

Erol, El, Suzgun, Yuksekgonul and Zou · 2025 · arXiv
  • Defines cost-of-pass: the expected money spent to get one correct solution. It is the same idea ParetoOps calls cost per successful task.
  • Different kinds of models were the most economical choice in different domains: lightweight models for basic quantitative tasks, large models for knowledge-heavy tasks, reasoning models for complex quantitative ones.
  • For complex quantitative tasks, the cost of a correct answer roughly halved every few months over the year studied.

What it means in practice: The best model depends on your workload, and the answer goes stale quickly. It is worth re-measuring whenever models or prices change.

AI Agents That Matter

Kapoor, Stroebl, Siegel, Nadgir and Narayanan · 2024 · arXiv
  • Benchmarks that ranked agents on accuracy alone encouraged agents that were needlessly complex and expensive.
  • Optimizing cost and accuracy together cut cost substantially while keeping accuracy.

What it means in practice: Judge configurations on cost and accuracy at the same time, and look for the ones nothing else beats on both.

FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance

Chen, Zaharia and Zou · 2023 · arXiv
  • Sending each query to the cheapest model likely to handle it, and escalating only when needed, matched the best single model of the time at up to 98% lower cost.
  • Spending the same budget that way instead raised accuracy by 4% over that best single model.

What it means in practice: The strongest model is not automatically the right one for every request. A mix chosen per task can beat any single model on cost, accuracy or both.

RouteLLM: Learning to Route LLMs with Preference Data

Ong, Almahairi, Wu, Chiang, Wu, Gonzalez, Kadous and Stoica · 2024 · arXiv
  • Routing each query to either a stronger or a weaker model cut costs by more than half in some cases, without lowering response quality.
  • The routers kept working when the underlying models were swapped.

What it means in practice: Many requests do not need the most capable model. Knowing which ones do is where the savings are.

Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations

Miller · 2024 · arXiv
  • Treats an eval as an experiment on a sample of possible questions, so every score comes with uncertainty.
  • Gives methods for measuring the difference between two models, and for planning evals with enough statistical power.

What it means in practice: Comparing two eval scores by eye is not enough to say a change helped or hurt. The difference needs an error bar.

Noise Floor Audit for Agent Benchmarks

Chen, Qian, Wang, Peng, Xu, Wu and Sun · 2026 · arXiv
  • Rewording a prompt without changing its meaning moved results 11 to 58 times more than simply rerunning the same prompt (median paired standard deviation).

What it means in practice: Small prompt edits can swing results far more than run-to-run noise suggests. Regressions hide easily unless you compare task by task.

Efficient Benchmarking of AI Agents

Ndzomga · 2026 · arXiv
  • Keeping only tasks with a middling historical pass rate (30 to 70%) cut the number of eval tasks by 44 to 70% while keeping model rankings largely intact.

What it means in practice: A large part of an eval budget can go to tasks that do not tell models apart.

Taken together

  • Measure cost per correct answer, not price per token. List prices don't predict what a task costs.
  • Measure on your own tasks, and re-measure. The best choice depends on the workload and changes as models and prices move.
  • No single model is the right answer everywhere. The most economical choice depends on the task, and many requests don't need the most capable model, while hard ones may.
  • Look at cost and accuracy together. Options that are worse on both axes are common, and easy to miss when you look at one number.
  • Compare runs task by task, with error bars. Prompt edits and run-to-run variation make eyeballed comparisons unreliable.

ParetoOps applies these ideas to your own eval results. Read the docs ortry the calculator.