TL;DR — Key Takeaways
- Gartner predicts AI inference costs per agentic workflow will increase more than fivefold through 2028, despite falling token prices.
- More capable reasoning models can consume substantially more tokens and compute as agents work through multiple steps and decisions.
- Gartner recommends inference tiering, token limits and monitoring to keep agentic AI costs under control.
Falling token prices might not make up for rising inference costs as agentic workflows become more sophisticated. Gartner predicts AI inference costs per agentic workflow will increase more than fivefold through 2028, even as the underlying economics of running AI models continue to improve. The firm released new research today arguing that declining token costs are driving the use of more capable models, which in turn consume more tokens and push total inference costs higher.
In this “inference paradox” dynamic, AI agents go well beyond basic prompt-and-response interactions, repeatedly reasoning through tasks, calling tools, evaluating results and deciding next steps. Gartner said routing a task to an agentic reasoning model can cost a provider at least five times as much as a basic chatbot interaction, as each additional reasoning step and tool call increases the compute needed to finish the task.
The fivefold forecast follows a Gartner prediction from March that running inference on a one-trillion-parameter LLM will cost providers more than 90% less in 2030 than in 2025. The firm anticipates improvements in semiconductors, infrastructure, model design and chip utilization to drive those reductions. But while those efficiency gains may lower the cost of individual tokens, today’s forecast suggests that ever more sophisticated AI workloads could absorb much of those savings.
To keep costs under control, Gartner recommends optimization approaches like inference tiering, which routes different parts of a workflow to models based on the capabilities they require. Simpler tasks go to smaller, cheaper models, while more advanced work is reserved for models that cost more to run. Gartner said this type of cost control will matter for ROI, since reasoning agents will need to generate exponentially higher returns than basic models, raising the bar for workloads that can justify the added expense.
One area where Gartner expects those cost pressures to become especially visible is software development. In June, the firm predicted AI coding costs will surpass the average developer’s salary by 2028, citing rising token use and a shift from seat-based to consumption-based pricing. Gartner said that shift makes costs more variable and harder to forecast, especially when vendors do not clearly show how they calculate and bill token usage. The firm also noted that coding agents can burn through more tokens when given excessive autonomy or unnecessary context, and recommended token thresholds and monitoring to flag high-consumption workflows.
Even with those controls, lower token prices do not necessarily translate into lower operating costs for agentic applications. A more meaningful measure may be the cost of completing a workflow instead of the price of the individual tokens it consumes.
Will Sommer, a senior director analyst at Gartner, said product leaders cannot rely on cheaper tokens alone to make advanced AI cost-effective because each step up in capability can require more, and often pricier, inference. Sommer also warned that relying on general-purpose autonomous intelligence can cost far more than building AI products around carefully routed and optimized combinations of models. For enterprises, added intelligence will only pay off if the business value keeps pace with the inference bill.

