Improved unit profitability causes the overall cost of AI to skyrocket without guaranteeing proportional and predictable value
Agent-based AI will not benefit from economies of scale, warns Gartner. The cost of AI inference varies considerably depending on the model and the level of reasoning and planning required by agent-based AI systems. This is the “Inference Paradox.”
A drop in token prices will not be enough. It may not offset the rise in inference costs as agent workflows become more complex. Gartner predicts that AI inference costs per agent workflow will increase more than fivefold by 2028, despite continued improvements in the underlying costs of running AI models.
Earlier this week, the firm published a new study demonstrating that lower token costs encourage the use of more powerful models… which consume more tokens and drive up total inference costs!
According to William Sommer, Quantitative Modeling and Economic Forecasting Expert at Gartner, “each new generation of AI capabilities will require more tokens, which are often more expensive. ” He also estimates that no single, reliable, and cost-effective model is on the horizon. “Developing competitive AI products will require the implementation and maintenance of complex multimodal ecosystems.”
Increasingly Sophisticated Workloads
It’s important to understand that AI agents go far beyond simple question-and-answer interactions, Gartner notes. They reason iteratively about tasks, call upon tools, evaluate results, and decide on next steps. This will inevitably cost more than an interaction with a traditional chatbot. And for good reason: each additional step in the reasoning process and each additional tool call increases the computational power required to execute the task. Gartner refers to this as the “Inference Paradox.”
While the firm anticipates improvements in semiconductors, infrastructure, model design, and chip utilization to account for these reductions—and while these efficiency gains may lower the cost of individual tokens—the current forecast suggests that increasingly sophisticated AI workloads could absorb a large portion of these savings.
Cost Control Is Essential
To control costs, William Sommer recommends optimization approaches such as inference prioritization, which involves routing different parts of a workflow to models based on their required capabilities. “Simpler tasks are assigned to smaller, less expensive models, while more complex tasks are reserved for more expensive models. ”
This type of cost control will impact return on investment, as reasoning agents will need to generate returns exponentially higher than those of base models, which will raise the bar for workloads capable of justifying the additional expenses.


