Seville, Spain
Seville, Spain
+(34) 624 816 969
In the fast-paced world of artificial intelligence, a narrative constantly repeats: the per-token costs of large language models (LLMs) are plummeting. Providers boast of ever-increasing efficiencies, and analysts project reductions of up to 95% by 2030. Yet, companies already deploying AI agents in production are seeing their bills grow alarmingly. How is it possible that if each token is cheaper, total spending skyrockets? The answer, according to a recent Gartner analysis, lies in what they call the "inference paradox": although the unit price drops, token consumption multiplies exponentially, and the most advanced models used for complex tasks are much more expensive per token than basic chatbots.

Gartner analysts Will Sommer and Sabine Zimmerhansl have studied this dynamic in depth and reached a striking conclusion: "the pace of innovation is outpacing the cost curve." The market, they warn, is dominated by an "illusion of token deflation." Buyers dangerously assume that per-token savings will automatically transfer to their development budgets, but the reality is quite different. Modern AI agents, capable of reasoning, planning, and collaborating with each other, consume a huge number of tokens per task, and this consumption skyrockets when we talk about "swarms" of autonomous agents working in the background.
Table of contents [Show]
To understand this paradox, compare what happens with a traditional chatbot and an advanced agent. A simple chatbot receives a request, interprets it, and generates a "probably reasonable" response in seconds. In contrast, a sophisticated AI agent must think, question its own conclusions, adapt when something fails, validate the accuracy of its results without human intervention, and communicate with other agents. All this process requires much greater computational power and, above all, a number of tokens that can be 5 to 30 times higher than a chatbot for managing equivalent tasks.
Gartner has quantified this phenomenon: advanced AI agents with reasoning capabilities can cost up to 150 times more per single task than a basic chatbot. And it's not just the cost per token: training these models is also more expensive. Hardware costs for training medium-sized agent models with advanced reasoning capabilities are 2.5 times higher than training simple chatbots of similar size, and inference costs are 5 times higher. When hundreds of agents executing dozens of tasks each hour are put into production, the result is a bill that can be "mind-blowing," in the analysts' words.

To analyze the real impact, Gartner has created a "tokenomics" model based on various training and inference scenarios. They simulated 12 types of AI models with different capabilities and obtained revealing figures: basic workflows cost around $0.05 per inference token; knowledge synthesis and retrieval cost approximately $0.10; more complex workflows, such as planning and learning, reach $0.40 per token. That is, the cost per token for the provider in planning tasks is 8 to 10 times higher than basic workflows.
These numbers have direct implications for companies building AI solutions. If your agent needs to plan, reason, and learn continuously, the cost per token skyrockets. And even if provider prices drop, additional consumption can more than offset that reduction. As Sommer and Zimmerhansl warn: "costs will inevitably increase, and as they do, there is no guarantee that value will grow proportionally." The return on investment for each new generation of technology will be hard-won.
Given this scenario, companies cannot sit idly by. Gartner proposes several strategies to keep token costs in check without giving up the benefits of AI. The first is intelligent orchestration: instead of letting agents invoke cutting-edge models by default, develop systems that route each query to the most cost-effective model. "Inference tiering" can improve costs and performance, assigning simple tasks to small models and reserving large models for cases that truly require them.
Another key recommendation is to move from flat compute fees to tiered plans that adapt to real needs. Adopting usage-based pricing allows you to adjust spending to demand and avoid surprises at the end of the month. Additionally, companies should establish minimum standards for AI execution: define success thresholds from the outset, as well as risk mitigation and compliance costs. Do not greenlight implementations until scenarios have been stress-tested against fluctuations in token prices and compliance expenses.

Gartner also advises integrating value per outcome into product planning. This means requiring each AI feature to forecast and track its spending against a "clear outcome metric," such as automated tasks or successfully closed cases. This way, it is possible to identify low-performing workflows that need improvement or should be eliminated. This practice, combined with continuous model updates (treating each release like a new car that loses value from day one), can help maintain a balance between innovation and profitability.
The good news is that return on investment is "eminently possible," but it requires significant effort. Companies cannot simply rely on traditional systems. As the analysts warn: "defaulting to generic autonomous intelligence will lead to unlimited costs, an order of magnitude higher than optimized product ecosystems." In other words, the key lies in optimizing and governing AI workflows.
At ForgeNEX, we have already seen how implementing generative AI in workflows can transform operational efficiency, but we also know that without a well-defined cost strategy, the project can become unsustainable. Managing permissions and agent execution is a critical factor, and Anthropic's revenue surge demonstrates that the enterprise AI market is booming. But there are also lessons to learn from history: as our analysis on the Luddite revolts points out, technology advances faster than our ability to adapt, and that has a cost.
In summary, the inference paradox is a reminder that nothing in AI is free. Token prices drop, but consumption skyrockets. Companies that want to harness the full potential of AI agents must be aware of these costs and adopt a strategic approach. Orchestration, tiering, usage-based pricing, and ROI measurement are essential tools for navigating this new era. And as always, the key is finding the balance between innovation and economic sustainability.
Original source: ComputerWorld. Analysis and adaptation by ForgeNEX.