For the past few years, enterprises have ploughed ahead with an “AI everywhere” approach. Much of this was driven by performance anxiety as every organization wanted to demonstrate successful adoption that was promised to be a fast track to innovation.
Teams were given free rein to experiment with GenAI and see how it could be used to support daily workflows, from engineering to sales and marketing.
In the rush to demonstrate progress with AI and beat out competitors, success with AI was focused on proving it was indeed being adopted.
The pressure to adopt AI outpaced serious consideration of long-term ROI. In the absence of true metrics, usage soon became a proxy for value.
Today, at the midpoint of 2026, it looks like this year will go down as the year of tokens. What started as tokenmaxxing, where employees demonstrated their value to leaders through the high consumption of AI tokens, has moved on to the great token reckoning, where usage-based enterprise contracts are leaving leaders in a state of bill shock due to this rampant token consumption.
Uber has already become one of the most infamous examples of this trend, after its team of 5,000 engineers exhausted the annual development budget in a few months. What has followed in the wake of the token reckoning has been a wave of usage caps and cost control efforts.
But what will these spend caps mean for enterprise innovation? Spending caps and strategies like using cheaper AI models may reduce costs temporarily, but they don’t offer a long-term solution because they don’t get to the crux of the issue at hand.
While soaring costs associated with high token usage are being blamed, this is more than a billing issue. As 2026 is quickly becoming known as the year of “token reckoning,” enterprise leaders need to understand this isn’t about the cost of models but an AI architecture issue. Here are three things to watch out for.
The Compound Interest of AI Context Across the Enterprise
The real challenge associated with token costs is linked to how many tokens are charged for each query. AI models will use tokens for input prompts, outputs, and caching information.
However, the same AI model can use a different number of tokens to reach the same output based on the way the prompt was written. In fact, a study from Stanford Digital Economy Lab found the number of tokens used to complete the exact same task varied by as much as 30 times.
This is something developers and engineers are familiar with. When usage was the primary metric for success, they maximized their consumption of tokens to boost productivity rankings. Since the focus has flipped to conservative use, ‘caveman’ prompts have helped to slash token usage by cutting down on non-essentials like verbose language.
On an individual level, this shows that token usage can be controlled to a point. However, it’s a different story on an enterprise scale. The anxiety that led leaders to apply AI everywhere, all at once, to be seen as an organization at the forefront of innovation has led to architectural and cultural habits that are adding to the severity of the token reckoning.
Employees were encouraged to use AI, and the user-friendly interface of current AI tools meant that this went far beyond the engineering team. This means regular employees often provide extensive documentation and conversational definitions regardless of whether they’re relevant to the prompt.
In many cases, leaders have let AI parse internal knowledge banks and historical conversations from tools like Slack so the technology could provide more useful support.
The rise of agentic AI has only added to the challenge. These agents often operate around the clock to keep business moving forward, and the way that they autonomously collaborate and produce their own prompts to complete multistep workflows is another layer of token usage without clear controls.
Taken together, these factors show that a usage cap will have limited success in controlling costs without limiting AI innovation. Enterprises must look at their AI architecture and the way context is given to tools and models in use to find a more efficient path forward.
After pushing AI adoption, enterprises are now discovering that token consumption scales much faster than expected, especially with AI agents that repeatedly call models, retrieve documents, and generate long outputs. As enterprises scale AI across thousands of employees and autonomous agents, this architectural inefficiency becomes a significant operational expense.
Model Orchestration or Context Orchestration?
Understanding the ROI of AI is the first step to building an effective cost-benefit analysis for the enterprise-wide use of AI. Once the habits that drive token consumption due to inefficient context have been identified, better context frameworks need to be applied.
Here, leaders are told about the benefits of model orchestration, using more expensive frontier models for complex innovation and cheaper models for simple, everyday tasks. This strategy is a solid step towards a more efficient usage model for AI, but it doesn’t get to the crux of the issue with context.
Model orchestration needs to be combined with context orchestration to truly get ahead of the token cost crisis and build a sustainable framework for AI use at scale across the enterprise, and this comes in the form of a strategic AI architecture.
Skylar Roebuck, CTO at Solvd, argues that to counter this, enterprises ought to build composable AI architectures rather than tightly coupling applications to a single model provider.
“By separating orchestration, context management, and application logic from the underlying models, organizations can gain the flexibility to evaluate and replace models as pricing and performance evolve,” explains Skylar Roebuck, CTO at Solvd.
By building composable AI architectures rather than prioritizing a single model provider, enterprises can future-proof their ability to control spend significantly.
In 2026, data strategy is increasingly about governance and cost controls. A 2026 report from Ness Digital Engineering found that the ability of an enterprise organization to generate business value from AI is directly tied to five pillars that cover architecture modernisation, data quality and reliability, governance and ownership, treating data as a product, and security and privacy.
This report also mirrored the advice from Roebuck at Solvd, highlighting the importance of architectures that offer scalable cloud-native systems and reduce dependence on direct integrations.
Enterprise AI value increasingly comes from architecture rather than raw model capability. This is where enterprise engineering practices are becoming increasingly important.
Rooting AI in Data Economics
When AI is layered on top of existing workflows as a general-purpose productivity tool, token spend has no anchor. When it’s embedded in a specific process tied to a specific outcome, the economics change.
Cesar D’Onofrio, CEO and co-founder of Making Sense, believes that legacy modernization is now a recurring structural barrier to AI delivering real economic impact in this segment.
“AI exposes the complexity that already exists. When workflows are fragmented, ownership is unclear, and teams operate under competing priorities, advanced models will not compensate; they will tend to amplify those inefficiencies, leading AI initiatives to lose momentum,” he explained.
These legacy issues stand in the way of extracting real value from AI and make it harder to determine when to deploy frontier models or cut back on token usage. Addressing them is essential to moving beyond pilot initiatives and scaling AI without losing control of costs.
Further, the actual ROI of AI can’t be measured without a clear definition of goals. This isn’t about using the cheapest models possible, but the appropriate models for the tasks at hand. Frontier models have a higher cost, but they can be justified if the output will have a notable impact on business performance.
“AI cannot generate measurable ROI if it operates outside the core systems of record. To deliver tangible business impact, AI must be integrated into the processes that directly drive revenue, operational efficiency, and customer experience,” he added.
As Srinivasan “KG” Govindarajan, Business Head for Science & Research, Digitalized Operations at Straive, puts it, “The real measure of AI success is not cost per token. It is the cost per successful business outcome,” a shift that forces leaders to judge models by the downstream impact they unlock, not just the unit price on an invoice.
Efficiency Becomes the Next Competitive Advantage
The organizations that succeed over the next several years are unlikely to be those generating the highest token volumes. They will be the ones that build AI systems capable of delivering measurable business outcomes while minimizing unnecessary computation.
The enterprise AI reckoning isn’t signaling the end of large language models. It marks the beginning of a more disciplined era—one where AI architecture, orchestration, and operational efficiency become just as valuable as model intelligence itself.
As enterprises move beyond experimentation toward production-scale deployments, reducing token waste may prove to be one of the biggest competitive advantages in enterprise AI.

