Challenges amid transformation
Artificial intelligence (AI) is transforming industries. From banking to healthcare, companies globally are investing billions to integrate AI across their operations, seeking to boost efficiency and gain insights that create competitive advantages.
Yet the results haven’t been linear, underscoring the need to prioritize spending on opportunities that can deliver the highest returns.
The following highlights the challenges companies face.
AI is seeing massive investment as illustrated below:
- In 2025, 80% of banks increased their AI budgets, with 41% of C-level executives expecting AI to exceed $25million1, 2.
- Global corporate AI investment surged to $582 billion, with generative AI alone reaching $110 billion in trailing 12-month revenue—growing at triple the pace of prior technology waves3, 4.
However, companies are struggling to immediately achieve the planned results:
- Only 13% of organizations report significant enterprise-level value from their GenAI investments.
- 40% of implemented projects fail to meet planned results5, 6.
Clearly the gap between investment and value is widening—and a significant reason lies in tokenomics, a strategic enabler for sustainable AI adoption.
What are tokens?
Tokens are the bricks of AI—small, discrete units that form the foundation of every interaction. Just as a building isn’t constructed from a single slab, language models process text by breaking it into tokens, each representing roughly three-quarters of a word.
These tokens are more than technical abstractions; they are the economic currency of AI, behaving like volatile input costs rather than predictable subscriptions7, 8.
Efficiency in token usage mirrors construction: just as architects optimize brick counts to balance cost and design, organizations must refine prompts and instructions to maximize value. The shift from possessing intelligence to applying it efficiently has turned tokens into a competitive battleground—a resource that demands strategic allocation, much like capital or talent9.
The race to the bottom—and what comes next
The AI pricing war of 2023–2025 compressed token costs dramatically. Research shows every 10% price reduction yields 12–18% more tokens consumed, with total spend still rising4. Enterprise AI budgets grew ~320% between 2024 and 202610.
But this compression was deliberate market strategy—not equivalent reductions in compute costs. We are already seeing the correction:
- A leading e-commerce player removed token leaderboards.
- A major tech firm cancelled an AI code subscription.
- Organizations questioning frontier model economics at scale11.
The dynamics are non-linear: tokens consumed per task have surged even as unit prices fell, keeping effective cost per task often flat or rising12.
The hidden cost problem
Context inflation
A single enterprise query can consume over 6,000 tokens in system prompts and retrieved documents before the user’s question is even processed—over 85% of the cost remains invisible to the end user8, 13.
Agentic multiplication
Agentic AI—systems chaining multiple autonomous steps—compounds this dramatically7. As organizations move from chatbots to autonomous workflows, consumption grows exponentially.
The pricing cliff
Frontier deployment is increasingly constrained by cost, capacity, and marginal returns. The shift is from “what models can do” to the “the price and scarcity of inputs”11. Organizations built on artificially low prices will face a significant adjustment.
What does this mean for the adoption race
Controlling token economics
Organizations should manage tokens like capital—tracking return on intelligence for each AI project and allocating spend to opportunities with the highest returns9. Token governance should match the governance of capital and revenue13.
Small models are a strategic hedge
Compact models can outperform larger frontier models on specialized tasks like mathematical reasoning. Running it on dedicated hardware costs ~$50/day for 100M tokens versus ~$1,560/day on a frontier API — a 32x difference10. The practical lever is rightsizing models to tasks rather than defaulting to the most capable option 12.
On-premises deployment is returning
Self-hosted inference offers potentially over 50% cost savings versus API approaches over three years14. For high-volume, predictable workloads, the economics are compelling—and provide greater control over cost predictability.
Conclusion
The AI adoption race is entering its second phase. The first rewarded velocity—deploying models at scale. The second will reward discipline:
- Who sustains investments without succumbing to cost inflation?
- Who demonstrates genuine ROI beyond pilot projects?
- Who navigates the non-linear economics of tokens, compute, and marginal returns?
Tokenomics is not a technical curiosity. It is the economic framework that will determine which strategies survive. Organizations that treat tokens like capital, models like strategic hedges, and governance like a competitive advantage will outlast those chasing speed alone.
The future belongs to those who measure, optimize, and govern, not just those who spend.
