The Growing Importance of Token Observability and Engineering-Led AI Optimization
In the News
The Growing Importance of Token Observability and Engineering-Led AI Optimization
The instinct in many boardrooms is to restrict token consumption. The more useful question is not how to restrict token maxxing, but how it can be leveraged to extract revenue growth or cost efficiency.
As AI evolves at record speed, competitive advantage no longer comes from choosing the best model. It comes from skilled engineers and trusted technology partners who can turn AI into measurable business value. Capability differences between frontier models narrow with each release cycle, meaning technology choices alone no longer create transformation. Real competitive advantage now accrues to organizations that can operationalize AI by embedding intelligent capabilities directly into production systems and sustaining them at a cost that supports a credible business case.
As AI moves into production, freedom of choice must be governed by economics. This reframes tokenomics as something far more strategic than a simple cost line. It is the most direct signal an enterprise has into how its AI operates in practice: revealing how model consumption, orchestration and human oversight combine across specific workloads, features and agent outputs. Few organizations can read that signal today, however, because attributing consumption requires governance and telemetry capabilities most enterprises never built. As a result, many are confronting this gap in flight as agentic workloads drive consumption well beyond their initial assumptions. Bridging this gap requires moving past the standard pattern where an engineer simply prompts an agent, reviews what comes back and moves on. The organizations pulling ahead are undergoing a true operating-model shift: building end-to-end software factories that use defined delivery pipelines built from reusable skills, subagents and governance rules. Delivering on this shift takes deep engineering expertise, a solid vertical AI stack and an architecture capable of measuring consumption against tangible business outcomes.
How the Industry Got Here
Understanding today's token crunch requires looking back at how enterprise AI adoption evolved through 2025. For most of that year, companies were still focused on a fairly basic question: how to push AI adoption to a meaningful scale across the workforce. Adoption itself, which is getting employees to use AI meaningfully, remained the dominant conversation in enterprise AI consulting through much of the year.
That changed around the turn of 2026. New frontier model releases toward the end of 2025 marked a visible leap in capability compared to what came before. Combined with continued internal pushes for adoption, this created what can only be described as a perfect storm. Employees who had been nudged toward AI finally began to see real value, and they gravitated toward more capable, more expensive and more token-hungry frontier models. Agentic workflows compounded the effect: agents operating in loops, calling tools and reasoning through multi-step tasks consume dramatically more tokens than a simple prompt-and-response interaction ever did. Budgets planned on 2025 usage patterns were simply not built for this shift.
Why the Spend Is Spiking
The instinct in many boardrooms is to frame this as a consumption problem that needs restricting. That framing misses the point. The more useful question is not how to restrict token consumption, but how to ensure proportional value, that shows up as revenue growth or cost efficiency, is being extracted from it.
Quantifying that value has historically been difficult, and AI has made it harder rather than easier. Software engineering productivity, for instance, has never had a clean, universally accepted metric: throughput, time to market, pull request velocity and code quality all serve as imperfect proxies. Layering AI token spend on top of an already fuzzy productivity picture means many companies simply cannot correlate rising AI bills with the business outcomes those bills are meant to produce.
Several industries offer a clear illustration of why consumption is rising so fast. Engineers who once sat on capped, predictable monthly per-seat AI plans are now driving significantly higher usage, for two reasons. First, internal adoption programs are actively encouraging broader usage across teams. Second, and this is often underappreciated, a large share of today's token consumption functions as an education tax. Building agentic solutions, learning new frameworks and adapting to a technology landscape that shifts every few months all require experimentation before it can be optimized. Only once employees move past that learning curve can organizations meaningfully start working backward to improve efficiency.
The allocation of token consumption matters as much as its total volume. As agentic tools accelerate code generation, enterprises must devote sufficient AI capacity to security testing, verification and remediation. Otherwise, higher development velocity can produce an equally rapid increase in exposure.
The Security Dimension Enterprises Cannot Ignore
The surge in AI-driven activity is far more than a resource or cost management challenge; it is a critical business continuity and cyber resilience issue. As frontier AI models and autonomous agents accelerate software development, they drastically compress the window between vulnerability discovery and exploitation to machine speed. Because AI can rapidly chain together lower-risk vulnerabilities into high-impact attack paths, malicious actors can now execute automated attacks as easily as defenders can respond. To match this speed, enterprises must move beyond reactive controls and treat security as an engineering capability: embedding continuous, engineering-led risk reduction and secure AI engineering directly into modern software delivery workflows.
This raises the bar for engineering organizations considerably. Security diligence, once treated as a secondary concern behind functional requirements, now needs to be a first-class citizen baked into releases from day one. Compounding this, many enterprises are still carrying significant technical debt on outdated frameworks and libraries. This is the debt that historically lost out to revenue-generating features in prioritization battles but can no longer be deferred given the exposure it creates. Sectors handling sensitive data, particularly financial services and healthcare, face the least room for delay. The silver lining is that agentic engineering itself can be turned toward remediating this debt faster than traditional methods allow, provided organizations invest in the adoption maturity to use it well.
Building the Discipline: What Companies Need to Get Right
Tackling rising token usage responsibly requires movement on three fronts.
The first is granular visibility. Broad, aggregate figures, like "this is what we spend with Vendor A or B," are not enough. Organizations need spend data broken down to the team level, and ideally to the feature level, tracking how much is spent at each phase of development: context discovery, implementation and verification. That telemetry should also capture the human oversight wrapped around agent output, like review, correction and approval time, as the true cost of a business outcome is the combination of model consumption, orchestration and the people governing both. Without this level of telemetry, no meaningful optimization is possible.
The second is a new skill set for engineering teams. Software engineers are no longer just building software; they are increasingly responsible for building the "software factories", which are chains of agents that build software on their behalf. That requires fluency in tokenomics principles: context management, caching, context compaction and retrieval optimization, including deterministic techniques like graph-based code search that reduce reliance on costly AI-driven search. Every organization needs at least a core group of engineers who own this discipline on behalf of their teams.
The third is the ability to translate spend into quantifiable business value, which is connecting AI investment to measurable productivity gains, revenue impact or reductions in cost of goods and services. Token consumption is only one line in that equation: a credible business case also accounts for development, integration and maintenance, along with the data, security and governance capabilities required to sustain the solution. This requires a level of metrics discipline that many companies have historically lacked, AI or otherwise, and building it now is non-negotiable.
Multi-Sourcing: Treating Tokens Like a Trading Desk
A useful emerging idea in enterprise circles is that comparable model capability can often be sourced through multiple routes. Organizations that build a trading-desk-like discipline to manage this by comparing price, availability and contractual terms across vendors can gain a genuine edge on cost and margin. The point is not to treat frontier models as interchangeable commodities; it is to make sourcing decisions deliberately, use case by use case. The right model is rarely the most powerful or the cheapest per token; it is the one that delivers the business case at the required level of quality, speed, security, governance and cost.
Several techniques are already converging to make this practical. Model routing is maturing quickly across major AI providers, automatically triaging requests by complexity and directing them to the most cost-appropriate model: a feature particularly valuable for non-technical users who should not have to make model-selection decisions themselves. More technically mature engineering teams are already codifying similar logic inside their agent factories, deploying more capable models for orchestration-heavy tasks and lighter, cheaper models for simple execution and validation steps. Many organizations are also blending frontier lab models with open-weight alternatives, which cloud providers increasingly make available alongside their proprietary offerings, for tasks where the additional intelligence of a frontier model is not required.
Latency is less of a constraint than it once was. The real trade-off for enterprises now sits between cost and capability, weighed alongside the quality, speed, security and governance each use case demands. Building the internal capability to make that trade-off deliberately, task by task, is becoming a core competency for technology leadership. This is also why many CIOs (Chief Information Officers) and CTOs (Chief Technology Officers) are increasingly designing their AI ecosystems to remain vendor-agnostic, preserving the flexibility to shift providers as pricing and capability continue to evolve.
From Cloud FinOps to Tokenomics
The parallel to cloud adoption is instructive but imperfect. FinOps discipline emerged only once cloud migration reached a certain maturity, once playbooks existed and patterns were well understood. Enterprise AI is still working through its equivalent learning phase, and unlike cloud, it is unlikely to reach a stable plateau any time soon. New model releases and new agentic patterns will keep forcing organizations to relearn.
That does not mean the tokenomics discipline needs to wait. The organizations best positioned over the next year will be those that start building granular visibility, developing in-house tokenomics expertise and negotiating multi-sourcing arrangements now, rather than waiting for the industry playbooks to fully mature. The token bill is not going away. The differentiator will be whether it is spent with intent or simply absorbed as a cost of doing business.
View the original article here.
Discover how EPAM can help redefine your enterprise and outpace your competitors in the era of AI.