Solutions Who We Serve Insights & Events About Contact
Published on August 27, 2026 8 min read

AI Token Cost Management

AI token consumption costs are increasing spending out of control

Summary: AI tokens have become a material operating expense with no clear owner. 98% of FinOps teams now manage AI spend, up from 31 percent two years ago, yet business units control most of the buying, and IT sees little of it. Applying cloud governance discipline, such as visibility first, then attribution, then optimization, turns AI spend from a surprise line item into an intentional investment.

Every prompt, every response, and every automated workflow runs on tokens. As AI embeds itself deeper into daily operations, those tokens accumulate quietly into a material operating expense. Most finance and technology leaders cannot yet answer the basic questions about that expense: How many tokens does the organization consume? Who consumes them? Does the output justify the cost?

The structural reason is that AI did not enter the enterprise through procurement. Zylo’s 2026 SaaS Management Index found that business units now control 81 percent of SaaS spend while IT directly manages just 15 percent. AI-native applications are the fastest-growing spend category in that dataset, up 108 percent year over year overall and 393 percent in organizations with more than 10,000 employees, frequently arriving through employee credit cards and expense reports before IT can respond.

This is the same position most enterprises occupied with cloud compute a decade ago. Usage grew faster than the controls around it, and the correction arrived as a surprise invoice rather than as a planned decision. The organizations that moved early on cloud governance did not necessarily spend less, but spent more deliberately.

Why Are Falling Prices Not Producing Falling Bills

The most common objection to AI cost management is that model prices keep dropping. Stanford HAI’s 2025 AI Index found that inference cost for a system performing at the GPT-3.5 level fell more than 280-fold between November 2022 and October 2024.

Furthermore, total bills are still rising. Gartner forecasts $2.59 trillion in global AI spend for 2026, a 47 percent annual increase, with spending on agentic AI software alone rising 141 percent to roughly $202 billion. Agentic workflows fire many billable calls to complete a single task, so consumption grows faster than unit prices fall.

In Zylo’s survey of 218 IT leaders, 78 percent reported unexpected charges tied to consumption-based or AI pricing models in the last 12 months, and 61 percent were forced to cut projects or initiatives as a result.

Tokenomics is Now a Named Discipline

The Linux Foundation announced its intent to launch the Tokenomics Foundation in June 2026, alongside FinOps X. It was formally launched shortly after, with the stated purpose of establishing open standards, benchmarks, and best practices for the economics of AI infrastructure.

Linux Foundation leadership described tokens as the new unit of technology spend and noted there is no shared way to connect AI spending to value. FinOps Foundation leadership put it more directly, characterizing token costs and efficiency as a CEO-level concern rather than an engineering footnote.

The FinOps Foundation’s sixth annual State of FinOps survey, covering 1,192 practitioners who steward more than $83 billion in annual cloud spend, found that 98 percent now manage AI spend, up from 63 percent in 2025 and 31 percent in 2024. AI cost management is the single most requested skill set in the survey, and many teams report being asked to self-fund AI investment through optimization savings elsewhere.

The Sequencing That Most Get Wrong

Most teams start with optimization by switching to a cheaper model, trimming prompts, or capping usage. They do it before they can attribute spend to a team or a use case. Without attribution, there is no way to tell whether the savings came from genuine efficiency or from someone quietly abandoning a workflow that was creating value. The right sequence is visibility, then attribution, then optimization, with governance as the sustaining layer.

That order supports a straightforward four-phase structure:

  • Discover: inventory all AI spend, including API consumption, seat licenses, AI features embedded in existing SaaS contracts, and shadow usage on personal or team cards.
  • Analyze: map spend to teams and use cases, and compare seats provisioned against seats actually active at 30 and 90 days.
  • Optimize: apply routing, caching, seat reclamation, and contract renegotiation, sequenced by effort against savings.
  • Govern: set per-team quotas, chargeback, anomaly alerts, and a quarterly review cadence.

What the Levers Return Once Visibility Exists

Two levers have the clearest published evidence behind them. The first is model routing, which classifies each request and sends it to the cheapest model capable of handling it. RouteLLM, published at ICLR 2025 by researchers at LMSYS and UC Berkeley, achieved roughly 85 percent cost reduction on the MT-Bench benchmark while retaining 95 percent of GPT-4 performance, sending only 14 percent of queries to the frontier model. Benchmark conditions are not production conditions, so treat that figure as a ceiling rather than a forecast.

The second is license reclamation. Zylo’s index, built on more than 40 million SaaS licenses and $75 billion in spend under management, found that organizations leave an average of 36 percent of their licenses unused when measured against recommended utilization levels. Median SaaS spend per employee sits at $9,455, so unused seats compound quickly at scale.

Neither lever is safe to pull before attribution exists, because a routing change or a seat reclamation without usage data is a guess with a savings number attached.

How This Looks Like Function by Function

AI cost management is not a single team’s responsibility, and the practical work differs by function:

  • Finance: token costs are not a single expense type. Tokens spent building internal capability behave like investment; tokens supporting internal work are operating expense by function; and tokens consumed inside a customer-facing product are cost of goods sold that directly affect gross margin.
  • Procurement: contracts written for predictable seat-based software do not hold up against variable consumption, which is why unexpected charges are now the norm rather than the exception.
  • Engineering: routing and caching only produce durable savings when paired with evaluation suites, which act as guardrails against silent quality regressions that a cost dashboard will never surface.
  • Governance: shadow AI is measurable. Ramp reports that the volume of AI-related reimbursement transactions tripled year over year, driven by twice as many companies processing them. With median business, AI spend grew fourfold between February 2025 and February 2026.
  • Leadership: because a misconfigured agent can generate an outsized bill in hours, monthly invoice review is too slow a control loop for this category of spend.

Four Questions Worth Asking This Quarter

The organizations thinking ahead are already asking a short list of questions. Consider the following questions for your organization:

1. Who has access to AI tools, and what are they using them for?

2. How are token costs allocated across teams and projects?

3. Are we on the right pricing plan for our actual usage?

4. Are there redundant or low-value AI calls we can trim?

If the answers are uncertain, it indicates the organization is still in the discovery phase, regardless of how mature the underlying AI deployment looks.

Where to Start

A practical entry point is a phased roadmap on a 30-, 60-, and 90-day cadence, sequenced by effort against savings. Quick wins come first, including seat reclamation and model downgrades on low-stakes workloads. Structural changes such as routing and caching follow, and governance comes last as the layer that keeps the savings from eroding. Each item should carry an owner, an estimated saving, and an implementation effort.

Alongside the roadmap, a one-page licensing review answers the two questions leadership asks first: what are we paying for, and what is it actually being used for? That page holds a current state table covering tool, vendor, plan tier, seats purchased against seats active, annual cost, cost per active user, and renewal date. It also captures how each team uses each tool, where two tools serve the same job, and where paid seats sit idle while another team pays API costs for the same capability.

Each recommendation then resolves to one of four actions: reclaim, downgrade, consolidate, or renegotiate, with a dollar impact and a renewal-driven timeline attached.

Final Thoughts: How to Build Visibility before You Cut

Organizations facing this problem tend to want a savings number first. Our position is that the number is not trustworthy until the instrumentation exists to produce it, because optimization applied to an unattributed baseline cuts whatever is easiest to reach rather than whatever is least valuable.

Tokens are the new unit of technology spend, and the organizations that treat them with visibility, accountability, and a governance mindset will be in a considerably stronger position than those who wait for costs to spiral. Ask yourself: is your organization already tracking token usage, or is this still an emerging conversation for your team?

How we can help

Aprio’s Data & AI Solutions team can help you proactively navigate the complexities of AI token cost management. We bring together integrated capabilities to take you from your first AI question all the way to measurable, enterprise-wide outcomes and business value. Connect with us

AI token consumption costs are increasing spending out of control