Token Economics: The New Enterprise AI Cost Frontier

T

Token cost is about to become the AI conversation every enterprise leadership team has, whether they choose to have it deliberately or have it forced on them by a quarterly invoice nobody can explain. Cloud spend went through this exact reckoning a decade ago: infrastructure that scaled faster than anyone’s ability to see what was driving it, followed by a wave of FinOps discipline built to catch up. Token spend is heading down the same road, faster, and most enterprises have neither the visibility nor the operating model in place to meet it.

The pattern is already visible in the field. AI programmes that started as a handful of pilots are now embedded in core workflows, and the token consumption behind them has grown in step, but almost invisibly. It sits inside application code, inside agent orchestration, inside third-party SaaS features nobody audited for their model calls. Finance sees a monthly total. Nobody can yet answer what drove it, which team owns it, or whether it produced anything worth the spend. That gap is not a temporary reporting lag. It is a structural gap, and it is about to matter a great deal more than it does today.

Two Gaps, Not One

The first gap is visibility. Most organisations cannot tell you, in anything close to real time, what they are spending in tokens this week, on which workload, on whose initiative. What exists instead is a vendor invoice that arrives weeks after the fact, aggregated to a number too coarse to act on. Business and IT leadership are making decisions about AI investment with the cost equivalent of a rear-view mirror, and the mirror updates monthly.

The second gap is optimisation, and it only becomes visible once the first gap closes. Visibility tells you what you are spending. It does not by itself tell you how to spend less without losing the capability you are paying for. Right now, few enterprises have a deliberate practice for this: choosing the right model for the task rather than defaulting to the largest one, designing prompts and context windows deliberately rather than accumulating them by accident, building caching and retrieval architectures that avoid paying twice for the same reasoning. Optimisation without visibility is guesswork. Visibility without optimisation is an expensive dashboard.

Why This Is Not Just FinOps With a New Line Item

It is tempting to assume the cloud FinOps playbook simply extends to cover this, and it does not, cleanly. Cloud spend is provisioned: a team requests infrastructure, someone approves it, and the cost accrues to a resource with an owner and a tag. Token spend is behavioural. It accrues call by call, inside code paths that were never designed with cost as a first-class concern, shaped by how a prompt is written, how much context an application decides to send, how many times an agent decides to retry or re-plan before it produces an answer. Two applications doing the same job can differ in token cost by an order of magnitude based purely on implementation choices nobody thought of as financial decisions at the time they were made.

That difference matters because it means the organisations that built strong cloud FinOps practices do not automatically inherit an equivalent capability for AI. The metering has to be rebuilt for a cost that is generated inside application logic rather than inside a provisioning system. The ownership model has to answer a harder question than “which team’s infrastructure is this,” because the token spend of a single feature can be driven by product decisions, prompt engineering decisions, and architecture decisions made by three different people who never spoke to each other about cost.

The Cost of Waiting

The honest reason to take this seriously now, rather than when the first uncomfortable board question arrives, is that the problem compounds. Every new AI feature shipped without cost instrumentation is a feature whose token behaviour nobody can later explain. Every agentic workflow deployed without a metering and optimisation discipline underneath it is a workload that multiplies its own token consumption as it grows more capable, because agents that plan, retry, and call tools do not spend tokens the way a single request-response call does. The organisations that build the visibility and optimisation muscle now are building it while the spend is still small enough to be a design decision. The organisations that wait are building it under audit pressure, against a baseline nobody can reconstruct.

The Conversation Technology Leaders Should Start Now

None of this requires an enterprise to slow down its AI investment. It requires treating token cost as an architecture and governance discipline from the outset, with the same seriousness that cloud cost earned after the first decade of uncontrolled growth taught the lesson the hard way. The technology leaders who get ahead of this will be the ones who can walk into a board meeting with a real answer to what the AI programme costs and what it returns, instead of an invoice and a shrug. The ones who wait will be explaining, after the fact, why nobody was watching.

About the author

Martijn Baecke

Add Comment

Martijn Baecke

About this blog

I am a technologist and strategic advisor specializing in multi-cloud architectures, security, AI integration, and modern IT operations.

This website is a dedicated space for sharing my knowledge, where I focus on translating complex engineering challenges into clear, actionable strategies that drive real-world business outcomes.

Disclaimer: All content and technical expertise are my own; AI is used solely for structural editing and formatting.