Give an enterprise real-time token visibility and a working optimisation practice, and it will still hit the same wall cloud FinOps hit a decade ago: nobody actually owns the number. IT operates the platform and receives the invoice. The business units request the features, drive the usage, and see none of the cost. That split worked, imperfectly, for infrastructure that scaled in predictable increments tied to provisioning decisions someone had to sign off on. It does not survive contact with token spend, which scales with usage patterns that no single approval gate controls.
Why the Old Split Breaks Down
The traditional IT cost ownership model assumes a provisioning gate: someone requests a server, someone approves the budget, the cost accrues to a resource with a named owner. Token spend has no equivalent gate. A single successful feature can 10x its token consumption simply by being popular, with no new approval required, because the cost is a function of how many times users invoke a capability that was approved once, at a much smaller assumed usage level. The team that owns the platform bill has no lever to control that growth, because they do not own the product decisions driving usage. The team that owns the product decisions has no visibility into the cost those decisions generate, because they have never seen a token invoice in their life.
This is not a hypothetical failure mode. It is the default outcome of applying an infrastructure-era ownership model to a cost that behaves nothing like infrastructure. The organisations that get token economics under control are the ones that recognise this early and build an ownership model suited to how token cost actually accrues, rather than forcing the cloud FinOps playbook onto a cost structure it was not designed for.
Chargeback, Adapted Rather Than Copied
Cloud FinOps solved its version of this problem with chargeback and showback models: allocating infrastructure cost back to the business unit whose workloads generated it, either as an actual internal billing mechanism or as a visibility exercise that makes the cost impossible to ignore. The same instinct applies to token spend, with an adaptation the tagging and attribution infrastructure from real-time observability makes possible: if every model call is tagged with the feature and business initiative that triggered it, the aggregated token cost by business unit is a natural output of the same instrumentation, not a separate reporting exercise.
What has to change from the cloud version is the granularity and cadence. Cloud chargeback typically runs monthly, against relatively stable infrastructure allocations. Token chargeback needs to run closer to real time, against a cost base that can move sharply within a single billing cycle, because a feature that goes viral this week generates the cost this week, and a monthly chargeback report that surfaces the number a month later gives the business unit no opportunity to respond to it while the pattern is still active.
The Accountability Structure That Actually Works
Ownership without accountability is just visibility with extra steps. The operating model that closes the gap needs three things: the business unit that owns a feature has to see its token cost close to real time, in terms it can act on; that business unit needs actual levers to influence the cost, meaning product and engineering decisions about the feature stay in the same accountable unit as the cost outcome, rather than being split across a product team and a platform team that do not share incentives; and there needs to be an escalation path when a feature’s token cost grows disproportionately to the value it produces, so that “it just kept growing” is never an acceptable explanation on its own.
This does not mean every product team needs to become a cost-optimisation team. It means the accountability for a feature’s token cost sits with the people who can actually change the feature, supported by the platform team’s tooling and expertise, rather than sitting entirely with a platform team that controls the infrastructure but not the product decisions driving demand on it.
Governance Without Bureaucracy
The risk in building any of this is over-correcting into a governance process so heavy that it slows every AI feature down with approval gates that recreate the very provisioning bottleneck cloud-native architecture was built to eliminate. The organisations doing this well keep the governance lightweight: clear ownership, visible cost, a defined escalation threshold, and genuine authority for the accountable team to act, rather than a committee that reviews every token-consuming feature before it ships. The goal is not to slow AI adoption down. It is to make sure that when adoption accelerates, the cost accelerating with it lands on someone with both the visibility and the authority to do something about it.
The Question Every Enterprise Should Be Able to Answer
The test of whether this governance model actually exists is simple: ask who owns the token budget for a specific AI feature, and see how quickly you get an answer. If the honest response takes a meeting to work out, the ownership model does not exist yet, regardless of how sophisticated the observability and optimisation practices built on top of it are. Those practices only pay off once someone with real accountability is positioned to act on what they show.
