Capitalize or Expense
A new industry standards body just conceded that the number every enterprise is governing is the wrong one. The organizations that cannot separate what their tokens build from what their tokens burn will simply receive a more precise account of a position they did not choose.
The Linux Foundation has launched the Tokenomics Foundation, backed by thirty founding members, with a mandate to build vendor-neutral standards for the cost and value of AI. The roadmap covers value frameworks, full-cost reference models, and token cost telemetry inside the FOCUS specification.
The most consequential item on that list is none of those things. It is a unit change. The Foundation intends to standardize cost to serve as cost per call rather than cost per token, so that the number maps to work actually performed.
That is not an accounting refinement. That is an industry conceding, in public, that the metric every enterprise has been governing does not describe the thing it is trying to control.
July's issue argued that token capital is a balance sheet asset, and that the question is not what AI costs but what the spending builds. This month takes the harder half of that argument, because the balance sheet framing only works if an organization can tell which half of its bill is accruing and which half is evaporating. Almost none can.
The Bill You Can Read Is Not the Bill That Breaks You.
Tokens are the most easily metered layer of AI spend. They are not the largest one, and they are rarely the one that ends the fiscal year badly.
A linear chatbot exchange and an orchestrated agentic workflow can resolve the same customer question. The second costs vastly more, and almost none of the difference is token price. It is structure: tool calls, reasoning loops, context revisited and reloaded, retries that nobody counted, subagents invoked by other subagents. EY's decomposition of a single production agent runs to seven distinct cost categories, and observes that most business cases capture only the first three.
Which means the CFO who asks what one agent costs receives a vendor invoice describing a fraction of the exposure, and reasonably treats it as the whole.
The Numbers Behind the Overrun.
Those four numbers describe one failure, not four. Cost escalates because the architecture generates it. Value is unclear because nothing is metered at the workflow. Controls are inadequate because the people who could impose them do not sit where the spend is created.
The 8% is the load-bearing figure. The discipline that governs AI spend sits almost entirely inside technology. Until that changes, the unit economics of every AI feature an enterprise ships will be set by people who are not accountable for the P&L, and a 2.4 times overrun remains the expected outcome rather than the exception.
Provisioned Spend Behaves. Generated Spend Does Not.
Mature cloud FinOps teams forecast within one to three percent of actual. Those same teams, applied to AI, miss by a factor of two to three. The discipline did not get worse. The spend changed shape.
Cloud spend is provisioned. A human commits capacity, and the bill follows the commitment. AI spend is generated. The workload writes its own bill at runtime, and a single request can trigger an unbounded chain of model calls as an agent reasons, invokes tools, and reloads context it already had. Every architectural decision is therefore a pricing decision, executed thousands of times a day, made by engineers whose performance metric is latency and uptime rather than margin.
Mustafa Suleyman named the turn from the supply side. Tokenmaxxing, he argues, has been the story of recent months, and token efficiency is the next one. His prescription is co-optimization of models, harnesses, and reinforcement learning environments, on the observation that frontier generalist models are not required for every task and that product-specific tuning can hold or exceed frontier performance at dramatically lower token cost.
The best real customer outcome per dollar invested. Mustafa Suleyman · Optimizing the Frontier Performance Curve · July 2026
That is the frontier lab telling enterprises there is a curve to choose a point on. The demand side cannot act on it. An organization cannot select a position on a performance-cost curve it has never plotted, and most have not, because the telemetry stops at the vendor boundary and the boundary is drawn in the wrong place.
Two Statements Exist. The Third One Reconciles Them.
Every set of financials has three statements. The token economy has two.
The balance sheet is token capital: the workflows, evaluations, traces, and routing policy an organization accumulates and holds. The income statement arrived with Sarah Friar's scorecard, measuring useful intelligence per dollar and cost per successful task.
The third remains unwritten, and it is the one governance actually needs, because it is the only one that reconciles the other two. The token economy's cash flow statement covers compute commitments, prepaid and reserved capacity, cached against uncached consumption, and above all the timing gap between when tokens are spent and when work is realized.
That timing gap is where the money disappears. An organization can post a healthy income statement on the workflows that succeeded while burning enormous sums on the retries, the abandoned agent loops, the context reloaded forty times, and the eval runs that produced no durable artifact. None of that appears as a line item. All of it appears in the bill.
Every Token You Spend Is Either Capitalized or Expensed.
Reduced to its essentials, token governance is the oldest question in accounting applied to a new unit. Every token an organization spends does one of two things.
What the spend leaves behind
A residue. Something durable that makes the next run cheaper, faster, or more accurate.
Nothing. One answer, one satisfied request, and then evaporation.
The artifact
A captured trace, a hardened prompt, an automated eval, a routing rule, queryable institutional memory.
Output. Real work, correctly performed, with no accrual.
Effect on next quarter's unit cost
Falls. The loop tightens with use.
Flat or rising. The organization pays retail again.
What it is
Capital formation.
Consumption.
Both look identical on the invoice. Neither is distinguished by any dashboard currently deployed in most enterprises. That indistinguishability, rather than the size of the bill, is the actual governance failure of this year.
The organization with the better-looking budget has the worse position.
An enterprise spending four million a year where most of it leaves residue is compounding. An enterprise spending one million where none of it does is renting intelligence at retail and will rent it again next quarter at a higher price. Only one of them is building anything, and it is not the one the board is congratulating.
Four Ways Enterprises Govern the Wrong Number.
Unit price instead of unit of work
Cost per token is a procurement metric. Cost per completed task is an operating metric. Optimizing the first while the second climbs is how falling prices and rising overruns coexist without anyone noticing the contradiction.
The vendor boundary instead of the workflow boundary
Spend arrives shaped like an invoice, organized by provider. Value arrives shaped like a workflow, organized by business process. Nobody can reconcile the two, so the conversation defaults to the shape that already has a number attached.
After the architecture rather than before it
By the time a cost review happens, the retry policy, the context strategy, and the model tier are in production and already executing. Cost governance applied downstream of architecture is not governance. It is reporting.
Reduction instead of formation
The most expensive error, because it looks like discipline. Cutting token spend indiscriminately cuts the eval runs, the trace capture, and the experimentation that constitute the learning loop. The bill goes down. The asset stops accruing. The organization has economized its way out of the compounding position.
Five Moves That Separate Capital Formation from Cost Reduction.
- Meter at the workflow, not the vendor. Allocation is the prerequisite for every other move on this list. Until spend is attributable to a business process rather than a provider, no value question can be answered and no architectural decision can be priced. This is unglamorous plumbing and it gates everything.
- Make cost per completed task the primary unit. Not cost per token, not cost per seat, not cost per call. Cost per unit of work a business owner would recognize as finished. The Tokenomics Foundation is standardizing toward exactly this. Organizations already measuring it will inherit the standard. Everyone else will retrofit to it under deadline.
-
Apply a residue test to every agentic workload. For each one, name the durable artifact it produces. If the honest answer is that it produces output and nothing else, the workload is a cost center by construction and should be budgeted and defended as one. That is a legitimate answer. Not knowing is not.
Worked example
A support triage agent runs 40,000 times a month at $0.38 per completed resolution. Monthly line item: $15,200. The residue test asks one question. What does the organization own at the end of the month that it did not own at the start?
If the answer is forty thousand resolved tickets, the entire line is expensed. Real work, no accrual. But if the agent also emits labeled failure cases into a weekly eval set, and that eval set produced a routing table that moved most traffic to a cheaper tier, then part of that spend bought an asset that lowers next quarter's unit cost.
Same invoice. Same vendor. Same number in the ledger. Entirely different position.
- Move the token decision upstream into architecture review. Retry limits, context strategy, caching, and model tier are cost decisions disguised as engineering decisions. They belong in design review with a number attached, not in a variance report ninety days later. This is also where the performance-cost curve becomes actionable, because tier selection is an architecture choice made once and paid for continuously. Field signal A global enterprise software provider running support at very high case volume found that its cost per resolution was driven less by problem solving than by information retrieval, case preparation, and coordination between engineers. Resolution effort, not diagnosis, had become the dominant cost. The response was not a headcount reduction or a spend cap. It was a workflow re-architecture that moved case summarization, historical incident retrieval, and knowledge access from human effort to system capability, with triage and routing handled upstream of the engineer entirely. Resolution time and handling effort improved by double digits, and first-contact resolution rose materially. The expertise stayed exactly where it was. What was removed was the friction preventing engineers from applying it. And what the organization now owns is a retrieval and triage layer that makes every subsequent case cheaper to resolve than the last, which is the difference between a cost reduction and an asset.
- Put finance on the architecture review board. Not to approve spend. To understand what generates it. This is the practical fix for the 8% problem, and unlike most items on most executive checklists, it is available to any organization willing to make a single calendar change.
Standards Measure What Is Already Instrumented.
Industry bodies do not form around solved problems. They form when a category of spend has grown large enough, fast enough, and opaque enough that no single vendor can be trusted to define its own measurement.
The Tokenomics Foundation exists because token spend outgrew its vocabulary. The standards will come, and they will be better than what most enterprises have built internally. That is the good news and also the trap. Standards measure what is instrumented. An organization that has not separated capitalized token spend from expensed token spend by the time those standards land will simply receive a more precise account of a position it did not choose.
Gartner's projection that more than forty percent of agentic projects will be cancelled by the end of 2027 names three causes: escalating cost, unclear business value, inadequate risk controls. Those are not three failures. They are one failure, observed from three seats. Every one of them is what happens when an organization cannot say which half of its token bill was building something.
July's argument was to own the asset. This month's is to govern the spend, on the understanding that governing it well is not the same as spending less. Cutting the bill is easy and frequently destroys the position. The discipline is knowing, line by line, which half of the bill is capital.
In our next article, we take this into the sovereignty layer, examining what happens when the capability an organization has spent two years building sits entirely on infrastructure it does not control, and what it costs to move it.