The Exchange Rate
Your AI bill now arrives in two currencies. Only one of them has a rate you can look up.
Two line items on the same invoice. The first reads $6,200. The second reads 40,000 credits.
You know what the first one cost. You can check it against a rate card, forecast it, argue about it at renewal. The second one you cannot check at all, because you do not know what a credit is worth, and the company that sold it to you is the only party that does.
Most enterprises have accepted that second line as a billing detail. It is not a billing detail. It is a currency.
Your enterprise AI bill now arrives in two currencies. One is dollars. The other is a unit your vendor issues, exchanges at a rate it defines, and can reprice as the platform evolves.
That distinction is the entire subject of this article, because currencies behave differently than prices. A price you negotiate. A currency has an exchange rate, and an exchange rate is set by whoever issues it.
The market has started to notice half of this.
On August 27, 2026, Jared Spataro, Microsoft's Chief Marketing Officer for AI at Work, published an argument that enterprise AI spend has split into two very different bills. The first is the user subscription license, the flat per-seat fee that has priced enterprise software for twenty years. The second is usage-based billing tied to the work agents actually perform, scaling with the job rather than the headcount. IT owns the first and measures it on access and reliability. The business owns the second and measures it on outcomes. Judge either bill by the other's logic and neither makes sense.
He is right, and it matters that this is now the position of the largest vendor in the category. The split is no longer an argument. It is consensus.
What the piece does not say is that the two bills differ in kind, not just in size and owner. One is denominated in dollars. The other, at Microsoft, is charged in Copilot Credits, a shared currency that aggregates AI operations across the platform. Everything August argued about token spend, that some of it forms capital and some of it evaporates, depends on being able to read the meter.
And the meter is moving.
The Split Itself Is Correct.
Spataro's case rests on an observation this publication shares: most day-to-day knowledge work has already absorbed as much intelligence as it needs. Drafting a standard contract or summarizing a call comes out about the same whether you run it on this year's best model or on last year's. Past a certain point, paying for more capability stops changing the answer. That is what keeps the flat subscription sustainable. IT is not chasing a bar that keeps rising, because for most everyday work, the bar stopped rising a while ago.
Agentic work is different. It scales with the job, varies widely from one run to the next, and cannot be absorbed by a per-seat fee. Usage-based billing is the honest instrument for it. Funding the two bills separately, under different owners with different success metrics, is the correct operating conclusion, and organizations should adopt it.
The consensus stops there. What follows is the missing half.
What a Credit Actually Is.
Put a token and a credit side by side and the difference is easy to see.
Derivable and attributable.A token has a posted price. You can look it up, and it is the same number your vendor looks up. If you hold the records, you can add up what any workflow cost, run by run, and check the total against the invoice. You can derive the cost from records, and you can attribute it to the work that caused it.
A balance you cannot price.A credit is something you buy from the vendor and spend back to the vendor, and the vendor sets what it is worth on the way in and what it buys on the way out. You know how many you bought. You know how many are gone. What you do not know is the rate, and the rate is the whole thing. Change it, and your bill changes without a single line item moving.
The airport counter.Anyone who has bought foreign currency at an airport counter understands this. The number of bills in your hand is not the question. The question is what the counter charged you for them, and whether you could see the rate before you paid.
This is not a Microsoft observation. It is a category observation, and Microsoft is simply its clearest example, because Spataro states the arrangement plainly rather than burying it. Through 2026, developer and enterprise platforms have been moving their variable AI billing onto credit systems for the same underlying reason. Per-seat pricing breaks in the open once one person's agents do the work of five. Credits solve that problem elegantly, and they solve a second one at the same time: they let a vendor quote a number that feels fixed while metering everything underneath it.
Spataro's optimistic claim is that the economics keep moving in your favor, because falling model prices accrue to your subscription. On the subscription side, that holds. On the credit side, it holds only if you can read the exchange rate between the credits you buy and the tokens they actually purchase. Most enterprises cannot. The rate is not published, the aggregation is not itemized, and the price drop the vendor captures is not necessarily the price drop you receive.
The Bill Rises While the Price Falls.
If falling token prices were the whole story, AI budgets would be shrinking. They are doing the opposite, and the reason is now documented rather than anecdotal.
A field study published in Harvard Business Review in July 2026 compared knowledge workers using a conversational assistant against the same population using an autonomous agent. The agent completed work faster, at lower cost, and with higher satisfaction. It also did something the efficiency framing misses entirely: it expanded the scope of work people attempted. Users crossed occupational boundaries, took on more complex multi-step projects, and generated categories of tasks that rarely appeared with the assistant. The tool did not just accelerate the existing workload. It created workload that did not previously exist.
McKinsey's QuantumBlack practice, writing three days before Spataro, opened from the same fact: the cost of individual tokens continues to fall while enterprise spending on AI spikes. Unit price down, attempted scope up, total bill rising. This is why a per-task efficiency ratio can improve every quarter while the annual number blows through its forecast. An income statement lens sees each task getting cheaper. Only a cash flow lens sees the total moving.
The McKinsey piece also carries the strongest objection to this publication's premise, and it deserves a direct answer. For a customer service agent in banking, their analysis found token costs frequently represent just 20 to 25 percent of variable run costs, with human oversight accounting for 70 to 75 percent. If tokens are a quarter of the cost, why govern the token?
Because the token is not the biggest line. It is the only line you cannot check.
Oversight hours are salaries. You already know what your people cost per hour, and you already account for them. Infrastructure is a cloud invoice, itemized, in dollars. Those three quarters of the bill are things a finance team can audit today with tools it already owns.
The remaining quarter is the part that arrives as a credit balance, and there is no method by which anyone outside the vendor can verify it. McKinsey's prescription is to track the fully loaded cost of finishing a job, and that prescription is right. But a fully loaded number is only as trustworthy as its least checkable part, and that part is the credit line.
The size of the token line is not the point. What it is priced in is. The denomination problem
Residue Has an Address.
August's argument was that every token you spend does one of two things. It leaves something behind, or it burns off. What it leaves behind is a hardened prompt, a routing rule, an eval that now runs on its own, a captured trace. July's argument, building on the token capital framework, was that the residue is a balance sheet asset.
September's question is simpler. Where does that asset actually sit?
Everything on that list lives somewhere specific. A prompt library sits in your repository, or in a vendor's configuration store. An eval suite runs in your pipeline, or inside a vendor's studio. Agent memory accumulates in a database your team administers, or in a managed store you can only reach through someone else's API. In each pair, the left option is yours and the right option is rented.
That distinction is easy to defer, because both options work identically on a Tuesday afternoon. It stops being identical the day you want to move, or the day you want to price what you have built.
The exposure is not hypothetical. Teams building agents inside vendor-managed environments are writing institutional knowledge into configurations that do not travel, and enterprise surveys through 2026 have repeatedly found proprietary dependency in agent memory, model integration, and orchestration tooling among the leading stated concerns of agentic adopters. Research from Activant Capital puts the asymmetry plainly: token prices keep falling and model capabilities keep converging, but what does not narrow is who holds the record.
The counterweight is real and worth naming. Protocol-level work on agent context has been moving toward neutral governance and portable, tool-agnostic formats, which is the same standards impulse that produced August's subject. The plumbing is opening up. The question each organization has to answer is whether its own residue is being stored in a form that can use it.
Whoever Routes, Prices.
Spataro closes his piece with a preview: matching each task to the right kind of spend, automatically, is where he picks up next. That sentence deserves more attention than it will get, because automatic task-to-spend matching is a router, and the router is where the two bills are reconciled.
The infrastructure category already exists. Gartner defines the AI gateway as middleware between applications and models that handles routing, security, observability, and cost control, where cost control means tracking token usage and enforcing quotas. Gartner projects that 70 percent of engineering teams building multimodel applications will run one by 2028, up from 25 percent in 2025. The gateway is where metering physically happens. It is the cash register of the agent economy.
Which makes the strategic question a simple one: which side of the gateway does your meter sit on? If the gateway is yours, every request gets logged in tokens at published rates, and your cost per completed task is something you can calculate. If the routing happens inside the platform and gets priced in credits, the platform decides what runs on the expensive model and what runs on the cheap one, and you see neither the decision nor the rate. You are reading your own operating statement in translation.
This publication asked in April who owns the runtime. The router is how the runtime charges you. And notice that the saturation point, the very observation Spataro's whole argument rests on, is what makes routing worth fighting over. If most of your work no longer needs the most expensive model, then most of your work can run almost anywhere. The economics are already portable. Your organization may not be, and that is being decided right now, mostly by default.
Four Numbers Pointing the Same Direction.
The unit price is falling, the total is rising, and the meter is consolidating onto the vendor side of the stack at exactly the moment the work is expanding on yours.
Two Positions, One of Them Default.
| The Metered Position | The Denominated Position | |
|---|---|---|
| Unit of account | Tokens, at published rates | Credits, at an issuer-set conversion |
| Where the meter sits | Your gateway, logging every request | Inside the platform, reported as drawdown |
| Where residue lives | Exportable formats you administer | Configurations inside the vendor environment |
| Who prices the route | Your routing policy | The platform's orchestrator |
| What you can audit | Cost per completed task, derived from records | A balance, reported to you |
Neither position is free. The metered position costs engineering effort, gateway operations, and the discipline of maintaining your own records. The denominated position costs nothing today and builds up quietly. You find out what it cost at renewal, which is the one moment you need a number you never collected.
Five Moves.
- Put the meter on your side of the gateway. Route model traffic through infrastructure you administer, even when the models themselves are vendor-hosted. The goal is not multi-vendor for its own sake. It is that every request gets logged in a unit you can price independently.
- Price the conversion before you sign it. Any agreement denominated in credits should carry the disclosure in contract: the credit-to-token conversion, the notice period for revaluation, and export rights over the underlying usage records. A vendor confident in its pricing will produce the rate. A vendor that will not is telling you something.
- Give residue an address you control. Prompts in repositories, evals in your pipeline, agent context in portable formats. Build inside vendor studios when the speed is worth it, but treat anything that cannot be exported as rented, not owned, and price it that way in the capitalize-or-expense ledger.
-
Reconcile the reported bill against derived records. A credit drawdown is a claim. Instrumented session records are evidence. The reconciliation is one division.
Worked example
Go back to those 40,000 credits. Say they cost you $1,400. Now pull the token volume your own logs recorded over the same period and price it at published rates. Say that comes to $980. Divide the first by the second and your effective rate is 1.43.
That number is the only honest description of what you are actually buying. It should hold steady month over month. When it moves and your workload mix did not, the rate moved.
This publication's own thirty-day impact report, covering twenty-two sessions across four AI platforms, produced its figures from session records and invoices rather than from models. The method is public and MIT licensed. It exists because derived and estimated are different epistemic categories, and only one of them survives an audit.
- Route deliberately, and review it quarterly. If most workloads have a real choice of model, then routing policy is spend policy, whether or not anyone has written it down. McKinsey's recommendation is to put AI unit economics into the quarterly business review, and that is the right home for it. What belongs in that review is the question the next quarter of this publication takes up.
McKinsey closes its August analysis with a set of open questions.
How do chargebacks and cost coding have to change?
Who decides pricing when an agent works across three departments at once?
Each appears to be a budgeting question. In reality, each is a measurement question. Every answer depends on a number the organization can read for itself. That is why they remain open.
July said own the asset. August said govern the spend. September says both of those depend on something that comes first: the number has to be in a unit you control. You cannot value what you cannot price, and you cannot improve what you cannot measure.
Measurement is where this publication goes next. The fourth quarter of Token Exchange takes up the work the credit economy has made necessary: instrumenting AI at the source, calculating cost per completed task from records instead of balances, and building the reporting layer that lets an organization finally read its own statement.
The exchange rate already exists. The question is whether your organization can read it.