When the price of something collapses, the interesting question is never how far it falls. It is where the value goes — because value rarely disappears. It moves. Cheap electricity did not enrich the electrons; it enriched the operators who controlled supply and the appliances that turned watts into things people would pay for. Cheap bandwidth did not pay the pipes; it paid the platforms that rode them. The price of intelligence is now falling faster than either ever did — roughly a thousandfold in three years — and the same question applies. As tokens approach free, where is the money settling?

That is a more useful question than the one the headlines ask. A new frontier model lands every week or so; a new open-weight champion arrives from China every few weeks; the leaderboard trades places by the point. It is a genuinely good show, and it matters for capability. But for anyone deciding what to build or what to back, the model race is mostly upstream of the thing that determines who gets paid. To see who gets paid, you have to map the supply chain the tokens actually travel — and mark, at each layer, whether the operator captures the deflation or is dissolved by it.

The clearest way into that map comes from one of the best-designed businesses in the stack.

Start with the toll that is levied on dollars, not tokens

Consider OpenRouter, the largest independent marketplace routing developer traffic across 400-plus models. Its design contains one quietly brilliant choice: it takes its fee — roughly 5% — on the dollars of inference that flow through it, not on the tokens. It marks up nothing per token; it charges a percentage of spend.

Follow what that single decision does. A fee on spend is, by construction, indifferent to which model wins. It tracks the value of the work, not the volume of the tokens. Cost-effective open-weight models can be the majority of the tokens moving through the platform and only a small share of its revenue; a premium lab can be a sliver of the tokens and the bulk of the dollars. On the platform's own reported usage, and consistently across several independent readings of it through mid-2026, one premium lab sits at roughly 12% of the tokens but about 46% of the dollars — with the cost-effective open-weight models showing the mirror image: most of the tokens, a minority of the money. Take the precise split as an estimate that drifts a few points quarter to quarter; the direction is not a coincidence. It falls straight out of taxing dollars instead of tokens, and it would hold whatever the exact numbers. It is not confined to one platform, either: a separate AI-gateway index reports open-weight models climbing toward a third of all volume even as blended price-per-token flattens — the same commodity-in-volume, value-in-spend split showing up wherever someone meters real traffic.

When tokens are free, the money doesn't vanish — it migrates: toward whoever prices the task and owns the demand, and away from whoever just moves the tokens.
Exhibit 1
A toll on dollars, not tokens — so the price war passes it by
Share of tokens Share of dollars spent
A premium closed model (e.g. a US frontier lab)12% → 46%
Cost-effective open-weight models (the commodity flow)most tokens, little spend
From reported OpenRouter usage (mid-2026), corroborated across several independent analyses of the platform's public rankings and billing data. The ~12% tokens / ~46% dollars split for a premium lab is an estimate that drifts a few points quarter to quarter; it is directionally robust because a fee on spend, by construction, tracks the value of the work, not token volume. A separate AI-gateway index shows the same volume-vs-spend split independently.

The implication is worth sitting with. The token price war — the thing every headline is about — does not touch a toll that is levied on spend. It presses on the businesses whose whole product is selling tokens cheaply. A marketplace built this way is, in effect, long the premium work and short the commodity flow, and it never had to pick a winner in the model race to get there. That is not a criticism of the sellers or a victory lap for the marketplace; it is simply what the design does. And it is the cleanest illustration of the migration this piece is about. Now walk the rest of the chain and the same pattern keeps appearing.

Follow the money down the supply chain

Walk the chain the tokens travel, from the metal at the bottom up to the application at the top, and mark at each layer whether the operator captures the deflation or absorbs it. That one distinction sorts the winners from the dissolved.

Exhibit 2
Where value accrues as tokens go to zero — the supply-chain map
LayerWhat it sellsMargin realityDeflation: capture or absorb?
Pure token resellerstokens, at spot pricethin, undisclosedAbsorbs it — dissolves. Same open weights, same open-source server, same rented chip as everyone else.
Neoclouds (GPU cloud)GPU-hours55–65% → ~14–16% after depreciationDepends. The moat is power and contracts, not the chip.
Serving / inference platformstokens + endpoints + efficiency~50% grossCaptures it — if it owns throughput efficiency, volume, and enterprise lock-in.
Custom-silicon speedfaster tokenscan be high (one is profitable)Best technology, most concentrated business — single-customer risk.
The router / marketplacea ~5% toll on spend + billing + dataasset-light, high grossCaptures it relatively — long the premium, short the commodity.
Margins: neocloud gross-to-post-depreciation from McKinsey analysis (relayed, Nov 2025); serving-platform ~50% gross is an analyst estimate (Fireworks, Sacra). The structure is the point; exact figures are directional.

Read the column on the right and a shape appears. Value is leaving the middle — the layers that just move a commodity — and pooling at the two ends: the layer that prices the task and the layer that owns the demand. Everything in between is a toll-free road. Start at the bottom, where the most capital and the most confident narratives sit, and work up.

The GPU is a melting ice cube — and why the moat is power

The neoclouds — CoreWeave, Nebius, Crusoe, Lambda — rent GPU-hours, and their headline gross margins look like software (55–65%). Then depreciation lands on an asset that is obsolescing on a shortening clock, and the real number falls to roughly 14–16%. The cause is plain enough: on-demand H100 rentals fell from $8–10 an hour in early 2024 to under $3.50 by 2026 as three hundred new neoclouds piled in. Rent a depreciating chip in a crowded market and you have a cost, not a moat.

The GPU is a melting ice cube. The durable moat is the gigawatt — then the software.
Exhibit 3
The melting ice cube — renting a depreciating chip is not a moat
H100 on-demand rental, $/hour
Early 2024$8–10
2026~$2–3.50
−64% to −75% in about two years.
Neocloud gross margin
Before depreciation55–65%
After depreciation~14–16%
The chip is a cost, not a moat.
Rental trend: Silicon Data / neocloud market reports (secondary). Margin: McKinsey analysis, relayed (Nov 2025). Illustrative ranges.

So what is the moat down here? It is worth being precise, because "the moat is power" is easy to say and easy to wave past. Power matters for two distinct, compounding reasons.

First, power is a direct and recurring input to the cost of a token — one of the few an operator can actually lock in. Break the cost of serving a token to first principles and it is the amortised cost of the chip, plus electricity, plus operating overhead, all divided by throughput. The chip line converges across everyone, because everyone buys from the same vendor at similar prices, and it depreciates whether you use it or not. Electricity is different: the price of a kilowatt-hour varies severalfold by geography, contract, and grid, and it recurs for the entire life of the asset. An operator who secures power at a structurally lower price — or better, generates it — carries a structurally lower cost floor for every token it will ever serve, for years. That is a durable advantage in a way that renting the same chip as a competitor never is.

Second, power is the binding constraint on scale, so it prices scarcity as well as cost. You can order more chips; you cannot run them without electricity and the grid interconnect, substations, and cooling to deliver it — and those take years and permits, not a purchase order. Power is what actually gates how many chips can be switched on. When the scarce input is gigawatts, whoever holds the long-dated, take-or-pay power contracts holds the real capacity, and that scarcity is contractible and defensible in a way the chip supply is not.

Put the two together and the moat resolves cleanly: power is simultaneously the cost lever (cheaper, owned energy lowers the floor under every token) and the scarcity lever (gigawatts gate who can scale at all). The operators pulling ahead reflect exactly this — Crusoe owns its energy; CoreWeave sits on a roughly $100B contracted backlog, though it also carries a concentration risk, with about two-thirds of revenue from a single hyperscaler. Above the metal sits the second, softer moat: the serving software that squeezes more sellable tokens out of each powered chip. A neocloud with neither power nor contracts is not really an infrastructure business; it is a leveraged bet on next quarter's spot price.

Capture the deflation, or absorb it

One layer up sit the serving platforms — Together, Fireworks, Baseten — and here the economics turn on a single question.

Every operator sees the same deflation. A few capture it; the rest absorb it.

Every serving business buys the same GPU-time at roughly the same cost and sells tokens at a market-cleared price. The gap between the two is throughput per dollar: how many sellable tokens you extract from one GPU-hour. Continuous batching lifts utilisation from 15–30% to 60–80%; speculative decoding adds more again on the output-heavy, agentic workloads that are becoming the norm. That efficiency is not a nicety — it is the gross margin, which lands these platforms around 50%, well short of the 80–90% of classic software.

The tension is that the base of that efficiency — the open-source inference servers everyone builds on — commoditises the gains almost as fast as they are found. So efficiency buys around 50% margins and a lead measured in quarters; it does not buy a permanent moat. The durable lock-in is the enterprise wrapper — fine-tunes, SLAs, governance, data gravity — reinforced by one powerful tailwind: as open-weight models proliferate, now led by Chinese labs, the serving layer is becoming the managed way to run open weights without owning GPUs — the Red Hat of open models. That demand accrues to breadth of catalogue and reliability, not to the lowest spot price. The pure per-token reseller underneath all of this is selling a commodity it doesn't make, on hardware it doesn't own, at a price someone undercuts next week; it absorbs the deflation and disappears into it. (The custom-silicon speed players are a revealing footnote: the best technology, the most concentrated business — one is genuinely profitable, but on the back of a single mega-customer, and the fastest of them was absorbed, team and all, by Nvidia rather than beaten in the market.)

Where the aggregator is strong — and where it is exposed

Which brings us back to the top of the map, and the most interesting business on it. On the argument so far the router looks close to unbeatable: asset-light, high gross margin, long the premium and short the commodity, and — the crucial point — its value rises precisely as the model field fragments.

Fragmentation is the aggregator's oxygen; consolidation is its test.

When six frontier labs reshuffle weekly and a flood of interchangeable open weights sits beneath them, switching models becomes a string change — which makes the layer that abstracts the switching more valuable, not less. Own the developer's single integration point, own the one billing relationship, and accumulate the asset no individual provider can build — a cross-provider view of what actually wins on cost, quality, and latency this week — and you have real gravity. That data flywheel is why the broader data-infrastructure stack (Google's growth fund led its Series B; Nvidia, Snowflake, Databricks, and MongoDB all participated) has bought optionality on the layer, and why it now attracts multibillion-dollar takeover interest at a steep premium to its roughly $1.3B valuation.

The exposure is the mirror image of the strength, and it is a structural tension the whole aggregation layer faces, not a flaw in any one operator. The moat is strong on distribution and data and thin on defensible technology and pricing power. The core routing function is a solved, open-source problem; rival gateways already give routing away and have competed the per-token markup to zero. A 5% toll is an inviting target that hyperscalers can bundle toward nothing and that a capable team can sidestep by self-hosting. And there is a genuine reflexivity to sit with honestly: by making models more substitutable, an aggregator helps drive the very commodity prices that thin the dollar pool its percentage is levied on. None of this makes the business fragile; it makes it dynamic. The resolution is the tell — the neutral middle rarely stays a neutral pipe. The operators that endure climb off the commodity floor into governance, routing intelligence, and spend management, where there is pricing power; and the layer often finds its most natural home inside a larger platform that wants an AI control plane. A major security vendor recently acquired one of these gateways for exactly that — not the routing, the control plane. That is not decline; it is where this kind of value tends to come to rest.

What founders should build

Strip it to a build order.

  • Think twice before building a token reseller or a generic router. Both are commodities the moment they ship — the reseller races the most cost-effective open-weight host to zero; the router competes with free tools and hyperscaler bundles. It is a zero-markup contest with the largest companies on earth on the other side.
  • Build what the pipe lets you own, not the pipe itself: the routing intelligence that provably saves money on real workloads (sell the arbitrage — "this query needs a $15 model, that one a $0.10 model" — not the plumbing); the spend-governance and billing layer (be the Ramp of tokens); the agent security and observability an enterprise cannot deploy without; or a vertical or region where the hyperscalers are thin and data residency matters — a compliance-native layer for a regulated Asian workload is a real opening.
  • Model the business as a fintech taxing gross spend, and assume your take-rate halves. If the unit economics only work at 5%, don't build it. Your cost of goods is throughput per dollar, so serve the commodity 80% on self-hosted open weights (now Chinese-led and near-frontier) and reserve the premium APIs for the hard tail — but keep the moat where it is defensible: your data gravity and your fine-tunes, never the model itself.

What investors should back

The same logic, as an underwriting rule.

  • Value in this layer accrues to demand aggregation, cross-provider data, and task-pricing — not to routing technology or GPU ownership. The neutral middle trends toward acquisition rather than durable independence; back it for the strategic exit and underwrite take-rate compression, not a permanent 5%.
  • Serving platforms: underwrite them at ~50% gross margins, not 75% SaaS, and stress-test the roughly 20-times-ARR valuations against the day the efficiency edge commoditises. These names are priced for both the volume tailwind to persist and the moat to hold; only one of those is safe to assume.
  • Neoclouds: back them only where power and take-or-pay contracts are locked; underwrite to the 14–16% post-depreciation margin, the depreciation schedule, and customer concentration. No power, no contracts — that is a levered spot-price bet; pass. Custom silicon is venture-binary — back it for the acqui-exit, not the standalone franchise.
  • Be wary of anything whose thesis is a token markup (already competed away), any enterprise-only router competing head-on with hyperscaler marketplaces, and any pure reseller wearing a software multiple.

Coda: who gets paid when the token is free

When the price of anything collapses, the lazy read is that the value collapses with it. It never does — it moves. Cheap electricity did not enrich the electrons; it enriched the operators who controlled supply and the appliances that turned watts into something people would pay for. Cheap bandwidth did not pay the pipes; it paid the platforms that rode them. Tokens are on the same path, and the value is settling the same way: with whoever prices the task, owns the demand, and governs the spend — and away from whoever merely moves tokens or rents the metal.

So let the model race keep the headlines; watch it for the capability, and enjoy the sport. But the money is quietly settling somewhere else. For a founder, that is a build order: own a scarce input, not a commodity you resell. For an investor, it is an underwriting rule: back the toll and the intelligence, not the traffic — and price in the day the toll itself gets cheaper. The question was never who builds the best token. It is who gets paid when the token is free.