The $11 trillion AI capex mirage (September 2026)

A “two-lab monopoly” extracting the surplus of the automated global workforce is a fantastic science-fiction pitch. It’s also a bet against a hundred and fifty years of microeconomics. Here’s the arithmetic the thesis quietly needs to defeat.


Dwarkesh Patel recently sat down with SemiAnalysis’s Dylan Patel, and out of the conversation came a number I keep seeing quoted: $11 trillion of AI infrastructure buildout, most of it captured by two labs — presumably OpenAI and Anthropic — which then hold a permanent monopoly on the automated global workforce.

I want to take the number seriously, because claims this large deserve arithmetic, not vibes. When I sat down with the arithmetic, four separate microeconomic forces show up, each of them individually sufficient to sink the thesis. Together they leave essentially nothing standing.

The forces are old and boring: diminishing marginal returns, Bertrand competition, the competitive fringe, and consumer surplus. That’s the whole post, plus a final section cleaning up the best objection to it — that all four together still might not be enough. It’s the same first-year IO reading list the internet ran through in 1999, and the outcome is going to rhyme.

A companion post, A Beautiful World, and Nobody Owns It, comes at the same target from the production-function side: how large the cognitive slice of the economy turns out to be depends entirely on whether thinking substitutes for physical work or merely complements it — and the answer for the labs is the same either way.


The number needs $2T of gross profit, every year, forever

Start by pricing what $11T of capex actually demands. Assume a twelve-year useful life for the hardware and a 10% cost of capital. The annualized cost of that stock, treating it as a level annuity, is roughly:

\[ \text{Annual cost} = 11{,}000 \cdot \frac{r}{1 - (1+r)^{-n}} \approx 11{,}000 \cdot 0.147 \approx \$1.6\text{T/yr}. \]

Add operating cost (power, cooling, staff, opex on the fabs), and to earn back capital plus keep the lights on you need something like $1.7-2T/year of gross profit — for a decade, in perpetuity.

Two things about that number before we go on.

First, it is deliberately generous. A twelve-year useful life for an AI accelerator is a fiction that flatters the buildout; the realistic economic life is closer to five years, and at \(n=5\) the annuity factor goes from 0.147 to 0.264 and the hurdle goes from $1.6T to $2.9T/yr. Every argument below uses the friendly number.

Second, and more important: $1.7-2T/yr is the break-even, not the jackpot. It is what the buildout has to clear merely to return capital at a perfectly ordinary 10% — no monopoly, no excess profit, nobody owning the automated workforce. Hold onto that distinction. It turns out to be the whole ballgame, and I come back to it after the four forces.

For scale: Alphabet’s entire operating income in 2024 was ~$100B. Global enterprise software revenue is ~$800B. Total U.S. payroll is ~$10T. The thesis is that two firms extract 20% of U.S. payroll, at monopoly margins, and keep doing it forever. Every argument below is a separate reason that doesn’t happen.

Loss goes down. Value doesn’t.

The scaling laws are real. Cross-entropy loss as a function of training compute is a power law with a small exponent:

\[ L(C) = k \cdot C^{-\alpha}, \quad \alpha \approx 0.05 \text{ (Kaplan / Chinchilla range)}. \]

Small exponent means you pay a lot of compute for a little loss. But the deeper problem is that loss isn’t revenue. Revenue is bounded above by the value of the task being automated. A model that does junior-analyst work perfectly cannot charge more than a junior analyst’s fully-loaded hourly rate, no matter how many gigawatts it took to train. The value function is asymptotically bounded:

\[ V(C) = V_{\max}\left(1 - e^{-\lambda C}\right), \quad \frac{dV}{dC} > 0, \quad \frac{d^{2}V}{dC^{2}} < 0. \]

The two curves compose the wrong way. Compute cost grows roughly linearly in the FLOPs bought; capability grows as a small power of compute; revenue is a saturating function of capability. Stitch them together and the marginal revenue per marginal capex dollar is:

\[ \frac{dV}{d\$} \;=\; \underbrace{\frac{dV}{d\text{cap}}}_{\text{saturating}} \cdot \underbrace{\frac{d\text{cap}}{dC}}_{C^{-(1+\alpha)}} \cdot \underbrace{\frac{dC}{d\$}}_{\text{roughly constant}} \; \to \; 0. \]

Concretely: pushing a coding agent from 95% task completion to 99% can require ~10× the compute. What does that buy?

The obvious objection — the one João Sedoc led with, reading the first draft — is that it buys a great deal, because reliability compounds. Chain a per-step success rate across a fifty-step agentic workflow:

\[ 0.95^{50} \approx 7.7\%, \qquad 0.99^{50} \approx 60.5\%. \]

Four points of per-step reliability is an eightfold improvement in end-to-end success — the difference between a demo and a deployment. Willingness to pay for that is discontinuous, not marginal. So much for the saturating value curve.

The objection is right about the value and wrong about who has to pay for it, in three steps.

Retries collapse the cliff, and retries are cheap. If you can detect a failed step, one retry takes a 95% step to 99.75%, and \(0.9975^{50} \approx 88\%\) — comfortably better than the 60% you get from a perfect single-shot 99% model. In coding agents you very often can detect it: compilers, type checkers, linters and test suites are failure oracles you already own. The cheapest route across a deployment threshold is a verifier loop burning commodity inference, not a frontier model burning 10× pretraining. The nonlinearity is real, and it routes demand toward scaffolding and away from the $11T.

The extreme version of this is the one I’ve been living in. On the Lean project the oracle isn’t a linter, it’s a kernel: a proof either type-checks or it doesn’t, with no false positives and no judgment call. Against an oracle that good, per-attempt reliability nearly stops mattering — six failed tactic rounds cost minutes of wall clock, and the attempt that finally compiles is exactly as correct as a first-try success would have been. That is an unreliable model emitting verified output, and what made it affordable was cheap inference against a perfect checker, not a better frontier model. l3m, the coding agent built on that machinery, is the same trick turned on the agent’s own sandbox.

And the sting is in what turns out to be scarce instead. Once the retry loop is cheap, the bottleneck moves off proving and onto stating: saying precisely what you want becomes the expensive step, and that is a human design cost no quantity of compute relieves. If the binding constraint on shipping reliable agents is specification rather than capability, then $11T of pretraining is buying down the wrong number.

Independence is doing enormous work in that arithmetic. The \(p^{50}\) model treats step failures as independent draws. They aren’t. The same model, meeting the same ambiguous spec with the same blind spots, fails in correlated clusters — which makes the true end-to-end rate better than 7.7% and the cliff shallower than the formula advertises.

A threshold doesn’t repeal the asymptote; it paints a staircase on it. Cross 99% and the workflows that needed 99% light up. The next tranche wants 99.9%, which is another 10× for a smaller step. \(V(C)\) is still bounded above by what the task is worth.

And then the part that actually matters for the thesis: none of this touches who keeps the money. Grant that reliability is worth a discontinuity. Thresholds are precisely where you should expect both labs to arrive within a quarter of each other — that is what a shared scaling law and a shared talent pool produce. A step change in value that two firms reach simultaneously is a step change in consumer surplus, which is the subject of the next three arguments.

The wall is still a wall. It just has a staircase painted on it, and the last dollar of the $11T is still buying fractional cents.

Bertrand eats the margin

When two firms sell a homogeneous good at marginal cost \(c\), Bertrand’s 1883 result is one line: undercutting drives price to marginal cost. If Firm 1 posts \(p_1 > c\), Firm 2 posts \(p_1 - \epsilon\) and takes the whole market. Iterate:

\[ p^* = c, \qquad \pi^* = 0. \]

The thesis needs frontier models to stay meaningfully differentiated forever. Look at the actual differentiation. Whatever OpenAI, Anthropic and Google shipped this quarter sits within a handful of points of the others on almost every reasoning benchmark. Naming the current three would date this paragraph within weeks, which is itself the point: the durable observation is that the scoreboard has looked like that for three years running, and every time one lab opens a gap, the others close it in months. Third-party routing layers — OpenRouter, LiteLLM — already treat the frontier as interchangeable and route on price. And the price has been collapsing: the API cost of GPT-4-class capability fell roughly 80% between 2023 and 2025. That’s not a monopolist’s margin curve. That’s the leading edge of Bertrand doing exactly what Bertrand does.

The switching cost between labs is an SDK config change. When the switching cost is a config change, the cross-price elasticity of demand approaches infinity, and \(p \to c\) is not a limit case; it’s next quarter.

The fringe is already sipping the bottom

Textbook fringe model: a dominant producer’s residual demand curve is the market demand minus everything the competitive fringe supplies at each price:

\[ D_r(p) = D(p) - S_f(p). \]

The fringe here is Llama, Mistral, Qwen, DeepSeek — every open-weight checkpoint on Hugging Face. Fringe capability trails the frontier by roughly 12-18 months and shrinking, and — this is the part the $11T number ignores — the shape of enterprise demand is heavily concentrated on the easy end. Classification, summarization, extraction, RAG, code completion: something like 80% of enterprise AI queries are handled acceptably by a 70B open-weight model running on hardware the enterprise already owns, at fractional pennies per million tokens.

Which means the frontier labs are not pricing into the full demand curve \(D(p)\). They’re pricing into the residual \(D_r(p)\) — the top few percent of queries where the extra capability actually matters. That’s a real market. It’s not a $2T/year market. You cannot service $11T of debt on the top-of-funnel while the fringe drinks the bottom 80%.

Producer surplus vs. consumer surplus: the real fatal one

This is the argument that has nothing to do with competitive dynamics and everything to do with who owns the value that a general-purpose technology creates.

Total surplus \(TS\) from a market decomposes cleanly into producer and consumer parts:

\[ TS = PS + CS, \] \[ PS = P^*Q^* - \int_0^{Q^*}\! MC(Q)\,dQ, \qquad CS = \int_0^{Q^*}\! P(Q)\,dQ - P^*Q^*. \]

Under competition, price collapses to marginal cost, so:

\[ \lim_{P \to MC} PS = 0 \;\;\Longrightarrow\;\; CS \approx TS. \]

The historical scoreboard on producer capture for general-purpose technologies is brutal:

Technology Est. total surplus Producer capture
Internet infrastructure ~$10T cumulative ~2-5%
GPS ~$1.4T/yr <1%
Container shipping trillions negligible
Grid electricity civilization-scale ~1%

The pattern isn’t an accident. General-purpose technologies leak nearly all of their surplus to consumers, because that’s what competition does to producers of general-purpose goods. If AI truly automates $10T of labor per year, competition will drive most of that $10T into the pockets of the buyers — cheaper software for consumers, higher margins for the businesses deploying it, price competition that hands the savings downstream and eventually to the customer.

The producers get what a competitive market pays producers: enough to cover marginal cost, plus whatever narrow rents they can defend for as long as they can defend them. Which is not zero. It’s just not $2T/year, and it doesn’t compound for a decade.

Zero profit is not the same as zero recovery

Here is the objection that nearly took this post down — João’s second, and the better of the two. It is a good one. Everything above establishes that competition drives economic profit to zero. But the $1.7-2T/yr hurdle in the first section was never a monopoly number — it was the break-even, the annuity that returns capital at an ordinary 10%. Zero economic profit means precisely that you cleared it. So a careful reader can grant all four forces and reply: “fine, the $11T earns its cost of capital, nobody gets rich, and there was never a mirage here at all.”

Competition eliminating excess returns and competition preventing recovery of capital at a normal required return are different claims. Conflating them is the difference between a post that proves something and a post that proves nothing, so let me separate them.

The four forces establish the first one flatly. The second needs one additional observation: in this industry marginal cost is nowhere near average cost. A trained model serves its next token for about the price of the electricity, while essentially the whole $11T sits in fixed, sunk, rapidly obsolescing capital. Bertrand pushes price toward marginal cost; recovery requires price at average cost. When those two numbers differ by orders of magnitude, “price falls to marginal cost” and “the capital comes back” are not jointly satisfiable, and the industry has to land on one of two branches:

Which branch we land on is an empirical bet, and I don’t need to win it, because the thesis dies on both. The claim under examination isn’t “a lot of GPUs get bought.” It’s “two firms own the automated global workforce.” Branch one says the buildout was overstated. Branch two says the buildout was real and somebody else kept the surplus. Neither ends with a two-lab monopoly.

It matters enormously for who eats it, though, and that is where the financing stops being a footnote. The $11T is usually described as something like $5T of cash off hyperscaler balance sheets and $6T of debt. Cash-funded overbuild is a quiet disappointment: shareholders earn 4% where they wanted 10%, an analyst writes a sad note, life goes on. Debt-funded overbuild is a different animal, because a contractual claim does not care how much consumer surplus you generated. That $6T is the mechanism that converts branch two from an underwhelming decade into a credit event. The railroads are the precedent there too, and not the flattering part of the story.


The bibliography, both sides

This fight is loud and roughly a decade old. What I want to do is locate this post inside it, because “the $11T number is wrong” is a big tent and the tenants disagree with each other for very different reasons.

The bulls — who need the number to be roughly right.

The bears — who agree the number is wrong, for different reasons.

Fellow travelers on the specific argument. The closest work in shape is the emerging academic literature on AI as a general-purpose technology with consumer-captured surplus — Stanford Digital Economy Lab’s “What is Generative AI Worth?” (Apr 2026) and Luis Garicano’s “Three theses on AI value capture” both apply the same producer/consumer surplus decomposition I use here. Neither pairs it with a four-forces breakdown, and neither prices the $11T number against the required annuity of gross profit.

The claim of this post, if it’s original at all, is the joinery: price the annuity on the capex, show that four independent microeconomic forces each individually block the monopoly the number depends on, and then separate the two things “zero profit” can mean so the result survives an economist reading it. Every piece is someone else’s; the assembly I haven’t seen elsewhere.


The rhyme

None of this is exotic. It’s what every general-purpose technology has done since roughly forever. Infrastructure gets built — often overbuilt, often by capital that gets wiped out in the shakeout — and the value flows to the users of the infrastructure, not the owners of it. The railroads didn’t own the 20th century economy; they owned the narrow slice of it they could keep charging monopoly rents on, and that slice kept shrinking. The internet is running the same play right now.

The $11T thesis requires two labs to defeat, simultaneously, diminishing returns on compute, Bertrand competition, an open-source fringe, and 150 years of consumer-surplus arithmetic. Any one of those is fatal to the number. All four together is a story that does not survive contact with a first-year syllabus.

The capex might still get spent — capex often does, that’s what capex bubbles are for. But the “permanent two-lab monopoly extracting 20% of global payroll” story is not how competitive markets end. It’s how they begin.