The Signal
Frontier labs have solved scaling. The real constraint is now physical: Google processes 3.2 quadrillion tokens monthly at 12 gigawatts; AI will exhaust global data center capacity in ~3 years without hardware innovation. Energy is 50% of ChatGPT's token costs. This reframes the entire moat question—it's no longer "who builds the best model" but "who controls the power contracts and the silicon that runs inference at sub-millijoule efficiency." Margin compression in open-source inference (10-20% vs. closed at 80%) becomes irrelevant if you can't secure power or can't run models efficiently enough to make margins matter.
IMPORTANT
The bottleneck shifted from GPU scarcity to energy scarcity; whoever solves 1000x efficiency gains in the next 3.5 years owns the infrastructure layer.
What's Moving
- Energy-first data center design — All-in podcast flagging 4D computing (oscillator dynamics merging compute and memory) achieving 500 nanojoules per image vs. GPUs at millijoules. Product ships in 2 years. Removes von Neumann bottleneck. (via @allin)
- Sparsity as efficiency multiplier — Remove 90% of neural network connections, reduce compute quadratically, paradoxically improve trainability and performance. This is the immediate lever frontier labs will pull when energy becomes the binding constraint. (via @allin)
- Power contract leverage becomes real estate — @emostaque noting inference margins collapse when models satisfice; Anthropic's $8B+ valuation hinges on closed-model premium lasting. Once commodity models handle 99% of tasks (6 months per @bindureddy), power cost + hardware efficiency determines who survives. (via @emostaque, @allin)
- Routing layers entrench faster than models — @bindureddy already routing personal agents/RAG to open-source, hard agentic tasks to Fable 5.1. As energy economics bite, this routing becomes essential—you route to the model + hardware combo with best cost-per-task, not best absolute capability. (via @bindureddy)
Crosscurrents
- Frontier lab pricing power fragile — @emostaque's observation that Anthropic margins look "crazier" when you split subscription vs. API revenue admits the tension: closed-model premium only holds if open models can't run at scale due to energy/hardware constraints, not intelligence gaps. Once sparsity and 4D compute mature, pricing floor drops hard.
- 3-year energy deadline overstated or accurate? — All-in's "exhaust global capacity in 3 years" assumes linear token growth and no efficiency gains. But sparsity + new architectures could buy time. The real signal: energy innovation is now required, not optional.
Tradecraft
BULL
New compute architectures (4D, sparsity-aware training) shipping in 18–36 months could unlock 1000x gains and reset the moat from "models" to "power contracts + efficient hardware." First-mover advantage is massive.
BEAR
If energy constraints bite before efficiency solutions scale, inference margins collapse industry-wide. Closed-model premium evaporates. Frontier labs lose pricing power before they recoup training costs.
WATCH
Vera Rubin and Blackwell deployment timelines; sparsity adoption rates in production inference; power contract announcements from frontier labs (Meta, OpenAI, Anthropic).
Desk Notes
- @allin — Energy and thermodynamic efficiency as the new paradigm; 3.5-year timeline to biological efficiency parity
- @emostaque — Margin collapse in open inference forces routing logic; closed-model sustainability questioned
- @bindureddy — Hybrid stacks (cheap commodity + expensive reasoning) already default; energy drives next layer of routing
- @ylecun — Continuous representation space as more efficient than token-space reasoning (reinforces efficiency narrative)