The Signal
@emostaque's casual observation that "DeepSeek runs an 80% margin" on inference while "Fireworks and others hitting $1bn run rate" reframes the entire stack: the money moves to running open models, not training them. @bindureddy's K3 open-weight drop (Monday, "3x faster") accelerates this. Closed-model labs (OpenAI, Anthropic) are now in the position of funding their own replacement—they train expensive frontier models that get distilled into cheaper open weights, then watch margin-maximizing inference providers (Fireworks, Together, DeepSeek) own the cash flow. The OpenAI security incident from last week compounds this: closed models now carry operational liability that open inference avoids entirely.
What's Moving
- Distillation as competitive weaponization — @bindureddy's blunt pushback ("everyone distills, even Anthropic trained on user agentic loops") signals the industry has normalized knowledge extraction. If US labs can't legally prevent distillation, they've already lost the margin game. Open weights will be cheaper and comparable within 12 weeks. (via @bindureddy)
- K3 Monday release is the pivot moment — @bindureddy's explicit confidence ("Kimi K3 will be truly open-weight and 3x faster on Monday") collapses the closed/open performance gap. Anything routing to Sonnet or GPT 5.5 becomes economically irrational. Inference providers will immediately rebrand around K3. (via @bindureddy)
- $1bn+ inference runs on open weights (real) — @emostaque naming Fireworks at "$1bn run rate" validates that open-weight inference is not theoretical; it's the largest emerging margin pool in AI. DeepSeek's 80% operating margin proves the unit economics. (via @emostaque)
- Memory + open weights collapse context-window costs — Combining persistent agent memory (from yesterday's dispatch) with open-weight inference means teams no longer pay for repeated inference on the same problem. Closed models' context-window pricing becomes a relic. (via @bindureddy implicit)
Crosscurrents
- Closed labs' venture model breaks — If inference margin collapses to open-weight commodities, where does capital return? Anthropic and OpenAI are now subsidy networks for open-source infrastructure. @bindureddy's acid question ("Why would you fund someone who will eventually kill you?") about Google + Anthropic is the real tell.
- Distillation legality remains contested — @bindureddy's "China distilled model outputs is SUPER LAME" argument works domestically but doesn't block regulators from carving exceptions. European framing of distillation as "legitimate open innovation" (per @emostaque's July 22 cite) may create regulatory moats anyway.
Tradecraft
Desk Notes
- @bindureddy — K3 Monday release is the inflection; everything at Sonnet/GPT 5.5 level becomes a cost arbitrage play against closed models.
- @emostaque — Inference margin capture ($1bn Fireworks run rate, 80% DeepSeek margin) is the real business; training is subsidy.
- @svpino — Skills + MCP scaffolding now adds value above inference layer; agents self-pay via x402 protocol (autonomous commerce).