The Signal
@ylecun just made explicit what the inference-time optimization dispatch missed: the field is bifurcating on foundational architecture, not just training efficiency. His positioning of Energy-Based Models and gradient-based planning as the path to human-level AI directly contradicts the auto-regressive LLM scaling narrative that still dominates closed labs. Meanwhile, @bindureddy's August open-source flood (GLM 5.5, DeepSeek Pro, Qwen 3.8-Max) is executing the post-training-wins playbook at scale—but on fundamentally different architectures than what OpenAI and Anthropic are optimizing. This is not competition within a paradigm. It's two different bets on what architecture wins.
What's Moving
- World models as the research legitimacy play — @ylecun positioning ODAI and EBM as foundational theory for inference-time optimization, not a hack. The signal: continuous gradient-based planning over discrete token prediction is the architecture of the next phase. This reframes the last six months of "inference search" posts as scattered experiments pointing toward a coherent paradigm. (via @ylecun)
- Qwen 3.8-Max arrives as the new open-source anchor — 2.4T parameters, Sonnet-class or better, $2/$6 pricing, caching at $0.25/M. @bindureddy's reading: "move from K3 to Qwen because it's 2x cheaper." This collapses the frontier window further and signals open-source has solved the post-training problem sufficiently that pure scale no longer buys you the margin. (via @bindureddy, @svpino)
- DeepSeek Flash hype collapse clarifies the routing signal — @bindureddy's correction (Flash is "way worse than Grok 4.5," not Fable-class) is critical tradecraft: small models win on cost, not capability. Enterprises routing by task complexity, not model prestige, is now the operational posture. (via @bindureddy)
- Auto-regressive limits becoming explicit constraint — @ylecun's thread on token-level information density (20 bits per token vs. high-dimensional brain state) is the theoretical underpinning for why scaling text prediction plateaus. This is not new critique; it's now the baseline technical assumption among frontier researchers. (via @ylecun)
Crosscurrents
- Closed labs still own speed-to-capability, but timing window is shrinking — @bindureddy concedes 12–16 week lead still exists, but @emostaque's sub-$0.28/M Opus-class model (July) and Qwen 3.8 (August) suggest the lead is measured in weeks, not quarters now. Anthropic has no Luna-class hedge.
- "Better at benchmarks" is now the pejorative — @bindureddy's framing of DeepSeek Flash as "benchmark-maxxing" vs. real capability signals a shift in how practitioners evaluate models. Benchmark superiority no longer indexes to production performance.
Tradecraft
Desk Notes
- @ylecun — Defending world models + ODAI as the path to human-level AI; LLMs are text interface layer, not the core engine.
- @bindureddy — Routing by task + cost; Open-source lead closing; Qwen 3.8 is the new price-performance anchor.
- @svpino — Leaving Anthropic over lobbying stance; betting on open-source + multi-model routing stacks.