Summary

The Moonshots episode four days earlier, covering the week both frontier labs shipped: OpenAI's GPT-6 Astra and Anthropic's Fable 5.1 and Mythos 5.1, released about thirty days apart. The same five-person panel works through the benchmarks and reaches a more interesting conclusion than which model won — that the lead itself has stopped being worth much, because it now lasts weeks. Also covers Tesla's Cybercab event and the first signs of what the panel calls data starvation. Read alongside the 9 September episode, which is where the Navier-Stokes result and the agent breakout land.

Why it matters

The headline is a benchmark race, but the finding that survives it is Blundin's: a frontier lead is now worth about thirty days, with Chinese open models sixty days behind. If that holds, capability stops being a moat and the real contest is distribution — locking up partners, power, chips and states while briefly ahead. That reframes almost every 'X is now the best model' story the knowledge base will collect from here.

The second thing to carry forward is that two labs, independently and in the same week, shipped their most capable models behind capability gates — Anthropic splitting Fable from Mythos, OpenAI tiering cyber access. Whatever the labs say publicly about risk, their release engineering now assumes some capabilities are too dangerous to hand out. That is a stronger signal than any statement, and it sits directly alongside the CNN interview where an Anthropic researcher says alignment for superintelligence is unsolved.

It also gives the Handbook a caution about benchmarks: two reputable suites rank the same two models in opposite orders, because they weight mathematics and broad economic work differently. Any future claim that a model is 'the best' should name the benchmark or be treated as marketing.

1DB_QDiviH4-transcript.txt