Summary
The Moonshots episode four days earlier, covering the week both frontier labs shipped: OpenAI's GPT-6 Astra and Anthropic's Fable 5.1 and Mythos 5.1, released about thirty days apart. The same five-person panel works through the benchmarks and reaches a more interesting conclusion than which model won — that the lead itself has stopped being worth much, because it now lasts weeks. Also covers Tesla's Cybercab event and the first signs of what the panel calls data starvation. Read alongside the 9 September episode, which is where the Navier-Stokes result and the agent breakout land.
- GPT-6 Astra's published benchmarks, read from OpenAI's release page: Frontier Math Tier 4 saturated at 98%, ARC-AGI 3 at 99.9%, ExploitBench at 100%. OpenAI frames the release as being about efficiency rather than raw intelligence, claiming a new Pareto frontier for intelligence index against output tokens — it uses fewer tokens to do the same work, which changes the economics as much as the capability.
- A hallucination figure was read out as falling "nearly by half, from 92 to 51%" with accuracy rising. Recorded as stated, but the baseline is not given and 92% does not obviously describe a general hallucination rate — treat it as unexplained until the release page is checked directly.
- Anthropic released two models with the same underlying intelligence and different safety envelopes: Fable 5.1 broadly available, Mythos 5.1 held back for tightly controlled cybersecurity and life-science programmes. OpenAI did something structurally similar with Astra, rolling out tiered cyber access to trusted partners first. Two labs independently arriving at capability-gated release is the pattern worth noting, not either release.
- Fable 5.1 scored 60.9% on Humanity's Last Exam without tools and 65% with — described on the episode as the highest published score of any frontier model on HLE, a benchmark built for expert-level reasoning across many fields. Its terminal-bench science score doubled to 52.6.
- The two benchmark suites disagree about who leads, which is the useful part. On Epoch's Capabilities Index, which leans mathematical, GPT-6 Astra is first. On Artificial Analysis, weighted towards broad economically valuable work, Fable 5.1 is first and Astra trails even Meta's Muse Spark. Wissner-Gross's judgement: Fable 5.1 is the strongest generally available model, while Astra is faster and stronger at mathematics.
- Mostaque read Astra's uneven profile as a sign it is "the first non-benchmaxed model" — a genuinely new pre-train rather than one tuned to score well. An unusual reading of a model losing benchmarks, and worth revisiting when independent evaluations arrive.
- Wissner-Gross's structural guess at what changed inside Astra: looped transformers — one transformer stacked on itself with tied weights and run recurrently — and he notes Chinese labs injecting recurrence elsewhere, such as Kimi's linear attention. If right, it would mean a new scaling axis, depth rather than size. Explicitly his inference from public comments, not anything OpenAI has stated.
- Blundin's framing of the race, and the most durable point in the episode: Fable 5.1 sits a notch above Astra, but they are thirty days apart and Chinese open-weight models roughly sixty days behind that. A frontier lead now lasts about a month, so the value is not the lead but what you lock up while holding it — partnerships, real estate, generators, chips, whole states and governments. The competition has moved from capability to distribution.
- The panel expects data starvation next. Maths and coding had abundant training data and are now, in their phrase, cooked; physics follows. Architecture and drug design are data-starved, and Blundin reports every company he is involved with that gathers proprietary data growing faster than anything he has seen.
- Tesla's Cybercab event: Musk's stated intention is to sell at about $30,000 so that buyers can put several on the road as revenue-earning robotaxis. The panel expects parts of cities to exclude human drivers on efficiency grounds. A stated intention, not a shipped price.
- Mentioned in passing and worth its own line: Fermat's Last Theorem has been formally verified in about 13 million lines of code, proving some 29,000 theorems along the way.
Why it matters
The headline is a benchmark race, but the finding that survives it is Blundin's: a frontier lead is now worth about thirty days, with Chinese open models sixty days behind. If that holds, capability stops being a moat and the real contest is distribution — locking up partners, power, chips and states while briefly ahead. That reframes almost every 'X is now the best model' story the knowledge base will collect from here.
The second thing to carry forward is that two labs, independently and in the same week, shipped their most capable models behind capability gates — Anthropic splitting Fable from Mythos, OpenAI tiering cyber access. Whatever the labs say publicly about risk, their release engineering now assumes some capabilities are too dangerous to hand out. That is a stronger signal than any statement, and it sits directly alongside the CNN interview where an Anthropic researcher says alignment for superintelligence is unsolved.
It also gives the Handbook a caution about benchmarks: two reputable suites rank the same two models in opposite orders, because they weight mathematics and broad economic work differently. Any future claim that a model is 'the best' should name the benchmark or be treated as marketing.
1DB_QDiviH4-transcript.txt