Summary
A rumour audit of what comes after GPT-6 Astra, built around an explicit badge system - confirmed, credible leak, speculation, debunked - applied to every claim. The channel grades its own earlier wrong call on air and declines to read leaked spec tables. It also carries a long, undisguised advertisement for the presenter's own platform in the middle. Treat the discipline as genuine and the enthusiasm as sponsored.
- Confirmed specifications: Astra shipped 3 Sep as GPT-6 under that API label, about a million-token context, knowledge cutoff end of April, text and images in and only text out. Priced at $10 per million tokens in and $50 out.
- That last point kills a large part of the omni story on its own: Astra takes no audio and no video in either direction. As the presenter puts it, 'the model that got called the arrival of the AGI era cannot hear you.'
- The independent scoreboard contradicts the launch framing. On Artificial Analysis 4.1.1 Astra is reported at 61.2 against Fable 5.1 at 65.7, Opus 5 at 63.1 and Fable 5 just above 62 - fourth place. Both things are said to be true: it wins the evals OpenAI published and loses the one it did not run.
- Two conflict-of-interest disclosures that matter more than the scores. Epoch AI is reported to have disclosed that OpenAI funded Frontier Math and holds exclusive access to part of it; and OpenAI is reported to have noted its Claude comparison runs used modified eval settings. Neither makes a number false; both make it a company number.
- Regressions are itemised rather than glossed: roughly 80 Elo down on GDPval against Sol, two to three points given back on banking-agent and side-code tasks, a similar dip in long-context reasoning. About 10% fewer output tokens, but at 2.5x the per-token price - a net of roughly 75% more expensive per average task. The verdict offered is 'a side grade with a great launch video'.
- The one number the presenter says genuinely impresses him is the omniscient hallucination rate falling from about 92% to about 51% at maximum reasoning effort with accuracy up about four points - and he argues that is worth more than any saturated maths benchmark.
- Four frontier labs shipped flagships in one week: Anthropic's Claude 3.1 and Mythos 5.1, Meta's Muse Spark 1.3, Astra, with Gemini Omni 1.1 Flash a week earlier. Meta reports 75.4% on a deep coding benchmark, Astra 74.1, Opus 5 around the same - which the presenter calls a statistical tie reported as three separate claims of the lead.
- On Fermat's Last Theorem this source materially disagrees with the Diamandis entry in this batch: here it is Claude working autonomously for 11 days, not 13 million lines of code proving 29,000 theorems. Both are second-hand. Record the conflict rather than picking one.
- The leak chain is dismantled carefully. A model code-named Bell, reported north of 10 trillion parameters, traces to a single post by one account, reshared by four outlets. 'Four headlines from one tweet is not four sources.' A commissioned 19-source research pass found no publicly verifiable document about any model beyond Astra. Badge: credible leak, not confirmed.
- The nuance most coverage flattened: Bell is described as a base model after GPT-6, not as GPT-7. 'OpenAI finished pre-training a very large base model' and 'GPT-7 exists' are different sentences with different consequences.
- A usable test for any leak, offered as the channel's rule: ask what access the claim would require. A parameter count needs training infrastructure access; a benchmark table needs eval harness access. If the claim needs access the leaker demonstrably lacks, the claim is the product.
- Precedent cited for restraint: a rumour wave earlier in the year priced a model code-named Spud as GPT-5.7; it shipped as GPT-5.5. Prediction markets also mispriced Astra's own release window until very late.
- Altman is quoted saying the brake is being applied deliberately - slowing things as needed so alignment work can be done - and separately that he expects an internal system by year end that he would personally call AGI. Internal is not shipped, and the presenter argues that gap explains why the buyable model feels like a side grade.
- Open weights are the quiet story: GLM 5.3 Flash is reported at around 7 cents per million tokens in, roughly 140 times cheaper than Astra. For high-volume moderate-difficulty work the presenter argues the frontier stopped being the right answer some time ago.
- One underrated signal, badged credible leak: OpenAI is reported to have shut down Sora and the Atlas browser this year and moved that compute to core models and agents. The argument made from it is sound - a company shelving its video product for compute reasons is not weeks from shipping native video.
Why it matters
This is the most rigorous single source the Handbook has taken in, and it is worth keeping as much for its method as its content. The badge system, the 'what access would this claim require' test, and the on-air correction of its own earlier wrong call are exactly the discipline Section 6 asks for, arriving from a YouTube channel rather than a journal. Substantively it settles two things: Astra is not omni and cannot take audio or video, which invalidates a large slice of current commentary; and the referee problem is now explicit, with the funder of a benchmark holding exclusive access to part of it. The Fermat disagreement with the Diamandis entry is a live conflict to carry forward, not a detail - two sources in one batch describing the same reported achievement incompatibly.
CN-H2WI3dHk-transcript.txt