Summary
Matt Wolfe's Friday roundup for the week ending 4 September 2026, covering a week in which FOUR frontier labs shipped state-of-the-art models: Claude Fable 5.1 (with Mythos 5.1, restricted to cybersecurity testers), Gemini 3.8 Flash, Meta's Muse Spark 1.3, and GPT-6 Astra. The model news is deliberately compressed because Wolfe made separate deep-dive videos on three of them. The substantive contribution of THIS entry is not the launch list but the argument he builds across it: that the two benchmarks he has relied on most have come apart from his own hands-on results, and he no longer trusts them. He says so explicitly — 'the benchmarks I've relied upon the most, I feel like I can't really trust even those anymore'. Note on sourcing: the automatic transcript garbles the benchmark names badly (one appears as 'Deep Seek', 'Deep Suite', 'Deep Sweet' and 'Deep Swell' in the same video — almost certainly SWE-bench; another as 'Busy Bench', 'Beauty Bench' and 'BU & C Bench'). The NUMBERS are consistent throughout and are recorded here; the benchmark NAMES should be verified against the video before being quoted anywhere.
- The week's four releases, in order: Claude Fable 5.1, Gemini 3.8 Flash, Meta Muse Spark 1.3, GPT-6 Astra. GPT-6 Astra was a limited rollout at recording — 'rolling out today to a limited set of organizations' — and Wolfe had early access rather than general availability.
- Fable 5.1 led Artificial Analysis on release and still led it at the end of the week, even after the other three shipped. Claimed scores: scientific research 52.6 (against Souls 22.4, Opus 29, Fable 24.7); Terminal Bench 55.8%.
- Anthropic's pricing claim versus Wolfe's measurement — the clearest example of the entry's theme. Anthropic states Fable 5.1 costs 'an estimated 25% less than Fable 5 for typical workloads'. On Artificial Analysis's own cost-per-task measure Wolfe reports Fable 5.1 as the MOST expensive model tested at $3.69 per task, against $3.14 for the model it is supposedly 25% cheaper than. Both figures are reported by Wolfe from third-party dashboards, not independently verified here.
- Anthropic also claims improved safeguards 'blocking 60% fewer false positives than before'. Reported, not tested.
- Gemini 3.8 Flash is the value story. Listed at $0.75 per million input tokens and $3.75 per million output, against Opus 5 at $5/$25 and Fable 5 at roughly $10/$50. It scored 73.7% on the coding benchmark where Opus 5 scored 74% — a 0.3-point gap — at $2.36 average cost per task against Opus 5's $11.84 and Fable 5's $21.63.
- THE CENTRAL OBSERVATION. Muse Spark 1.3 scored 75.4% on the coding benchmark, which would make it the best coding model released, and placed third on Artificial Analysis ahead of both GPT-6 and Gemini 3.8 Flash. On Wolfe's own two practical tests it was the worst of the four: 20th on the code-generated-image benchmark, and its game-clone test produced 'a cube with a little cylinder and I'm shooting other cubes' against recognisable characters and environments from the others. This is a first-hand result and it is the reason for the entry.
- GPT-6 Astra reportedly saturated ARC-AGI 3 at 99.9%, which Wolfe reads as retiring the benchmark: 'This benchmark is useless now at this point.' It scored 74.1% on the coding benchmark — below Muse Spark's 75.4 — while placing first on his code-generated-image test and producing what he judged the best game clone.
- A cost datum worth keeping for anyone budgeting agentic coding: Fable 5.1 worked about two hours on the game-clone task, exhausted the daily allowance of Wolfe's $200/month Anthropic 20x plan, ran into overage, and cost him roughly $120 for one generated game. The same task on Gemini 3.8 Flash via Cursor used 'a very, very small percentage' of his credits.
- Nvidia's acquisition of Hugging Face is CONFIRMED this week, having been rumour in the previous week's episode. Wolfe's reading — his interpretation, not company statement — is that it is a bet on open-weight models: as Meta, OpenAI and Google build their own silicon, Nvidia's growth case shifts to enterprises and individuals running open models on their own hardware.
- Atlas, from World Labs, generates a navigable three-dimensional environment from one to a few dozen still images with camera control, rather than generating video. Early access; Wolfe had not used it at recording.
- Runway's Solaris allows real-time manipulation of a generated scene — dragging objects, swapping clothing — with lighting and shadow updating in response. Built on Gen 4.5, treating clicks and drags as conditioning for the next frame. Early access, not tested by Wolfe.
- MiniMax released a video model generating 15 seconds of video in 13 seconds — faster than playback — which has produced continuously generating interactive streams (fal.live, Peter Levels's infiniteslop.ai, reportedly 37,000 concurrent viewers on one day). Wolfe's own verdict is sceptical: 'just because we can, does that really mean we should?'
- Two transcription models in three days: Meta's Muse Voice Transcribe (1 September, aimed at streaming/real-time) and Microsoft's MAI Transcribe 2 (3 September), the latter claimed as fastest, most accurate and cheapest. Vendor claims, uncompared.
- Policy and legal, both worth flagging for the Handbook's ethics thread: ChatGPT conversations are not privileged and are reachable in court; and New York's Mayor Mamdani has barred AI use in school for K-8 students — not a ban on children using AI, a ban on its use during schooling, on the reasoning that fundamentals should be learned before the tool.
- Smaller items: ChatGPT can finally connect more than one Google account; Gemini voice interaction arrives in Gmail, Docs and Keep for paid AI Plus/Pro/Ultra tiers; OpenClaw shipped 2.0, which Wolfe has not used because Codex and Cursor replaced it for him. Closing gadget: Dyson's Cam Jet, a $500 camera-guided AI toothbrush that water-jets between teeth while brushing.
Why it matters
This is the entry to reach for when a benchmark score is offered as evidence. A single video captures the decoupling cleanly and from one consistent tester: the model topping the coding leaderboard produced the weakest practical output of the four, and the model whose vendor claimed a 25% cost reduction measured as the most expensive per task. That is not a claim about which model is best — it is a caution about the instrument, and it comes from someone who says he has relied on these two benchmarks more than any others and is now abandoning them. Treat the specific scores here as reported rather than verified; the durable content is the method, which is that a hands-on task you can judge with your own eyes caught something the leaderboards did not. Secondarily, this is the entry that confirms the Nvidia/Hugging Face acquisition previously carried as rumour, and it dates four frontier releases to a single week — useful context for how compressed the release cycle had become by September 2026.
GfPZm9yucQo-transcript.txt