Summary
A daily briefing episode dominated by the Navier-Stokes credit dispute, followed by a roundup of four model and product releases. The host's standing thesis is the move away from picking one best model towards routing between models by job and cost. The reporting on the dispute is even-handed to the point of concluding nobody comes out of it well.
- OpenAI published a solution to the Navier-Stokes problem, one of the seven Millennium Prize problems, of which only one has been solved in the 26 years since the prize was established. It used an internal model described as significantly more capable than GPT-6 Astra. Noam Brown is reported as saying the result cost several million and took a week or two.
- The dispute, from NYU professor Tristan Buckmaster's published account: he and an Anthropic employee had worked on the problem for over a year, using several models including Codex, and had found novel solutions to related problems - stepping stones, not the prize itself.
- Buckmaster says he asked OpenAI whether the model had been trained on or had access to the Codex sessions holding all their drafts; he was told the model did not look up user data, and says he asked again about training and got no answer. He reports being offered co-authorship on condition his Anthropic collaborator was removed, and quotes the reply to his threat to go public as 'why would you ruin your career?' and 'if you don't want me to be nice, then I don't have to be nice.' This is one side's account; OpenAI's Sebastian Bubeck called the allegations false and inflammatory and published part of the message chain.
- OpenAI's position as quoted: researchers and agents did not see the work until it was public, no specific user data was accessed, but 'while unlikely, we cannot rule out that deidentified data derived from the usage of our products helped improve our models.' Mark Chen separates the two questions explicitly - no human or agent looked at user data for this effort; yes, deidentified data improves the products, as at every LLM company.
- The question the host flags as more relevant than the drama, quoting a former DeepMind employee: 'Can these labs see all your work and scoop you when the stakes are high enough?' A mathematician's follow-up about opting out of training while pasting a trade secret is reported as unanswered at recording.
- The strategic question underneath: whether labs sell the inputs to innovation or start selling the outputs - doing the drug discovery themselves and taking the patents rather than selling scientists the tools.
- A class action has been filed on behalf of Claude Max subscribers alleging the $100 '5x' and $200 '20x' plans do not deliver five and twenty times a $20 plan's usage, because of how five-hour and weekly limits are calculated. The host is unsure it goes anywhere but reads it as representative of scrutiny to come.
- ElevenLabs hired a CFO from Adyen and is reported to be exploring an IPO, on track for $600m annualised revenue by year end from $350m, reportedly profitable with more than half of revenue from large enterprise.
- Cognition raised $2bn at a $48bn valuation, up from $26bn in May, with revenue run rate reported going from $492m to almost $900m. It stays independent, and the host reads that as informed by watching OpenAI cut off access to Cursor customers after SpaceX's acquisition.
- Gemini 3.8 Flash: 73.7% on a coding benchmark against Opus 5's 74%, but terminal-bench 4.0 collapsing to 19.1% against Opus 5's 51.8%. Artificial Analysis scored it 59, seventh, then revised its index and it fell to twelfth. Still described as the undisputed speed leader and 'the cheapest we've measured at this level of intelligence'.
- Meta's Muse Spark 1.3 scored 68 on the coding agent index at max settings, tied first with Opus 5 at the time. SemiAnalysis's critique is the important part: Gemini 3.8 Flash and Muse Spark 1.3 are called 'two of the most clearly benchmaxed models we've seen yet', on the argument that labs need not train on public benchmark tasks directly because they can buy data from RL-environment startups designed to mimic them. The tell is failure to generalise from terminal-bench 2.1 to 4.0.
- Meta's Alexander Wang conceded the point in public: 'We don't claim Muse Spark 1.3 is as strong as Astra or Fable 5.1, but it is significantly more cost effective.'
- Meta launched Muse, a personal agent, with a security design worth recording: each instance in its own isolated VM, a separate system called Sentinel checking every action before anything leaves it, and the agent never seeing actual passwords or card numbers.
- The adoption worry about Muse is not capability but trust, quoted from A16Z's Olivia Moore: 'I was more reluctant to press the connect email button on Muse than on 10+ startup agent products I've tried.' Meta's distribution advantage cuts both ways.
- ChatGPT Images 2.5 shipped in two variants, Flare for speed and Sunburst for controlled editing, with a sketch input feature and a claimed 50% latency reduction. OpenAI reports users generating more than 3 billion images a week.
Why it matters
Two things here change the picture rather than adding to it. The first is the benchmaxing mechanism SemiAnalysis describes: a lab need never train on a public benchmark to beat it, because it can buy environments built to imitate it - which means a high score on a public benchmark is now weak evidence by construction, and the only useful signal is a benchmark too new to have been farmed. That is a sharper version of the Handbook's existing principle about benchmark rotation and should replace it. The second is the Navier-Stokes dispute, which is the first concrete instance of the question every professional user now has: whether work done inside a lab's product can end up benefiting the lab in competition with you. OpenAI's own careful wording - cannot rule out that deidentified data derived from product usage helped improve the models - is the part to keep, because it is the answer, and it is not no.
j90tdo5Tjes-transcript.txt