Summary
AI Samson's survey of what people are doing with GPT-6 Astra and OpenAI's new Images 2.5 model, published 11 September 2026, a week after the Astra rollout began. READ THIS AS A CURATED FEED, NOT AS TESTING. The large majority of the examples are other people's demos collected from social media and replayed with commentary — he credits the creators and links them — so almost nothing here is independently verified, and demos posted by their own authors are a selected sample by construction. A minority ARE his own hands-on work and those are marked as such in the points below; they are the parts worth weighting. The video is also sponsored by Artlist, and the extended segment praising Artlist's interface for Images 2.5 is paid placement. SOURCING WARNING: the automatic transcript mangles product names throughout — 'GPT-6' appears as 'GBT6', 'GBG6' and 'GB6'; 'Nano Banana 2' as 'Nana Banana 2'; 'Navier-Stokes' as 'Navia Stokes'; and at one point 'OpenAI' appears as 'Pinei'. Creator names are unreliable. Verify before quoting.
- THE BIGGEST CLAIM IN THE VIDEO, AND IT IS CONTESTED. OpenAI is said to claim a solution to the Navier-Stokes Millennium Prize problem — roughly 90 years unresolved — produced by a group of agents running on an unreleased model 'significantly more capable than GPT-6 Astra'. The controversy is reported alongside it: the effort reportedly began from a rumour that two mathematicians were close to cracking it; OpenAI states nobody saw their work but adds that it cannot rule out that data from people using its products helped train the model; and one of those mathematicians (transcript: 'Tristan Bookmaster', very likely Tristan Buckmaster) has said publicly that he is furious about how it was handled. AI Samson explicitly declines to adjudicate — 'nobody outside of those rooms knows yet'. File as a live dispute, not as a result.
- Real-time video understanding is the capability thread running through the first half, and the priced example is the useful one: player and ball detection with possession tracking at roughly 20 cents per second of video. At that rate a 90-minute match is about $1,080, which is the number that decides whether any of these applications are real.
- Video-analysis demos shown, all third-party: warehouse tracking of every person and box turned into structured data to identify theft; retail and traffic monitoring; sports technique analysis; an app overlaying muscles, tendons and bones on a moving body. The surveillance implications are raised only as use cases — there is no discussion of consent or misuse anywhere in the video, which is itself worth noting.
- A robotics demo (third-party) in which Astra was given a robot arm, a paintbrush and a camera, and improved its brushwork across iterations by looking at the result. AI Samson's framing is the substantive claim: 'this is not a generated image, this is physical paint corrected by looking' — a perception-action loop achieved by a general model rather than a robotics-specific one.
- 3D reconstruction from ordinary footage: a home studio rebuilt in Blender from five photos and three panoramas; a short video clip converted into an editable Blender scene with characters, set and camera moves; a 3D reconstruction of an air-crash site built from newly released NTSB footage. The reverse path — from finished video back to editable scene — is the genuinely new direction here.
- Game generation, the section with the most examples and the least verification: a playable Zelda: Ocarina of Time recreation claimed at 60 seconds, a browser-based Need for Speed clone, a GTA 3-style open world (with visible physics failures, a car driving through a barrier), and a Miami open-world game that took its author about 90 hours. Note the spread — 60 seconds to 90 hours — and that every duration is self-reported.
- A systems-engineering example that is more interesting than the game demos: a user had Astra tune their Mac to raise Age of Empires from about 8 frames per second to as much as 150. Whether or not the figures hold, it points at the genuinely new thing — specialist work that previously needed expert knowledge becoming a prompt.
- IMAGES 2.5 ships in two variants: 'flare' for speed and high volume, 'sunburst' for precision and detailed multi-turn edits. Output up to 4K.
- The real advance in Images 2.5 is edit stability, not image quality. Changing one element of an image while leaving everything else untouched was previously unreliable — the subject would shift between generations — and the comparison shown against GPT image 2 is close to imperceptible change. That is what makes consistent stop-motion, sprite sheets and brand mockups possible, and it is a more consequential improvement than realism.
- HIS OWN WORK, so weight these above the rest: magazine-cover treatments holding layout exactly while changing context; a barber scene where an apron's colour and pattern were altered and the change tracked correctly into the mirror reflection; UI screens converted light-to-dark with the design preserved; localisation of UI into languages needing more space; dense magazine layouts where the body text is coherent English rather than filler.
- HIS OWN COMPARATIVE TEST, and the most honest moment in the video: a poker-table prompt specifying exactly what each of three players does with their hands. NO model tested completed it correctly — Images 2.5 sunburst gave a figure two right hands, Flux gave a player five cards instead of two. He notes Nano Banana 2 produces sensible headlines but gibberish body text at small sizes. Hands and fine text remain unsolved.
- A workflow point he returns to and that generalises beyond these tools: designing the interface with the image model FIRST and only then asking Astra to build it produces better results than asking Astra to program the app directly. His reasoning is that the image model is better at the aesthetic judgement than the coding model is.
- Scientific applications listed at the close — protein design in minutes, lowering the barrier to complex scientific software — are asserted without sources and should not be carried forward without them.
Why it matters
Useful as a capability census rather than as evidence: it captures, in one place and on a dated day, the range of things people believed GPT-6 Astra could do a week after launch. Treat it that way and it is valuable; treat any individual demo as established and it will mislead, because these are self-published successes with no failure rate attached and no independent replication. Three things do carry real weight. The 20-cents-per-second figure for video analysis is the kind of number that decides whether a use case is a business or a demo. The edit-stability improvement in Images 2.5 is a genuine capability change with a clear before-and-after, and it unlocks consistency-dependent work — animation frames, sprite sheets, brand systems — that was previously out of reach. And his own poker-hand test is the entry's best contribution precisely because it failed on every model he tried, which is the sort of result nobody posts. The Navier-Stokes claim needs watching rather than recording: if it stands it is among the most significant results in the field, and the training-data question attached to it is unresolved and serious.
eUFdtZLDOo8-transcript.txt