Summary
A hands-on test of OpenAI's GPT Image 2.5 by Theoretically Media, run through the reviewer's standing set of comparison prompts. His verdict is in the title of the update rather than the video: it is a point-five release, and the image quality gains are modest. The finding worth keeping is elsewhere — with GPT-6 Astra behind it, the image model can now go and research before it draws, and the reviewer's working method shifts from writing better prompts to pointing it at references and arguing with the result.
- Two variants: Flare (fast) and Sunburst (heavy). Available to ChatGPT, GPT and Codex users, but OpenAI does not say which is the default in each — the reviewer infers Flare from behaviour. Both are exposed through the API with quality settings from low to max. Comparing Flare against Sunburst at maximum quality, he found the difference subtle enough to be closed by a creative upscaler on a low setting.
- The grain and noise pattern that dogged GPT Image 2 is reduced but not eliminated, and still shows up on cinematic prompts. It can be removed by running output through a denoiser, which is an extra step. Output in ChatGPT and Codex comes at 1672x941, so anything for real use needs upscaling.
- Why his wine-glass test matters, and it generalises: these are reasoning models rather than diffusion models, so they follow instructions that contradict their training data. Asked for a glass filled to the top and a clock reading 5:15, it delivers both — where a diffusion model like Midjourney regresses to what its training images show, which is under-filled glasses and clocks set to 10:10. The trade is that Midjourney is more beautiful and wrong in the details, while this is blander and right.
- The strongest demonstration is conversational editing. Given the finished image and asked to empty the glass, advance an hour, and leave evidence someone made bad decisions, it moved the clock forward correctly, added the empty bottle, and unprompted added a drunk text to an ex and a receipt for a life-size bronze goose at $2,800.
- The four-angle grid test — one 21:9 image split into four quadrants showing alternate camera angles on the same scene — held character consistency better than Nano Banana 2 or Nano Banana Pro. Astra then noticed the feet were cropped and produced a wider bonus generation without being asked.
- The behaviour that actually changes the workflow: it researches. Asked for a Times Square image with the reviewer's own billboards, it worked out who he was and included his real short film and website branding. Given GTA 6 screenshots and a poor first attempt, it went and found who plays Jason and Lucia and used the actors as photographic references, producing "a photograph that was never taken". Asked for the colour grade of Spielberg's War of the Worlds, it reproduced the blown skies, crushed blacks and halation.
- Where it still fails: the mirror test passed on reversing text but could not keep the reflection's pose or braid consistent with the subject. The compound prompt — pelican on a bicycle at 5:15 holding wine — broke down into a duplicate digital clock, flamingo-like legs and both pedals on the same side of the bike.
- The reviewer's practical conclusion is that default output is bland and the fix is not a better prompt but an argument: telling it the images feel like promotional comps rather than film stills produced a dramatic improvement with no extra specificity. Working through Codex or Astra so it can fetch references beats prompt-craft.
- Aside worth recording: he says Astra replaced Fable 5.1 as his daily driver within a weekend, and that he rebuilt his oldest AI short film — originally made in the Veo 2 era — in a single day using Astra and computer use.
Why it matters
The image model is the less interesting half of this. What the entry records is a change in how these tools are used: an image generator with a reasoning model attached stops being a prompt box and becomes something you brief. It looks up who you are, finds the actors, studies a film's colour grade, and notices a compositional mistake you did not mention. Prompt-craft as a skill is being displaced by reference-giving and iteration.
It also gives the Handbook an independent, hands-on read on GPT-6 Astra from someone with no stake in the benchmark argument running through DK-91 and DK-93. His verdict — first model in a long time to genuinely surprise him, while explicitly declining to 'drink the AGI Kool-Aid' — is a more useful data point than either the panel's enthusiasm or the leaderboards.
The reasoning-versus-diffusion distinction in the wine-glass test is the durable technical point and belongs in the Handbook's content-tools section: it explains why these models obey awkward instructions and why their output looks duller.
VsmL_KROmyI-transcript.txt