Summary
A Bloomberg profile interview with Fei-Fei Li — the Stanford computer scientist who built ImageNet and is widely called the 'Godmother of AI' — about World Labs, the startup she co-founded in 2024 to build world models rather than language models. Li's argument is that spatial intelligence, not language, is the next frontier, and that the industry's focus on LLMs leaves out most of what intelligence has to do. She frames this explicitly as a complement rather than a rival: 'not about anti-LLM. It's about the next frontier.'
- Li defines a world model as three distinct functions, and says the term is now overloaded: RENDERING, which outputs pixels for humans to look at (she names OpenAI's Sora as the example); SIMULATION, which captures the geometric structure of the world for machines rather than people; and PLANNING, which tells a robot what to do next. The three are usually conflated in coverage.
- World Labs' product is Marble, which generates an explorable, editable 3D world from a single image or text prompt. Named users are virtual production in film, game developers, and an NVIDIA collaboration using Marble environments to augment robot training.
- The numbers Li gives: World Labs has raised $1bn and has around 50 people. She puts total investment in world models across the field at $3bn and growing, and says humanoid robotics funding at $6bn is 'too small' set against what self-driving and language models absorbed.
- Asked whether world models are where chatbots were in 2019 — everyone chasing it, nobody having cracked it — Li agrees, and says the field is 'a lot earlier compared to LLMs' and has not yet agreed on how to build them.
- Her policy position is specific and repeated: root regulation 'in science, not science fiction'. She argues that Silicon Valley conversation about human extinction and AGI overlords actively distracts from real policy work, and that the substantive asks are resourcing the public sector and STEM education.
- The historical anchor: ImageNet was 14 million images across more than 21,000 categories, built from 2006. The 2012 turning point was Geoffrey Hinton's Toronto team entering AlexNet, running on NVIDIA GPUs — the combination of large data, neural networks and GPU compute that she calls the golden recipe for modern AI.
- On risk she is neither dismissive nor alarmed: mis- and disinformation, robots weaponised by bad actors, and students using the technology 'as a lazy crutch instead of an empowering learning tool', which she says she worries about a lot. She sets these alongside electricity, cars and the internet as technologies that could also have gone wrong.
- Asked about AI CEOs accused of God complexes, she says it is 'dangerous for any individual to think they know better than everybody else' and that it is not a founder's position to make every decision for people — while declining to name anyone.
Why it matters
This is a well-credentialled dissent from the assumption that scaling language models is the road to general capability, and it comes from the person whose dataset started the current era. Two things are worth carrying. First, 'world model' is not one thing — rendering, simulation and planning are different products with different customers, and coverage that treats a Sora demo and a robotics planner as the same advance is not tracking anything real. Second, Li's own framing is that the field is pre-breakthrough and has not settled on an approach, which is the opposite of how world models are currently being sold. Set against the same week's Musk and Gates material, hers is the only voice arguing the policy conversation is being actively harmed by extinction talk.
ITxsc3mgqts-transcript.txt