Slop TVNewsLatest
Explainers

World models guide: Genie 3, GWM and open weights, updated weekly

SLOP TV's running guide to world models: what runs in real time, what you can call today, the open weights, the hardware and SLOP TV's coverage.

Illustration: World models guide: Genie 3, GWM and open weights, updated weekly
AI-generated illustration for SLOP TV News with GPT Image 2, showing the Google DeepMind and Runway marks. Logos are trademarks of their owners.

Key takeaways

  • A world model simulates how an environment develops and how an action changes it, and Google DeepMind's Genie 3 announcement says that makes it possible to train AI agents in rich simulated environments.
  • Runway's GWM Worlds 2, shown as a research preview on 3 September 2026, generates continuous 720p video at 24 fps with 48,000 Hz audio and lets a text action be aimed at a subject or at the scene itself.
  • Runway Characters and Odyssey's Odyssey-2 Pro can be called through an API today, while Genie 3 and Runway's GWM Worlds 2 are research previews; Characters runs at 24 fps with 37 ms of model time per frame and 1.75 seconds of server-side turnaround.
  • Astronex-World 1.0, submitted to arXiv on 17 September 2026, is a 5B model with Apache 2.0 weights that streams 832x480 video at 24 fps on a single Nvidia L20 48GB GPU.

Updated 26 Sep 2026 · next update 3 Oct 2026

Key facts as of 26 Sep 2026

  • What one is: a world model is an AI system that simulates aspects of the world so an agent can predict both how an environment will develop and how its actions change it, which is how Google DeepMind describes the category. (source)
  • Genie 3: DeepMind's interactive world model generates a navigable environment at 720p and 24 frames per second from a text prompt and holds consistency for a few minutes, with its own limits listed as a small action space, weak multi-agent interaction and short sessions. (source)
  • Runway's GWM Worlds 2: shown on 3 September 2026 as a research preview, it generates continuous 720p at 24 fps with 48,000 Hz audio, is steered by text actions and camera motion, and has no preset session length. (source)
  • Runway Characters: built on Runway's GWM-1 general world model, it animates a single reference image at 24 fps with an effective 37 ms of model time per frame and 1.75 seconds of server-side turnaround from the end of speech to the first response frame. (source)
  • PixVerse R2: announced on 22 September 2026, it makes input persist so an action shapes what happens later in the same session, and PixVerse says R1, launched in January 2026, was the first real-time world model. (source)
  • World Labs: a spatial intelligence company whose first product, Marble, generates persistent 3D worlds from text, images, video or 360 panoramas, with an API listed alongside it. (source)
  • Odyssey: raised a $310 million Series B, describes its models as foundation world models and names Agora-2 among them, and its team comes from DeepMind, Waymo and Tesla; its API documentation offers Odyssey-2 Pro through real-time interactive streams. (source, API docs)
  • Open weights: Astronex-World 1.0, submitted to arXiv on 17 September 2026, is a 5B model built on the Wan2.2-TI2V-5B prior that streams 832x480 video at 24 fps on one Nvidia L20 48GB GPU and scores 73.5 on WBench Navi and 70.0 on WBench Full. (source)
  • Hardware: AWS's Neuron Science team and Reactor generated real-time video above 16 frames per second on a single Trainium2 chip, with one kernel dropping from five seconds to 1.8 milliseconds. (source)

What separates these systems is what you can do with them: Runway Characters and Odyssey's Odyssey-2 Pro can be called through an API today, while Genie 3 and GWM Worlds 2 are research previews. (Runway, Odyssey)

What a world model is

A world model is an AI system that predicts how an environment behaves. Google DeepMind's Genie 3 announcement puts it plainly: world models use an understanding of the world to simulate aspects of it, so agents can predict both how an environment will evolve and how their actions will affect it. That makes them both a product category and a piece of research infrastructure, because a world model is a place to train agents without a physical world.

The distinction that matters commercially is between generating a clip and running a world. A video model returns a fixed result: change something and you prompt again. A world model keeps generating while you are inside it, takes new input mid-session and holds its state. PixVerse's R2 announcement describes the difference from the inside: in earlier real-time generation an input affected the current moment and then passed, and in R2 an input updates the running state of the world and shapes what follows.

Three things have to be true at once for this to work, and they explain why the field is hard. The world has to stay consistent over time, the model has to run fast enough that a late frame does not break the experience, and the session has to persist. Runway's engineering post on Characters puts the frame budget at 42 milliseconds at 24 fps, before audio processing and network transit.

What world models cost

Costs split by whether you are running one, calling one, or building on one, and most of the field is not yet purchasable.

  • Research previews: Genie 3 and Runway's GWM Worlds 2 are shown rather than sold. Genie 3's announcement describes access as a limited programme, GWM Worlds 2 carries a research-preview label on Runway's own page, and neither publishes a price. (Genie 3, GWM Worlds 2)
  • APIs: Runway Characters is available through the Runway API and the Runway web and mobile apps, and Odyssey's API documentation offers Odyssey-2 Pro, a general-purpose world model, through interactive streams for real-time generation as well as viewable streams and batch simulations, with a developer account and API key. Both can be called today. (Runway, Odyssey)
  • Open weights: Astronex-World 1.0 is free to download under Apache 2.0 from the authors' release, and the model streams on one Nvidia L20 48GB GPU, so the cost is hardware and time rather than a rate card. (arXiv)
  • 3D output: World Labs' Marble generates and exports 3D worlds rather than a video stream, and the company lists an API beside the product without publishing a price. (World Labs)
  • Infrastructure: serving one at scale is a separate problem. Reactor's chief technology officer, Bryce Schmidtchen, told Amazon Science: "It's about real time, it's about low latency, and it's about doing that as efficiently as you can at scale." The AWS and Reactor work on Trainium is published research, not a priced product. (Amazon Science)

Previews with no published price are the normal state here, so any comparison of what a world model costs is really a comparison of access: an API you can call today, a preview with limited access, or weights you can run.

What world models do

They simulate rather than render. Genie 3 builds a navigable world frame by frame from a description and the user's actions, and DeepMind notes the consistency is emergent rather than the result of an explicit 3D representation like a NeRF or Gaussian splat. Actions include movement plus text-based events, which DeepMind calls promptable world events.

They keep state. PixVerse's R2 carries input forward, so a choice made early changes what happens later, and its example is a story encounter that branches live depending on what a player offers. Runway's GWM Worlds 2 turns the same idea into a directorial mode: a text action can be addressed to a subject or to the scene, so a campfire flares and a sunset fades to night in the same running world.

They generate sound as well as picture. GWM Worlds 2 generates audio at 48,000 Hz alongside the video, and a text action can be a spoken reply as well as a movement. (source)

They drive characters. Runway's Characters animates a single reference image at 24 fps with lip sync and head motion, with no fine-tuning, and 1.75 seconds of server-side turnaround, and Runway says Characters can join Zoom, Google Meet or Teams calls and respond in real time. (source)

They need the compute fixed. AWS and Reactor's work is a window into the real constraint: the teams used Rolling Forcing as a stand-in for autoregressive diffusion video, and the kernel-level rewrites took one operation from five seconds to 1.8 milliseconds and cache copies from 23 to 1.9 milliseconds per layer, to sustain above 16 frames per second on a single Trainium2 chip.

They are being made in 3D too. World Labs' Marble generates spatially consistent 3D worlds from text, images, video or panoramas and exports them in 2D and 3D formats, which is a different output shape from a streamed video world and a different set of customers.

Who world models are for

Today, four buyers. Researchers are the first audience, and they are the one group the previews actually serve: Genie 3's announcement frames the work around training agents, and the model was tested with DeepMind's SIMA agent navigating generated worlds. Robotics and autonomous driving are the second: Odyssey names both among the fields it wants world models to change, and Runway lists robotics and embodied-agent simulation among the uses for GWM Worlds 2.

Creators and interactive studios are the third, which is where Runway's interactive work and PixVerse's game engine sit. And infrastructure buyers are the fourth: Amazon's write-up ties the growth of world models to a growing need for hardware and infrastructure that can run them, and the AWS and Reactor teams describe their techniques as generalizing across models.

For a creator this week, most of these systems are not yet for sale. Two can be called through an API today, Runway Characters and Odyssey-2 Pro, Genie 3 and GWM Worlds 2 are research previews, and the open weights are a research artefact that runs on a 48GB card. What is worth watching is the direction of the numbers: real-time video above 16 frames per second on a single Trainium2 chip, and a 5B model streaming 832x480 video at 24 fps on one 48GB GPU.

The world models diary

Dated entries, newest first. Each entry is added once and never rewritten; corrections are made in a new entry.

26 Sep 2026

The hub opens. This first entry records the news the page starts from, all of it from the labs and the papers rather than from coverage of them.

On 22 September PixVerse announced R2, making input persist so that an action shapes a later moment in the same session, and said its Game Engine now runs on the model. Source: PixVerse.

On 17 September Astronex-World 1.0 was submitted to arXiv: a 5B controllable world model built on the Wan2.2-TI2V-5B prior that generates 832x480 at 24 fps, streams in real time on one Nvidia L20 48GB GPU, and scores 73.5 on WBench Navi and 70.0 on WBench Full. Source: arXiv.

Runway's engineering page for Characters, published on 4 May 2026, describes a real-time video agent built on its GWM-1 general world model that animates a single reference image at 24 fps with 37 ms of model time per frame and 1.75 seconds of server-side turnaround. Source: Runway engineering.

On 3 September Runway showed GWM Worlds 2 as a research preview, generating continuous 720p at 24 fps with 48,000 Hz audio and no preset session length. Source: Runway research.

On 5 August 2025 Google DeepMind announced Genie 3, its first world model to allow real-time interaction, at 720p and 24 frames per second with consistency for a few minutes, and listed its limits as a small action space, weak multi-agent interaction and short sessions. Source: DeepMind.

What changed this week

First edition, so there is no earlier edition to compare against. The page is built from DeepMind's Genie 3 announcement, Runway's GWM Worlds 2 and Characters pages, PixVerse's R2 post, the World Labs and Odyssey sites, the Astronex-World 1.0 paper and AWS's kernel write-up. The next edition will be measured against this one.

SLOP TV's world models coverage

Sources

  1. deepmind.google
  2. runway.com
  3. runway.com
  4. pixverse.ai
  5. worldlabs.ai
  6. odyssey.world
  7. arxiv.org
  8. amazon.science
  9. documentation.api.odyssey.ml