Earlier this month Runway published a research preview of GWM Worlds 2, which it describes as turning high‑fidelity video and audio generation into a real‑time interactive simulation. The company calls the approach an “autoregressive diffusion” model — autoregressive because it generates output across time, step by step.
A notable new feature is WorldPrompt, a proposed input format for specifying a generated world and the actions inside it. With WorldPrompt you can fix parts of a simulated environment (including the first frame) and provide a series of timestamped events; those events can be prompted in real time. WorldPrompt functions as a control layer for characters, cameras and environment elements, but it is a prompting mechanism rather than a scripting language — unlike games such as Minecraft or Roblox, GWM Worlds 2 does not expose a general scripting API or structured state.
To explore the implications of WorldPrompt we spoke with Kamil Sindi (Runway CTO) and Robin Kahlow (Principal Research Scientist for generative video and multimodal AI), and we include exclusive comments from Anastasis Germanidis (co‑founder & co‑CEO) drawn from a podcast with swyx and Vibhu.
Who’s building real‑time interactive world models?
Runway’s most recent funding round was $315 million in February, and media reports cite a company valuation of about $5.3 billion based on that raise. Runway’s first world model release, GWM Worlds, arrived last December. Other notable projects in this space include Google DeepMind’s Genie 3 (which also targets 720p at 24 fps), Odyssey‑2 Pro, and World Labs’ RTFM (Real‑Time Frame Model). All of these projects face limitations related to the complexity and latency demands of real‑time video and audio generation: for example, Google notes Genie 3 can currently support a few minutes of continuous interaction rather than many hours.
The central idea of WorldPrompt
WorldPrompt is intended to be a way to control objects and agents in a generated scene — similar in concept to game‑style control but expressed as prompts. According to Robin Kahlow, WorldPrompt lets you “control all the different subjects in the world”: NPCs can be instructed to approach, speak or act, cameras can be steered, and fine‑grained control over scene elements is possible.
Because WorldPrompt is prompting rather than scripting, it sacrifices explicit state control for flexibility: Runway positions promptable, on‑demand video and audio worlds as the benefit. At the same time, prompting has limits — Kahlow notes the research preview is not perfect and complex actions remain challenging; Sindi adds that more training data and larger models improve the model’s ability to follow instructions.
Turning a video model into a real‑time runtime
Runway’s long‑term promise is that world models will enable fully self‑generated, real‑time games and experiences. That ambition raises two major engineering challenges, per Kahlow: first, the model must not generate an entire clip at once but produce frame‑by‑frame (or few‑frames‑at‑a‑time) outputs; second, generation must be fast enough for real‑time interaction.
GWM Worlds 2 currently streams continuous 720p video at 24 frames per second with audio at 48,000 Hz. Runway reached this setup by fine‑tuning its foundational audio‑video model to the WorldPrompt format, post‑training it to operate autoregressively, and then making it real‑time through distillation methods. Anastasis Germanidis described this pipeline as starting from bidirectional diffusion (which generates whole video at once) and making it autoregressive so it can generate one frame or a few frames at a time. For distillation, they either compress a large model into a smaller one or reduce the number of diffusion denoising steps (for example from ~50 steps down to a handful), trading some quality for much faster inference.
The challenges of real‑time generation
Germanidis highlighted error accumulation as a fundamental issue for autoregressive models: generated frames are fed back into the model for the next step, so small errors can compound over time. Sindi pointed to the difficulty of handling effectively “infinite generations”: deciding which context to keep and what to discard without blowing up GPU memory. Another persistent limitation is long‑term memory: the model does not have perfect memory yet, which the team describes as an open research problem.
Beyond performance, world models must render plausible causal consequences of different user actions. Germanidis used a football example to show a training distribution bias: online video data contains more successful goals than failures, so a video model may render success more convincingly. For a world model, you want equally realistic counterfactuals — what happens if the user does A versus B — which remains a gap between current video models and goal‑oriented world models.
Runway runs automated verifiable tests for certain behaviors, but as GWM Worlds 2 is a research preview, Kahlow recommends users test the model themselves to find failure modes.
Applications beyond gaming: robotics and agents
While gaming is a clear use case, Runway points to robotics and large‑scale agent testing as important applications. Kahlow notes that simulated environments can help evaluate robot behaviors and allow thousands of simulated environments for agent training. In GWM Worlds there is no structured programmatic state for agents to read; agents perceive the generated pixels and audio the same way a real camera would.
Runway also sees potential for synthetic data generation to train agents, and Germanidis suggested combining reasoning/planning models with the diffusion head: reasoning models could plan scenes or actions, then a diffusion generator would render the pixels.
Technical and organizational context
- GWM Worlds 2 streams interactive worlds at 720p, 24 fps and 48 kHz audio.
- The development path involved fine‑tuning a base audio‑video diffusion model to follow WorldPrompt, post‑training it autoregressively, and then applying distillation (model size reduction or step reduction) to reach real‑time performance.
- Runway’s development trajectory moved from Gen‑1 (depth‑conditioned video) through Gen‑2 and Gen‑3 to Gen‑4.5 and world models; distillation and model serving infrastructure have been key investments.
What Runway’s leaders say
- Kamil Sindi (CTO) emphasized promptable, on‑demand video and audio worlds and the role of more data and scale in improving instruction following.
- Robin Kahlow (Principal Research Scientist) noted movement tends to be handled fairly reliably by the model but complex actions and long‑term interactions remain imperfect; he framed autoregressive conversion and speed as the two main engineering hurdles.
- Anastasis Germanidis (co‑founder & co‑CEO) described the bidirectional→autoregressive→distillation pipeline in detail, discussed distillation tradeoffs (fewer denoising steps versus model‑size distillation), and highlighted error accumulation and counterfactual generation as core research issues. He also discussed robotics use cases and an “interface world model” concept that renders software front‑ends as pixels and reacts to clicks, drags and scrolling in real time.
Snapshot and outlook
GWM Worlds 2 and WorldPrompt represent a research step toward interactive, real‑time world models. The prototype offers playable 720p@24fps worlds with synchronized audio, and the WorldPrompt interface aims to give fine‑grained, promptable control of characters and cameras. Major remaining problems include autoregressive error accumulation, limited long‑term memory, runtime latency and the generation of plausible counterfactual outcomes. Runway positions robotics simulation, agent testing and novel interface paradigms (pixel‑rendered interactive front‑ends) as early practical applications, and frames distillation and scaling as the short‑term technical path forward.
(Interviewees and quoted sources: Kamil Sindi, Robin Kahlow, Anastasis Germanidis; product: GWM Worlds 2 and WorldPrompt; sources: Runway announcements and the swyx & Vibhu podcast.)



