Model launches

Meta launches Muse Image and previews Muse Video—agentic models for media creation

Meta Superintelligence Labs released Muse Image, an agentic image-generation model that uses search, code, and self-refinement to improve factual accuracy and edit precision, and previewed Muse Video, a high-fidelity text-to-video model.

Meta launches Muse Image and previews Muse Video—agentic models for media creation

Meta Superintelligence Labs has released Muse Image, an agentic image-generation model, and offered an early preview of Muse Video, a text-to-video model built on the same pretraining base. Meta describes both systems as capable of invoking tools such as search and code, self-refining outputs, and combining reasoning with media generation.

Availability and core capabilities of Muse Image

According to Meta, Muse Image is their most advanced image model: it follows instructions precisely, performs targeted edits, composes content from multiple reference images, and leverages Instagram for social context. The model supports agentic tool use and is integrated with Muse Spark to enable shared tooling and joint planning between models.

Availability:

  • Muse Image is available now in the Meta AI app and on meta.ai.
  • It is available in Instagram Stories in the United States.
  • It is available on WhatsApp in limited countries.
  • It will be coming soon to Facebook.

Agentic tools: coding, search, and self-refinement

Rather than mapping prompts directly to pixels, Muse Image acts like an agent: it calls external tools to increase accuracy and refines its own outputs. Meta highlights three tool-driven behaviors:

  • Coding: during reinforcement learning the model learns to write and execute code that produces accurate plots or QR codes; it can condition on rendered figures to improve generated images. Combined with Muse Spark, code plus media generation enables animated GIFs, web pages with embedded images, and interactive visual games.

  • Search: Muse Image can perform web searches to ground its images with factual, real-time information and visual references. Enabling search improves factual accuracy on knowledge-intensive prompts, especially those involving current events or real-world facts.

  • Self-refinement: the model reflects on its chain-of-thought and improves its own outputs. This can take the form of a small local edit, a full regeneration, or switching tactics (for example, invoking a tool) when that yields higher reward. Meta says this behavior emerged during reinforcement learning because self-refinement produced better images.

Test-time compute and scaling behavior

Muse Image improves as it ‘‘thinks’’ more at inference time: allocating additional test-time compute increases reasoning steps, tool calls, and self-refinement iterations, which raises human-preference Elo scores. Meta reports an approximately log-linear relationship between combined compute and quality. They note that how the token budget is spent matters: Best-of-N (BoN) strategies that generate multiple variants and pick the best improve quality quickly but saturate, while spending the same compute on deliberate reasoning scales better. Combining reasoning with tool use yields complementary gains: tools provide external references or precise computations that reasoning alone cannot.

Editing, composition, and coherence

Muse Image performs precise edits requested by users and maintains coherence across multiple editing turns to support iterative refinement and brainstorming. It can compose elements from many input reference images—including people, objects, clothing, styles, and environments—and supports interleaving text and images inline in prompts for complex compositions.

Meta reports that Muse Image ranks No. 2 on Arena for text-to-image, single-image editing, and multi-image editing according to human-preference Elo at the time of writing.

Muse Video preview

Muse Video, built on the same pretraining base as Muse Image, is presented as delivering strong visual fidelity with native audio support. Meta says it is investing in areas where performance gaps remain—such as audio–video synchronization and physically accurate fast motion. Muse Video will be released soon to creators and in Meta AI. On Arena, Muse Video ranked No. 3 in human-preference Elo for text-to-video at the time of writing.

Content provenance: Content Seal and detection preview

To help verify whether an image was AI-generated, Muse Image includes Content Seal, Meta’s invisible watermarking system. Images produced by Muse Image in the Meta AI app and on meta.ai carry a hidden provenance signal that Meta says survives cropping, compression, resizing, and screenshots. The company plans to extend Content Seal to video in the future. Meta is also previewing a detection tool that allows users to check whether an image contains a Content Seal watermark as an initial way to indicate if an image was produced with Meta AI.

Integration with the Meta ecosystem and use cases

Muse Image is integrated with Meta’s ecosystem: combined with social tools in Meta AI, users can create images collaboratively or reimagine Instagram photos. Meta highlights use cases such as marketing assets for small businesses, personalized presets on Instagram, and dynamic content generation across Meta products.

Closing

Meta positions Muse Image’s agentic combination of coding, search, and self-refinement as key to improving accuracy and flexibility in media generation, while Muse Video represents a next step toward higher-fidelity video generation. Both models will be further developed and rolled out across Meta products over time.