Meta Superintelligence Labs today announced Muse Spark 1.1, an updated multimodal reasoning model designed for agentic workloads. The release follows Muse Spark and is intended to improve tool and computer use, coding performance, and multimodal understanding.
Availability and integration
- Muse Spark 1.1 is available now in “Thinking” mode in the Meta AI app and on meta.ai.
- Meta launched the new Meta Model API in public preview alongside this release; developers can access Muse Spark 1.1 through this API.
- The announcement coincides with this week’s launch of Muse Image, which Meta says together with Muse Spark 1.1 advances its vision for “personal superintelligence.”
Key improvements
- Agentic tasks: Muse Spark 1.1 delivers notable gains on personal, goal-oriented tasks that require planning and orchestration across multiple external apps and services. It zero-shot generalizes to new native tools, MCP servers, and custom skills.
- Multi-agent orchestration: trained to optimize end-to-end latency, the model can act as a main agent—gathering context, planning, and delegating to parallel subagents—or as a subagent that adheres to its role and escalates when needed.
- Long context handling: the model actively manages a context window of 1 million tokens, retaining and retrieving actions and information from much earlier work and compacting content to preserve critical steps for later use.
- Computer workflows: Muse Spark 1.1 performs well across workflows that span multiple applications with changing information, maintaining context across extended sessions, adapting to evolving requirements and navigating unfamiliar interfaces with minimal human intervention.
Automation versus direct interaction
Rather than simulating every mouse click, the model learns when to automate (for example, by writing scripts) and when direct interface interaction is simpler. Training emphasized writing scripts when automation is faster, clicking when direct interaction suffices, and generating batches of actions at each step.
Practical examples
- Agentic dinner organization: the model detects new context that changes the task while placing an order and updates the plan without user intervention.
- Facebook Marketplace agent: from smartphone video, Muse Spark 1.1 extracts useful photos, reasons about the product and operates the user’s browser to create a Marketplace listing on their behalf.
Coding and debugging
- Coding performance: the model shows substantial improvements on real-world tasks involving large, complex codebases—diagnosing and fixing complex bugs, adding features in enterprise systems, and performing large code migrations.
- Demonstration in OpenCode: Muse Spark 1.1 builds a chat web app, takes automated screenshots to identify user-visible failures, traces issues to relevant code, implements fixes and validates them—combining coding, multimodal understanding and tool use.
- Agentic coding setups: the model adapts to diverse harnesses and handles multi-turn dynamics, supporting features such as planning mode, goal conditioning, subagent delegation and context compaction.
Multimodal perception and action
- The model excels at perception, multimodal reasoning and tool use: visual-to-code artifact generation, ultra-descriptive image and video captioning, and agentic workflow execution for multimodal scenarios.
- Muse Spark 1.1 can inspect visual and audio inputs, preserve details across long workflows and use those details while operating computers on a user's behalf.
Evaluation and usage
- Internal use: Meta engineers and researchers are using Muse Spark 1.1 daily. On Meta’s internal coding benchmark (Meta Internal Coding Bench), Muse Spark 1.1 significantly improves over Muse Spark and is competitive with leading alternatives.
- Research automation: researchers are leveraging the model to automate model development and evaluation tasks.
- DeepSWE/OpenCode: the model ran self-evaluations on a subset of DeepSWE tasks and generated an analysis dashboard of results.
Safety and risk management
- Pre-deployment safety tests: Meta conducted extensive safety evaluations following its Advanced AI Scaling Framework, which sets evaluations, threat models and deployment thresholds for advanced models.
- Frontier risk categories: across Chemical & Biological, Cybersecurity, and Loss of Control categories, Meta’s evaluations indicate Muse Spark 1.1 operates within safe margins.
- Robustness: the model shows resistance to direct jailbreaks and indirect attacks from untrusted data, prompt injection and developer-prompt attacks, and exhibits lower hallucination rates and reduced sycophancy.
- Full safety posture is documented in the Muse Spark 1.1 Evaluation Report.
Partner feedback
Early partners describe Muse Spark 1.1 as a complete agentic foundation, combining million-token context handling with strong coding and reasoning capabilities suitable for large-scale agentic workloads. Yashodha Bhavnani, VP of AI Products at Box, said Muse Spark performed competitively on Box’s enterprise evaluation set and highlighted its strengths in structured, procedural workflows across industries such as professional services, public sector and industrial operations.
Outlook
Meta frames Muse Spark 1.1 as evidence of ongoing research momentum. The company says it has even more capable models in training and expects to share further developments as they become available. Developers can begin building with Muse Spark 1.1 now via the Meta Model API public preview.



