Skip to content

Muse Spark 1.1 (xhigh)

Available

Meta · 2026-07-09 · 32,000 tokens

An AI model from Meta, suited to a broad range of AI workloads.

Supported modalities:textcode

Quick Overview

Text Generation5/10
Code Generation7/10
Reasoning6/10
Multimodal4/10

Benchmark Results

Scores from leading benchmark suites.

artificial analysis intelligence53.2
artificial analysis coding71.3

Performance Metrics

Latency and throughput performance.

P50 Latency
0tokens/sec

Dive Deeper

AI model analysis

Muse Spark 1.1 Review: Fast Coding Performance with Preview-Stage Risks

Muse Spark 1.1 Review: Fast Coding Performance with Preview-Stage Risks
Summary

- **Where it stands:** Muse Spark 1.1 ranks 22 of 578 on the Artificial Analysis Intelligence Index at 50.6 - **Price:** $2 per 1M blended tokens - **Speed:** 166.23 output tokens per second, 0.3s to first token - **Pick it when:** You need a fast, affordable model for coding agents, tool calls, and multimodal experiments - **Watch out:** Evidence remains insufficient for production reliability, sustained throughput, and emotionally consistent creative dialogue

01

Muse Spark 1.1 at a glance

Muse Spark 1.1 is a fast and capable coding-agent candidate, but its public-preview status and uneven non-engineering evidence limit a universal production recommendation. Meta presents Muse Spark 1.1 as a multimodal reasoning model for agents, tool calls, computer use, coding, and multimodal understanding in its official announcement. The model documentation lists text, images, video, and PDF as inputs, with text output. The supported API surface includes OpenAI-compatible Chat Completions and Responses APIs, plus Anthropic Messages compatibility, according to Meta’s developer guide. The model ID to use is muse-spark-1.1, while xhigh names a reasoning setting rather than a separate model, as described in the reasoning documentation. The supplied snapshot places Muse Spark 1.1 near the top of the measured intelligence and coding fields. The research brief does not establish comparable production reliability across those fields. Data provided by https://artificialanalysis.ai/.

02

The short verdict

Muse Spark 1.1 offers one of the strongest value positions in this snapshot for tool-heavy software work. Muse Spark 1.1 ranks 21 of 202 on the Artificial Analysis Coding Index and 22 of 578 on the Artificial Analysis Intelligence Index. Those positions support a practical recommendation for coding agents and structured engineering tasks, not a blanket claim that the model is best at every kind of conversation or autonomy. The comparison below uses the supplied Artificial Analysis snapshot, with qualitative product facts cross-checked against Meta’s documentation.

Nearby reference What changes if you choose it instead
Claude Opus 5, adaptive reasoning, low effort A more expensive alternative with a different reasoning profile and slower measured output in the supplied snapshot.
GPT-5.5, medium A higher-cost option with a similar coding position, making platform fit or task-specific quality the deciding factor.
Gemini 3.5 Flash, high A faster measured alternative at a higher blended price, useful if response speed matters more than token economics.
Gemini 3.6 Flash, high A newer higher-cost reference with faster measured output, but no supplied evidence of stronger coding placement.
GLM-5.2, max A near-price alternative with stronger intelligence placement and weaker coding placement, making workload mix important.

Product maturity is the main qualification. Meta currently describes the API as public preview and limits access by region in its developer guide. The official model page also presents model information whose title and listed model column are not perfectly aligned. That does not prove an imminent replacement, but it does make model pinning, regression testing, and fallback planning necessary.

03

What the performance ranking means in practice

Muse Spark 1.1 is a high-ranked coding model whose speed and tool support matter more in agent loops than raw benchmark position alone. A coding rank of 21 of 202 places Muse Spark 1.1 close to the front of the measured coding field. Developers should interpret that as a strong reason to test repository edits, debugging, terminal work, structured outputs, and tool-mediated implementation. It is not proof that every large codebase task will complete without review. The intelligence rank of 22 of 578 gives the model broader credibility for planning and reasoning, but it does not remove the need to test domain-specific instructions, failure recovery, or factual verification.

Meta says Muse Spark 1.1 supports parallel tool calls, streaming tool parameters, structured tool use, web search grounding, and context continuation in its developer guide. These features fit workflows where the model must inspect state, call a tool, interpret the result, and continue. The snapshot also reports 166.23 output tokens per second and 0.3 seconds to first token. That combination should feel responsive in interactive agent loops, although median measurements do not establish guaranteed service levels under production load.

Reasoning control creates an important implementation detail. The reasoning documentation supports settings from minimal through xhigh, while the Chat Completions documentation states that reasoning_effort="none" is unsupported. Teams that need lower latency or lower output spend must select an official reasoning level rather than assume reasoning can be removed. Chat Completions is stateless, while the Responses API can continue a server-side chain through previous_response_id, so long-running agents need an explicit state strategy.

Evidence remains insufficient for sustained throughput, large-project coding reliability, and consistent multimodal accuracy. Meta’s benchmark page does not provide complete evaluation methodology, and the available Hacker News discussion reports a repeated voice in application-generation outputs without offering an independent model-level statistic.

04

Where the price helps, and where it does not

Muse Spark 1.1 is cheap enough for broad experimentation, but output-heavy agent workflows can erase much of that advantage. The supplied snapshot gives Muse Spark 1.1 a blended price of $2 per 1M tokens, lower than each listed nearby model. Meta’s public pricing lists $1.25 per 1M input tokens and $4.25 per 1M output tokens in the developer guide. The practical implication is clear: workloads that send substantial context but produce controlled responses should benefit most from the price structure.

Agent workflows often produce more than a final answer. Planning, tool arguments, intermediate explanations, retries, structured summaries, and repair attempts can all add output tokens. A low blended price therefore does not guarantee a low cost per successful task. The relevant unit for procurement is a completed, accepted result. Teams should compare successful patch cost, valid tool-call rate, retry frequency, and human review effort before treating the token price as the deciding factor.

The price is especially attractive for coding-agent pilots, document triage, visual inspection, and internal automation where moderate failures are recoverable. It is less automatically attractive for a critical workflow if one incorrect action triggers expensive review or rollback. A more expensive model may still win on total task cost when it reduces retries, although the supplied evidence does not quantify that tradeoff for Muse Spark 1.1.

Access conditions also affect the business case. Meta describes the service as public preview and notes that availability can depend on the developer’s region in the Model API overview. The official pages do not provide an independent SLA or long-term availability commitment for Muse Spark 1.1. That makes the current price useful for experimentation, while production budgeting should include a fallback model and the operational cost of migration if access conditions change.

05

Who should choose Muse Spark 1.1

Muse Spark 1.1 deserves a pilot for coding agents, multimodal workflows, and interactive tools, not an unconditional production default. The best fit is a team that values coding performance, fast responses, structured tool calls, and low token cost, while retaining control over evaluation and fallback behavior.

Scenario Recommendation Reason
Coding agents and terminal automation Choose for a pilot The coding ranking is near the front of the supplied field, and Meta explicitly positions the model for coding, agents, and tool use.
Repository maintenance and debugging Choose with review gates The ranking supports serious testing, but the evidence does not prove reliable autonomy across every large project.
PDF, image, and video workflows Pilot The input modalities are officially supported in the model documentation, but task-specific accuracy evidence is limited.
Interactive IDE assistance Choose if output style passes review The responsiveness is attractive, while the Hacker News discussion reports repeated phrasing in generated applications.
Roleplay or emotionally consistent character dialogue Avoid as the default An early, non-standard Reddit test reported weaker role consistency and emotional understanding.
Long-running stateful agents Use the Responses API Meta documents cross-turn continuation through the Responses API, while Chat Completions remains stateless in the developer guide.
Critical production paths Wait or run a contained pilot Public-preview access, regional availability, and the lack of an independent SLA create operational risk.

The strongest adoption path is a narrow pilot with fixed prompts, representative repositories, multimodal fixtures, and explicit tool-call checks. Measure accepted code changes, recovery from failed tools, output-format validity, review time, and style drift. Treat the model’s strong ranking as permission to test seriously, not as a substitute for task-level evidence. The community evidence also has clear limits. The Reddit reasoning discussion describes hidden reasoning behavior and custom-tag experiments, but it provides no standardized quality metric. Those reports should shape test design, not determine the final decision alone.

06

Questions to answer before adoption

Muse Spark 1.1 requires a short operational checklist before teams commit it to a critical path. The key questions concern API identity, reasoning control, hidden reasoning, multimodal inputs, regional access, and fit for creative dialogue. Meta’s developer documentation resolves the main API semantics, while community reports provide directional warnings rather than standardized proof. The answers below separate documented behavior from evidence that still needs validation.

Frequently asked questions

Is Muse Spark 1.1 worth using for coding agents?

Muse Spark 1.1 is worth piloting for coding agents because its coding rank is 21 of 202, its tool-oriented positioning fits agent loops, and its measured response profile is responsive. Final adoption should depend on repository-level tests.

Does Muse Spark 1.1 support multimodal inputs?

Muse Spark 1.1 supports text, images, video, and PDF inputs, then returns text, according to Meta’s model documentation. Teams should still validate accuracy on their own files.

Can developers disable reasoning?

Muse Spark 1.1 cannot use reasoning_effort="none"; developers must choose a supported setting from minimal through xhigh, as described in the reasoning documentation and Chat Completions documentation.

Can developers access the model's raw chain of thought?

Muse Spark 1.1 does not expose raw chain of thought through the standard response; the cited community discussion reports access to a summary, while custom-tag requests did not clearly improve visibility.

Is Muse Spark 1.1 ready for production?

Muse Spark 1.1 is better treated as a controlled pilot than a universal production default because public-preview access, regional availability, and independent reliability evidence remain limited, as reflected in Meta’s developer guide.

Is Muse Spark 1.1 suitable for roleplay?

Muse Spark 1.1 is not the safest default for roleplay or emotionally consistent character dialogue because an early, non-standard Reddit test reported weaker character consistency and emotional understanding.

Sources

  1. Introducing Muse Spark 1.1Official model positioning, multimodal reasoning, agent tasks, coding, and computer use.
  2. Build with Muse Spark, now available on Meta Model APIAPI model ID, compatible protocols, pricing, tool support, reasoning settings, public-preview status, and regional access.
  3. Muse SparkOfficial model page, evaluation listings, product-page status, and missing evaluation methodology context.
  4. Model API OverviewModel API availability and operational overview.
  5. Model API ModelsSupported input and output modalities.
  6. ReasoningSupported reasoning-effort settings and parameter semantics.
  7. Chat Completions APIReasoning disablement behavior, token parameters, and Chat Completions constraints.
  8. Muse Spark 1.1 Testing for RPEarly roleplay testing, character consistency, emotional understanding, and test-method limitations.
  9. Reddit reasoning discussionCommunity observations about hidden reasoning and custom-tag experiments.
  10. GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 appsApplication-generation discussion and repeated output-style feedback.
  11. Artificial AnalysisData attribution for rankings, pricing, latency, throughput, and model comparisons.

Published: