Meta's engineers measured Muse Spark 1.3 at 75.4% on DeepSWE v1.1, 88.8% on Terminal-Bench 2.1, and 98.5% on long-context retrieval inside a 1 million token context window. Those numbers travelled fast. The part that did not travel as fast: Meta lost every agent row on its own chart, the numbers come from a variant still in limited preview, and the comparison against its own predecessor changes the reasoning setting mid-table.
That gap - leading on coding benchmarks while trailing on agent benchmarks - is the most important thing to understand about Muse Spark 1.3 before you route any production agentic workload through it.
What Muse Spark 1.3 actually shipped
Muse Spark 1.3 is an agentic coding model Meta released on September 2, 2026. The model's weights remain closed, meaning it cannot be self-hosted. It is accessible only through Muse Code and Meta's API, or through third-party API aggregators. Pricing is unchanged from Muse Spark 1.2, at $1.25 per million input tokens and $4.25 per million output tokens.
Meta's engineers measured it finishing coding work with roughly 20% fewer tool calls and 25% fewer tokens than Muse Spark 1.2. That efficiency improvement is real and worth paying attention to: for long-horizon coding tasks where tool calls compound, fewer round-trips directly lowers latency and cost.
Muse Code is Meta's terminal-based coding agent, launched in beta alongside Muse Spark 1.2 in early August 2026, giving Meta a product to compete directly with Anthropic's Claude Code and OpenAI's Codex CLI.
Meta's first closed, directly monetised frontier line is shipping on a four-week cadence and pressuring every rival on price.
There is also a pricing tier most coverage skipped. The Contributor tier costs $0.10 per million input tokens in exchange for permission to train Meta's future models on your prompts and completions. That is 92% cheaper than Standard. It is also a meaningful data-governance decision, and teams handling customer conversations or proprietary code should read it before opting in.
Where Muse Spark 1.3 loses: the agent benchmark rows
Muse Spark 1.3 leads on narrow coding benchmarks such as DeepSWE v1.1 and Terminal-Bench 2.1. It trails Anthropic's Opus on the GDPVal-AA v2 agent benchmark and trails GPT-5.6 on DeepSearchQA and the Agentic IF Index, indicating a model tuned specifically for coding rather than general agentic breadth.
This distinction matters more than it sounds. Coding benchmarks measure a closed loop: given a repo and a bug, write a fix. Agent benchmarks measure something messier: planning across ambiguous steps, deciding when to call a tool versus reason it through, recovering from unexpected outputs. Those are the properties that determine whether a model is useful as the reasoning core of a multi-step agent, not a one-shot code completer.
On pure agent benchmarks, Opus 5 (max) still leads on four of six. And the headline numbers come with a catch: Meta's scorecard compares 1.3's max mode (still gated behind safety testing at launch) against 1.2's xhigh mode, so part of the jump is a reasoning-tier change, not a clean generational leap.
DeepSWE, released by Datacurve in May 2026, was built specifically to address contamination in older benchmarks. Its 113 tasks were written from scratch, with fixes not sourced from public pull requests and not merged into any repository, so the solutions could not have appeared in any model's training data. That matters: a high DeepSWE score is harder to game than SWE-bench Verified, where contamination has been an open problem. Muse Spark 1.3's coding scores are probably honest. The agent scores are where you should press harder before committing.
The open-weights question that does not have an answer yet
Teams that chose earlier Meta models partly because Llama weights could be self-hosted have a real decision to make. Muse Spark 1.3 is not currently an open-weights model. Meta has said that a Muse Spark open-weights release is coming soon, but no release date, model variant, or license has been confirmed.
The pattern is worth naming clearly: Meta has a documented track record of shipping an adjacent open-weight release (first Muse Glimmer) while the frontier model's weights remain closed. Muse Glimmer at 30 billion parameters is a genuinely useful open-weight model for local agentic tasks, but it is not Muse Spark - it is a smaller, distilled version of it.
Any infrastructure decision that depends on downloadable weights should be planned around what is actually available - currently, the API-only variant - and should treat future open-weight releases as upside rather than baseline.
This is not a knock on Meta. It is a planning reality. If your team's AI strategy requires weights you can audit, run on-premises, or fine-tune without a vendor agreement, Muse Spark 1.3 does not satisfy that today. Muse Glimmer 30B, which is open, might - but it benchmarks well below the frontier model.
The build-versus-buy pressure this creates is real. According to the McKinsey State of AI 2026 report, 32% of organizations have decided against buying off-the-shelf software, opting instead to build their own solutions using agentic coding tools. Many of those teams are betting on coding agents to reduce SaaS spend - and then finding the coding agent itself requires a SaaS commitment to access the best model. The share of companies attributing any EBIT impact to AI held flat at 37%, and only 6% clear McKinsey's high-performer bar of at least 5% of EBIT from AI. Swapping one subscription dependency for another while the profit line stays flat is not the escape route it looks like.
A teammate like Beagle can surface that scorecard comparison directly inside a Slack thread, without anyone needing to leave the conversation to dig through launch posts. The decision still takes judgment - but it should at least start with the right table.
Muse Spark 1.3 agentic coding: common questions
Does Muse Spark 1.3 beat Claude Opus 5 on agentic tasks?
On coding-specific benchmarks, yes. Muse Spark 1.3 scores 75.4% on DeepSWE v1.1 versus Opus 5's 74.0%, and ties GPT-5.6 Sol on Terminal-Bench 2.1 at 88.8%. On general agent benchmarks, Opus 5 wins four of the six rows in Meta's own comparison table.
Can I self-host Muse Spark 1.3?
No. As of September 2026, Muse Spark 1.3 is API-only, available through Muse Code and the Meta Model API. Mark Zuckerberg has confirmed that open weights for the Spark line are coming, but no date, variant, or license terms have been announced.
What is the Contributor pricing tier, and should I use it?
The Contributor tier prices input tokens at $0.10 per million - about 92% cheaper than the Standard tier. In exchange, Meta receives the right to train future models on your prompts and completions. Teams handling customer data, proprietary code, or regulated information should treat this as a data-governance decision, not just a cost decision.
How does Muse Code compare to Claude Code for real agent workflows?
Both are terminal-based coding agents. Muse Code runs on Muse Spark 1.3 and leads on long-context retrieval and autonomous code-fix tasks. Claude Code runs on Opus 5, which leads on broader agent benchmarks involving planning and tool-use decisions outside of pure coding. The practical answer: evaluate both on your actual task distribution, not the headline benchmark.
Is the 20% fewer tool calls improvement meaningful in practice?
For long-horizon coding tasks, yes. Fewer tool calls means lower latency, lower cost per task, and fewer points where an agent can go wrong. The caveat is that Meta measured this on coding tasks; the improvement may not transfer directly to agent workflows involving non-coding tools like search, APIs, or structured data lookups.