Meta’s Muse Spark 1.3 improves agentic and coding scores, but its low per-task cost may matter more than its still-unfinished frontier challenge.

Meta has released Muse Spark 1.3, positioning the latest model in its rapidly expanding series as a closer competitor to offerings from Anthropic and OpenAI. The model’s strongest gains appear in agentic work, coding, and professional task benchmarks, while its pricing gives developers a reason to consider it even though the most capable version is not broadly available.
The release is split across two performance tiers. The xhigh version is available through Muse Code and the Meta Model API, while the more powerful max version remains a limited partner preview pending additional safety testing, according to The Decoder’s report citing Artificial Analysis. Meta has also said an open-weights version is coming, although the evidence does not specify a release date, licensing terms, or whether it will match the preview model.
That availability gap is central to the story. Meta’s best benchmark results come from max, but most developers can currently access only xhigh. As a result, the model’s competitive position depends not only on test scores but also on which tier customers can deploy, what the max version will cost, and whether its safety evaluation expands access.
Muse Spark 1.3 is the fourth model in the series in five months. The series launched in April, followed by versions 1.1 in July and 1.2 in August, according to The Decoder. The short release cycle reflects how quickly leading model providers are iterating on reasoning, tool use, and software-development performance.
Meta has kept the standard token prices unchanged at $1.25 per million input tokens and $4.25 per million output tokens, according to the report. However, the model is more expensive to use than Muse Spark 1.2, which cost $0.40 per task in the comparison cited by The Decoder. Version 1.3 comes in at $0.55 per task on the same index-based estimate.
That figure is the release’s clearest commercial differentiator. Artificial Analysis found that no model scoring 59 points or more on its Intelligence Index was cheaper than Muse Spark 1.3. Comparable rivals in that performance range cost between $0.94 and $1.23 per task, the report said.
The comparison should not be treated as a universal production cost. Per-task economics vary with prompt length, output length, tool calls, retries, context windows, and application architecture. Still, the pricing gap could be meaningful for teams running high-volume coding agents or workflow automation, where inference costs can quickly exceed the cost of the underlying software.
Artificial Analysis gave Muse Spark 1.3 a score of 62 for max and 61 for xhigh on its Intelligence Index, up from 57 for version 1.2 and 53 for version 1.1. The increase is partly explained by the index’s weighting: GDPval-AA v2 accounts for 20% of the score, Terminal-Bench 2.1 for 16%, and τ³-Bench Banking for 14%.
Those are also the areas where Meta’s latest model improved most. In τ³-Bench Banking, which measures an agent’s ability to operate tools in a simulated banking environment, max scored 52%, the highest result reported in the comparison. The available xhigh tier scored 47%, tying Claude Fable 5.1 in its max configuration and GLM-5.3-Flash rather than leading them.
The model also improved on Terminal-Bench 2.1, a test of coding through a terminal. Xhigh rose from 80% with version 1.2 to 85%, while max reached 86%. Claude Fable 5.1 remained ahead at 91.4% in max, 91.0% in xhigh, and 89.9% in high, according to the cited Artificial Analysis results.
On GDPval-AA v2, which covers 220 professional tasks and uses a human-expert reference score of 1,000, Muse Spark 1.3 increased from 1,615 to 1,709 for xhigh and 1,754 for max. Claude Fable 5.1 max scored 1,853. The report also noted that max used 62% more reasoning tokens than xhigh, suggesting that part of its advantage comes from additional inference compute rather than an across-the-board capability change.
The available evidence supports a more measured conclusion than Meta’s claim that Muse Spark 1.3 rivals or catches Anthropic and OpenAI. Several tests show substantial progress, but the model does not lead consistently across the reported evaluation set.
On GPQA Diamond, an expert-level science benchmark, Muse Spark rose from 90% to 94%. That placed it near the top group, but below Gemini 3.8 Flash high at 95.3% and Grok 4.6 high at 94.9%. On CritPt, a research-physics test, the model improved from 18% to 26%, while GPT-5.6 Sol max scored 32.3% and Claude Fable 5.1 xhigh scored 31.1%.
Two reported measures moved backward. AA-LCR fell from 83% to 79%, while factual accuracy on AA-Omniscience declined by as much as three points. The Decoder attributed that decline to more frequent refusals when the model was uncertain. For enterprise deployments, that trade-off matters: a model that avoids unsupported answers may reduce some risks, but excessive refusals can also disrupt workflows that require useful escalation or clarification.
These are independent benchmark results reported by Artificial Analysis, not proof of real-world superiority. The broader claims that Muse Spark 1.3 competes with leading Anthropic and OpenAI models come from Meta and were echoed in coverage from Yahoo Finance Canada, Digital Trends, SiliconANGLE, and VentureBeat. VentureBeat’s headline also highlighted that the strongest results rely on a model developers cannot yet broadly use. Neither Meta nor Artificial Analysis has published a price for the max tier in the evidence available here.
For AI product teams, Muse Spark 1.3 looks most relevant where tool use and cost control matter more than absolute leadership on every knowledge or reasoning test. Coding assistants, terminal-based agents, and structured business workflows could benefit from the reported gains in Terminal-Bench and τ³-Bench Banking, particularly if xhigh is available through Meta’s API at the stated token rates.
The practical evaluation burden remains high. Teams should test task completion, recovery from tool errors, refusal behavior, latency, and the number of reasoning tokens consumed—not just an aggregate benchmark score. The 62% compute difference between max and xhigh is a reminder that a more capable tier can change unit economics even before a provider publishes a separate price.
The coming open-weights release could broaden the model’s impact, especially for organizations that need private deployment or tighter control over data and inference infrastructure. But without details on timing, weights, license, hardware requirements, and parity with the hosted model, buyers should treat that option as a future possibility rather than a current deployment path.
For rivals, the pricing is strategically important. Meta is not yet demonstrating uncontested leadership, but it may be lowering the cost of a model that performs well enough for many agentic workloads. That can pressure providers to compete on inference economics, not only on benchmark records.
The first signal will be whether Meta moves Muse Spark 1.3 max from partner preview to general availability after safety testing. Its eventual API price will determine whether the reported $0.55 per-task advantage survives outside the current comparison.
Developers should also watch the open-weights announcement for licensing and model parity. Independent evaluations on real production tasks will matter more than the Intelligence Index alone, particularly for factual reliability, tool-use recovery, and refusal rates.
Finally, the next version will show whether Meta can sustain its rapid release cadence without trading away reliability. The reported regressions on AA-LCR and AA-Omniscience make that balance a more important measure than another single headline benchmark.
Muse Spark 1.3 is best understood as a strong value challenge, not a definitive takeover of the frontier-model market. Meta has improved meaningfully in agentic and coding evaluations, but the top configuration is still restricted, its price is unknown, and several leading models remain ahead on important tests.
For builders, the release could still be consequential. If xhigh delivers dependable tool use at Meta’s stated rates, cost-sensitive teams may have a credible alternative to more expensive models. The decisive evidence will come from accessible max deployment, independent production testing, and the promised open-weights release—not from Meta’s positioning alone.