SPIRITT logoSPIRITTMetaMuse Spark 1.3

Muse Spark 1.3 makes a 20-point DeepSWE jump

Meta's September 2 release targets long-horizon coding, browser, and desktop agents. Its launch scorecard reports 75.4% on DeepSWE v1.1, 98.1% retrieval across 512K to 1M tokens, and 20% fewer tool calls. Muse Spark 1.3 is available in the SPIRITT model picker for agentic apps and workflows.

Muse Spark 1.3 across X

SPIRITT's launch reaction, the official announcement, product direction, benchmark coverage, and an early public coding-agent test.

A release built for longer, more efficient agent loops

Meta's numbers show a broad step up from Muse Spark 1.2. The strongest claims still need independent 1.3 testing.

What Meta launched, and what you can build in SPIRITT

Meta positions Muse Spark 1.3 as a major upgrade for coding, UI understanding, browser use, desktop automation, and multi-hour agentic workflows. The company reports more reliable tool use, better task planning, and 20% fewer tool calls than Muse Spark 1.2 on average.

Meta's launch scorecard reports 75.4% on DeepSWE v1.1, 59.4% on SWEAtlas CodeBase QnA, 66.9% on OSWorld 2.0, and 98.1% on MRCR's 512K-to-1M-token band. Those are vendor-compiled results drawn from a mix of official leaderboards, self-reported comparator scores, and Meta-run evaluations under the linked methodology.

The scorecard uses the model's max configuration, while Meta says max reasoning will arrive after additional safety testing. Launch-day product availability is also rolling out by surface, and the announcement does not publish a separate exact 1.3 price table.

Muse Spark 1.3 is available in the SPIRITT model picker. Select it inside a fully equipped workspace to build agentic apps and workflows with browser, terminal, files, tools, memory, and integrations.

75.4% DeepSWE98.1% long retrieval20% fewer tool callsBrowser + desktop agentsMax reasoning pendingAvailable in SPIRITT

GDPval-AA v2 Elo

Claude Opus 5 max
1824
Muse Spark 1.3 max
1754
GPT-5.6 Sol max
1710
Muse Spark 1.2 xhigh
1615

Meta launch scorecard. Muse Spark 1.3 max at 1,754 improves 139 points over 1.2 and sits between the displayed GPT-5.6 Sol and Claude Opus 5 results.

Where Muse Spark 1.3 advances, ties, and still trails

All 1.3 figures below come from Meta's launch scorecard and methodology. They are not independent SPIRITT reruns.

Vendor-compiled long retrieval
Winner

MRCR 512K–1M

Muse Spark 1.3 max
98.1%
GPT-5.6 Sol max
73.8%
Muse Spark 1.2 xhigh
55.5%

Meta scorecard. Claude Opus 5 is marked unavailable for this context band.

Vendor-compiled coding agent
Winner

DeepSWE v1.1

Muse Spark 1.3 max
75.4%
Claude Opus 5 max
74%
GPT-5.6 Sol max
73%
Muse Spark 1.2 xhigh
55%

Meta scorecard. The 75.4% result is 20.4 points above Muse Spark 1.2.

Vendor-compiled codebase knowledge
Winner

SWEAtlas CodeBase QnA

Muse Spark 1.3 max
59.4%
GPT-5.6 Sol max
53.5%
Claude Opus 5 max
52.7%
Muse Spark 1.2 xhigh
46.2%

Meta scorecard. Higher is better.

Vendor-compiled terminal agents
Winner

Terminal-Bench 2.1

Muse Spark 1.3 max
88.8%
GPT-5.6 Sol max
88.8%
Claude Opus 5 max
86.7%
Muse Spark 1.2 xhigh
82.9%

Meta scorecard. Muse Spark 1.3 ties the displayed GPT-5.6 Sol result.

Vendor-compiled computer use
Near frontier

OSWorld 2.0

Claude Opus 5 max
68.3%
Muse Spark 1.3 max
66.9%
GPT-5.6 Sol max
62.7%
Muse Spark 1.2 xhigh
47.6%

Meta scorecard. Muse Spark 1.3 nearly matches the strongest displayed result but does not lead.

Vendor-compiled job tasks
Near frontier

JobBench

Claude Opus 5 max
65.7%
Muse Spark 1.3 max
64.9%
Muse Spark 1.2 xhigh
61.6%
GPT-5.6 Sol max
45.4%

Meta scorecard. Muse Spark 1.3 is 0.8 points behind the displayed leader.

Vendor-compiled limitation

DeepSearchQA

GPT-5.6 Sol max
93%
Claude Opus 5 max
90.4%
Muse Spark 1.3 max
89.4%
Muse Spark 1.2 xhigh
85.9%

Meta scorecard. Strong absolute result, but behind both displayed frontier comparators.

Vendor-compiled automation
Near frontier

AutomationBench

Claude Opus 5 max
50.3%
Muse Spark 1.3 max
49.4%
GPT-5.6 Sol max
46.7%
Muse Spark 1.2 xhigh
38.2%

Meta scorecard. Muse Spark 1.3 improves sharply over 1.2 but remains just behind the displayed leader.

Source: Meta's September 2 Muse Spark 1.3 launch scorecard and linked multimodal evaluation methodology. Meta selects the highest comparable published score for each system from official leaderboards, self-reported results, or its own runs; every chart is labeled vendor-compiled accordingly. No independent Muse Spark 1.3 model page was publicly listed on Artificial Analysis at verification time. Evaluate the released configuration on your own workloads before production use.

How It Works

Choose Muse Spark 1.3 in the picker, then build the agentic workflow

01

Open a workspace

Open a workspace and land in a fully equipped cloud computer: browser, files, terminal, integrations, and memory. No local setup. No thin chat box pretending to be an agent.

Create a workspace in SPIRITT
02

Choose Muse Spark 1.3

Open the SPIRITT model picker and select Muse Spark 1.3. Then give it the tools and acceptance path for a real coding, browser, or long-horizon agent task.

Electric infinity-loop visual for the Muse Spark 1.3 model-picker step
03

Build or automate

Tell it what to ship or what to run. Muse Spark 1.3 can code, call tools, drive the browser, coordinate multi-step work, and keep going while you step away. The point is not another chat window. It is an agentic environment where Muse Spark 1.3 actually does the job.

Build an agentic app in SPIRITT

Build with Muse Spark 1.3 now.

Open a SPIRITT Workspace, choose Muse Spark 1.3 in the model picker, and turn the release into a working agentic app or workflow.

Questions

Muse Spark 1.3 FAQ

01What is Muse Spark 1.3?+
Meta's September 2026 upgrade to its Muse Spark model for coding, browser use, desktop automation, UI understanding, and long-horizon agentic tasks.
02How much better is it than Muse Spark 1.2?+
Meta's scorecard shows broad gains, including DeepSWE rising from 55.0% to 75.4%, OSWorld 2.0 from 47.6% to 66.9%, and 512K-to-1M MRCR from 55.5% to 98.1%. Meta also reports 20% fewer tool calls on average.
03Are these independent benchmark results?+
No. The displayed 1.3 scorecard is Meta-compiled. Its methodology combines official leaderboard results, self-reported comparator scores, and Meta-run evaluations. No independent Muse Spark 1.3 model page was publicly listed on Artificial Analysis when this page was verified.
04Is max reasoning available at launch?+
Meta says max reasoning will become available after additional safety testing. That matters because the launch scorecard labels Muse Spark 1.3 as max, so launch-access configurations may not reproduce every displayed result immediately.
05What are the main limitations?+
The release still trails the strongest displayed comparator on GDPval-AA v2, JobBench, OSWorld 2.0, DeepSearchQA, Agentic IF, and AutomationBench. Rollout is staged, exact 1.3 pricing was not separately published in the launch post, and independent 1.3 evaluation is still pending.
06Does Muse Spark 1.3 support very long tasks?+
That is a central release claim. Meta targets multi-hour workflows and reports 98.1% on its 512K-to-1M-token MRCR evaluation band. Real applications should still test retrieval and tool-state reliability across their own longest traces.
07Is Muse Spark 1.3 available in SPIRITT Workspaces?+
Yes. Open a SPIRITT Workspace, choose Muse Spark 1.3 in the model picker, and use it for coding, browser, desktop, and long-horizon agentic workflows.
Buy from builders who use what they sellBuilt usingSPIRITT