BENCHBAITSTORY · MONITORING

AI NEWS, TRACKED AND ROASTED.

meta reports muse spark 1.3 makes about 20% fewer tool calls than 1.2. the first doubt landed on the benchmark

Meta AI Research announced [95] Muse Spark 1.3 ↗ on September 2, 2026, and says it is rolling out today in Muse Code and Meta Model API. Previously available reasoning modes are available; max reasoning is coming after additional safety testing. The efficiency claim is Meta's own, from comparisons by Meta engineers against Muse Spark 1.2.

Tool calls
Figure
approximately 20% fewer
Tokens
Figure
approximately 25% fewer

The linked [96] Muse Spark 1.3 multimodal evaluation methodology ↗ says the Muse Spark 1.3 results were generated through Meta Model API, and it records what the comparison held constant and what it did not:

  • Muse Spark 1.3, Claude Opus 5, and GPT-5.6 Sol ran at max reasoning effort; Muse Spark 1.2 ran at xhigh.
  • OSWorld version 08.08 for Muse Spark 1.3, version 06.24 for Muse Spark 1.2.
  • Third-party model runs are best-effort and may not reflect provider-optimized performance.
  • Ungradable responses and refusals receive zero credit and remain in the denominator.

The announcement does not state a pricing schedule, a context limit, or an API model ID.

what people are saying

Both items below are attributed community material — the first a critique, the second a reaction — summarized as the packet records them rather than reproduced verbatim. Neither confirms model performance.

Early results look interesting — but is something off with the DeepSWE benchmark?

— Sergio (@sergooforai1), in [97] an X post questioning the DeepSWE benchmark ↗ published September 3, 2026. The post carries one image; its original media URL returned 404 during bounded inspection, so the image content was not used as evidence. Attributed reaction only, not independently verified.

Benchmark results are not meaningful until they are proven in practice.

— alexmil78, in [98] an r/AI_Agents thread asking for practical comparisons with GPT 5.6 Sol and Claude Opus 5 ↗ published September 3, 2026.

I have not tested Muse Spark. New-model benchmark hype and hallucinated syntax on a simple script can coexist.

— Loose-Finish-2133, commenting in [98] that same r/AI_Agents comparison thread ↗. An untested concern, not a test result.

claims

  • Launch (confirmed). Meta AI Research announced Meta Muse Spark 1.3 on September 2, 2026.
  • Availability (confirmed). The announcement says Muse Spark 1.3 is rolling out in Muse Code and Meta Model API.
  • Efficiency (supported). Meta reports approximately 20% fewer tool calls and 25% fewer tokens relative to Muse Spark 1.2, in comparisons by Meta engineers.
  • Evaluation caveats (supported). Meta's evaluation methodology uses differing configurations and says third-party comparisons are best-effort, so the published comparisons do not establish universal practical superiority.
  • Discourse (supported). The bounded community material adds a benchmark-to-practice skepticism angle beyond the launch facts; it is attributed reaction and contradiction, not confirmation.

receipts

LAST VERIFIED 2026-09-03

WHAT CHANGED

THE TIMELINE

REVISION 1 · LAUNCH · 2026-09-03ROAST / SATIRE

meta reports muse spark 1.3 makes about 20% fewer tool calls than 1.2. the first doubt landed on the benchmark

  • VERIFIEDMeta AI Research announced Meta Muse Spark 1.3 on September 2, 2026.RECEIPTS [95]
  • VERIFIEDThe announcement says Muse Spark 1.3 is rolling out in Muse Code and Meta Model API.RECEIPTS [95]
  • QUALIFIEDMeta reports approximately 20% fewer tool calls and 25% fewer tokens relative to Muse Spark 1.2 in comparisons by Meta engineers.RECEIPTS [95]
  • QUALIFIEDMeta's evaluation methodology uses differing configurations and says third-party comparisons are best-effort, so the published comparisons do not establish universal practical superiority.RECEIPTS [96][97][98]
  • QUALIFIEDThe bounded discourse adds a distinct benchmark-to-practice skepticism angle beyond the launch facts.RECEIPTS [97][98]

EVERY JOKE HAS RECEIPTS

SOURCES

  1. [95] PRIMARY · Meta AI ResearchIntroducing Muse Spark 1.3primary ↗
  2. [96] PRIMARY · Meta AI ResearchMuse Spark 1.3 Multimodal Evaluation Methodologyprimary ↗
  3. X reaction questioning the DeepSWE benchmark

    The post says early results look interesting and asks whether something is off with the DeepSWE benchmark. The post carries one image; the original media URL returned 404 during bounded inspection, so the image content was not used as evidence. Verification status: attributed reaction; not independently verified and not factual confirmation.

    contradiction ↗
  4. Muse Spark 1.3 vs. Claude Opus 5 vs GPT 5.6 Sol

    The post asks for practical comparisons with GPT 5.6 Sol and Claude Opus 5. The author says benchmark results are not meaningful until proven in practice. A commenter, Loose-Finish-2133, says they have not tested Muse Spark and predicts new-model benchmark hype can coexist with hallucinated syntax on a simple script. Verification status: attributed community reaction and untested concern; not independently verified and not factual confirmation.

    reaction ↗