Clinical Benchmarks

Muse Spark

Released 8 Apr 20261.0M contextproprietaryAlso written as Muse Spark (2026-04-08), Muse Spark Thinking, Muse Spark 1.0Compare with other models

Clinical Benchmarks Index
80.9rank 9 of 148; 80.9 × 1 = 80.9, from 5 of 10 boards
Boards
5 of 13
Results
5
Latest measurement
Sep 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

Clinical reasoning and knowledge

  1. HealthBench Professional
    length-normalized, GPT-5.4 low-reasoning grader; Muse Spark (1.0) column in the Muse Spark 1.1 Evaluation Report Figure 44
    0.541
    Rank 19 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Jul 2026
  2. MedXpertQA (MM)
    Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
    78.4
    Rank 5 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
    Vendor-reported
    Measured Apr 2026
  3. 0.621
    Rank 3 of 10 here, 11 models on the boardLeader Gemini 3.1 Pro (Preview) 0.652
    Official leaderboard
    Measured May 2026

Documentation and coding

  1. MedScribe (Vals AI)
    model ID meta/muse_spark; temperature=1; max_output_tokens=30000
    85.90
    Rank 22 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  2. MedCode (Vals AI)
    model ID meta/muse_spark; temperature=1; max_output_tokens=30000
    51.31
    Rank 14 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

Sources

Open a line for the quote and page.

  1. 19HealthBench Professional length-normalized, GPT-5.4 low-reasoning grader; Muse Spark (1.0) column in the… 0.541
    Printed as 54.1Vendor-reported, measured Jul 2026Configuration: length-normalized, GPT-5.4 low-reasoning grader; Muse Spark (1.0) column in the Muse Spark 1.1 Evaluation Report Figure 44
    Muse Spark 1.1 Evaluation Report model card, Meta, 9 Jul 2026. p. 101, Figure 44 (image), row HealthBench Professional, column Muse Spark; protocol p. 104 (printed 103)
    Health | HealthBench Professional | 59.3 | 54.1 | 41.6 | 55.8 | 51.8 (Figure 44 image table row; columns Muse Spark 1.1, Muse Spark, Gemini 3.1 Pro (high), Opus 4.8 (max), GPT 5.5 (xhigh))
    Every result from this document
  2. 22MedScribe (Vals AI) model ID meta/muse_spark; temperature=1; max_output_tokens=30000 85.90
    Printed as 85.90%Official leaderboard, measured Sep 2026Configuration: model ID meta/muse_spark; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 22 of 105 (Muse Spark), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["meta/muse_spark"].
    22 | Muse Spark | 85.90%±1.85 | N/A | 3m11s
    Every result from this document
  3. 14MedCode (Vals AI) model ID meta/muse_spark; temperature=1; max_output_tokens=30000 51.31
    Printed as 51.31%Official leaderboard, measured Sep 2026Configuration: model ID meta/muse_spark; temperature=1; max_output_tokens=30000
    Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 14 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["meta/muse_spark"].
    14 | Muse Spark | 51.31%±2.24 | N/A | 2m03s
    Every result from this document
  4. 5MedXpertQA (MM) Meta's Muse Spark launch table (Meta reports the better of vendor self-reports… 78.4
    Printed as 78.4%Vendor-reported, measured Apr 2026Configuration: Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
    Introducing Muse Spark: Scaling Towards Personal Superintelligence launch post, Meta, 8 Apr 2026. Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column Muse Spark Thinking; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'
    MedXpertQA (MM) | Medical Multiple Choice | 78.4 | 64.8 | 81.3 | 77.1 | 65.8 (image table row; columns Muse Spark Thinking, Opus 4.6 Max, Gemini 3.1 Pro High, GPT 5.4 Xhigh, Grok 4.2 Reasoning)
    Every result from this document
  5. 3MedHELM 0.621
    Printed as 0.621Official leaderboard, measured May 2026
    MedHELM leaderboard (medhelm.org), v5.0.0 official leaderboard, Stanford CRFM (MedHELM), 14 May 2026. medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
    3 Muse Spark (2026-04-08) Meta 0.621
    Every result from this document

Other Meta models: Llama 3.1 70B, Llama 3.1 8B, Llama 4 Maverick, Llama 4 Scout, Muse Spark 1.1, Muse Spark 1.2