Muse Spark
Released 8 Apr 20261.0M contextproprietaryAlso written as Muse Spark (2026-04-08), Muse Spark Thinking, Muse Spark 1.0Compare with other models
- Clinical Benchmarks Index
- 80.9rank 9 of 148; 80.9 × 1 = 80.9, from 5 of 10 boards
- Boards
- 5 of 13
- Results
- 5
- Latest measurement
- Sep 2026
Results
Each line is placed on its own board. Dark tick: the board leader.
Clinical reasoning and knowledge
- HealthBench Professionallength-normalized, GPT-5.4 low-reasoning grader; Muse Spark (1.0) column in the Muse Spark 1.1 Evaluation Report Figure 440.541Rank 19 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- MedXpertQA (MM)Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)78.4Rank 5 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
- 0.621Rank 3 of 10 here, 11 models on the boardLeader Gemini 3.1 Pro (Preview) 0.652
Documentation and coding
- MedScribe (Vals AI)model ID meta/muse_spark; temperature=1; max_output_tokens=3000085.90Rank 22 of 105Leader Claude Opus 5.5 91.43
- MedCode (Vals AI)model ID meta/muse_spark; temperature=1; max_output_tokens=3000051.31Rank 14 of 103Leader Claude Opus 5 63.57
Sources
Open a line for the quote and page.
19HealthBench Professional length-normalized, GPT-5.4 low-reasoning grader; Muse Spark (1.0) column in the… 0.541
Printed as 54.1Vendor-reported, measured Jul 2026Configuration: length-normalized, GPT-5.4 low-reasoning grader; Muse Spark (1.0) column in the Muse Spark 1.1 Evaluation Report Figure 44Muse Spark 1.1 Evaluation Report model card, Meta, 9 Jul 2026. p. 101, Figure 44 (image), row HealthBench Professional, column Muse Spark; protocol p. 104 (printed 103)Health | HealthBench Professional | 59.3 | 54.1 | 41.6 | 55.8 | 51.8 (Figure 44 image table row; columns Muse Spark 1.1, Muse Spark, Gemini 3.1 Pro (high), Opus 4.8 (max), GPT 5.5 (xhigh))
Every result from this document22MedScribe (Vals AI) model ID meta/muse_spark; temperature=1; max_output_tokens=30000 85.90
Printed as 85.90%Official leaderboard, measured Sep 2026Configuration: model ID meta/muse_spark; temperature=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 22 of 105 (Muse Spark), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["meta/muse_spark"].22 | Muse Spark | 85.90%±1.85 | N/A | 3m11s
Every result from this document14MedCode (Vals AI) model ID meta/muse_spark; temperature=1; max_output_tokens=30000 51.31
Printed as 51.31%Official leaderboard, measured Sep 2026Configuration: model ID meta/muse_spark; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 14 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["meta/muse_spark"].14 | Muse Spark | 51.31%±2.24 | N/A | 2m03s
Every result from this document5MedXpertQA (MM) Meta's Muse Spark launch table (Meta reports the better of vendor self-reports… 78.4
Printed as 78.4%Vendor-reported, measured Apr 2026Configuration: Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)Introducing Muse Spark: Scaling Towards Personal Superintelligence launch post, Meta, 8 Apr 2026. Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column Muse Spark Thinking; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'MedXpertQA (MM) | Medical Multiple Choice | 78.4 | 64.8 | 81.3 | 77.1 | 65.8 (image table row; columns Muse Spark Thinking, Opus 4.6 Max, Gemini 3.1 Pro High, GPT 5.4 Xhigh, Grok 4.2 Reasoning)
Every result from this document- Also printed in Muse Spark Eval Methodology: 78.4, p. 5, benchmark table image, HEALTH section, row MedXpertQA (MM), column Muse Spark Thinking
- Mirrored by MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai: 78.4%, BENCHMARK SCORE TABLE (8 MODELS), rank 3; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology
3MedHELM 0.621
Printed as 0.621Official leaderboard, measured May 2026MedHELM leaderboard (medhelm.org), v5.0.0 official leaderboard, Stanford CRFM (MedHELM), 14 May 2026. medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')3 Muse Spark (2026-04-08) Meta 0.621
Every result from this document- Also printed in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate): 0.6208333333333333, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate'
Other Meta models: Llama 3.1 70B, Llama 3.1 8B, Llama 4 Maverick, Llama 4 Scout, Muse Spark 1.1, Muse Spark 1.2