Clinical Benchmarks

Meta

Meta models with results here, each placed on its own board.

Models
7
Results
16
Boards
7 of 13

ai.meta.com

Models

ModelKindReleasedBoards
Muse Spark 1.2Model5 Aug 20262
Muse Spark 1.1Model9 Jul 20263
Muse SparkModel8 Apr 20265
Llama 4 MaverickModel5 Apr 20252
Llama 4 ScoutModel5 Apr 20252
Llama 3.1 70BModel23 Jul 20241
Llama 3.1 8BModel23 Jul 20241

Results

One line per result. Dark tick: the board leader.

Clinical reasoning and knowledge

  1. HealthBench Professional
    length-normalized, GPT-5.4 low-reasoning grader, xhigh reasoning via Meta Model API (Muse Spark 1.1 Evaluation Report Figure 44)
    0.593
    Rank 12 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Jul 2026
  2. HealthBench Professional
    length-normalized, GPT-5.4 low-reasoning grader; Muse Spark (1.0) column in the Muse Spark 1.1 Evaluation Report Figure 44
    0.541
    Rank 19 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Jul 2026
  3. MedXpertQA (MM)
    Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
    78.4
    Rank 5 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
    Vendor-reported
    Measured Apr 2026
  4. 0.621
    Rank 3 of 10 here, 11 models on the boardLeader Gemini 3.1 Pro (Preview) 0.652
    Official leaderboard
    Measured May 2026

Documentation and coding

  1. 90.06
    Rank 5 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  2. 88.89
    Rank 8 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  3. MedScribe (Vals AI)
    model ID meta/muse_spark; temperature=1; max_output_tokens=30000
    85.90
    Rank 22 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  4. MedScribe (Vals AI)
    model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_tokens=30000
    54.22
    Rank 103 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  5. MedScribe (Vals AI)
    model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max_output_tokens=30000
    50.59
    Rank 104 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  6. MedCode (Vals AI)
    model ID meta/muse_spark; temperature=1; max_output_tokens=30000
    51.31
    Rank 14 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  7. MedCode (Vals AI)
    model ID meta/muse_spark_1_2; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000
    49.35
    Rank 20 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  8. MedCode (Vals AI)
    model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_tokens=30000
    36.51
    Rank 74 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  9. MedCode (Vals AI)
    model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max_output_tokens=30000
    23.31
    Rank 100 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

Safety

  1. MedPIC
    zero-shot; independent questions; exact option-set match
    55.7
    Rank 15 of 28Leader Gemini-3.1-Pro 80.7
    Independent run
    Measured Aug 2026
  2. MedPIC
    zero-shot; independent questions; exact option-set match
    40.0
    Rank 25 of 28Leader Gemini-3.1-Pro 80.7
    Independent run
    Measured Aug 2026
  3. 79.7
    Rank 4 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2
    Official leaderboard
    Measured Aug 2026