Meta
Meta models with results here, each placed on its own board.
- Models
- 7
- Results
- 16
- Boards
- 7 of 13
Models
| Model | Kind | Released | Boards |
|---|---|---|---|
| Muse Spark 1.2 | Model | 5 Aug 2026 | 2 |
| Muse Spark 1.1 | Model | 9 Jul 2026 | 3 |
| Muse Spark | Model | 8 Apr 2026 | 5 |
| Llama 4 Maverick | Model | 5 Apr 2025 | 2 |
| Llama 4 Scout | Model | 5 Apr 2025 | 2 |
| Llama 3.1 70B | Model | 23 Jul 2024 | 1 |
| Llama 3.1 8B | Model | 23 Jul 2024 | 1 |
Results
One line per result. Dark tick: the board leader.
Clinical reasoning and knowledge
- HealthBench Professionallength-normalized, GPT-5.4 low-reasoning grader, xhigh reasoning via Meta Model API (Muse Spark 1.1 Evaluation Report Figure 44)0.593Rank 12 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- HealthBench Professionallength-normalized, GPT-5.4 low-reasoning grader; Muse Spark (1.0) column in the Muse Spark 1.1 Evaluation Report Figure 440.541Rank 19 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- MedXpertQA (MM)Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)78.4Rank 5 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
- 0.621Rank 3 of 10 here, 11 models on the boardLeader Gemini 3.1 Pro (Preview) 0.652
Documentation and coding
- 90.06Rank 5 of 105Leader Claude Opus 5.5 91.43
- 88.89Rank 8 of 105Leader Claude Opus 5.5 91.43
- 85.90Rank 22 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_tokens=3000054.22Rank 103 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max_output_tokens=3000050.59Rank 104 of 105Leader Claude Opus 5.5 91.43
- 51.31Rank 14 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID meta/muse_spark_1_2; reasoning_effort=xhigh; temperature=1; max_output_tokens=3000049.35Rank 20 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_tokens=3000036.51Rank 74 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max_output_tokens=3000023.31Rank 100 of 103Leader Claude Opus 5 63.57
Safety
- 55.7Rank 15 of 28Leader Gemini-3.1-Pro 80.7
- 40.0Rank 25 of 28Leader Gemini-3.1-Pro 80.7
- 79.7Rank 4 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2