Clinical Benchmarks

Llama 4 Scout

Released 5 Apr 202510M context109B A17BopenAlso written as meta-llama/Llama-4-Scout-17B-16E-Instruct, together/meta-llama/Llama-4-Scout-17B-16E-InstructCompare with other models

Clinical Benchmarks Index
25.1rank 119 of 148; 30.7 × 0.816 = 25.1, from 2 of 10 boards
Boards
2 of 13
Results
2
Latest measurement
Sep 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

Documentation and coding

  1. MedScribe (Vals AI)
    model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max_output_tokens=30000
    50.59
    Rank 104 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  2. MedCode (Vals AI)
    model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max_output_tokens=30000
    23.31
    Rank 100 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

Sources

Open a line for the quote and page.

  1. 104MedScribe (Vals AI) model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max… 50.59
    Printed as 50.59%Official leaderboard, measured Sep 2026Configuration: model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 104 of 105 (Llama 4 Scout), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["together/meta-llama/Llama-4-Scout-17B-16E-Instruct"].
    104 | Llama 4 Scout | 50.59%±1.90 | $0.18/$0.59 | 11.32s
    Every result from this document
  2. 100MedCode (Vals AI) model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max… 23.31
    Printed as 23.31%Official leaderboard, measured Sep 2026Configuration: model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max_output_tokens=30000
    Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 100 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["together/meta-llama/Llama-4-Scout-17B-16E-Instruct"].
    100 | Llama 4 Scout | 23.31%±1.75 | $0.18/$0.59 | 10.74s
    Every result from this document

Other Meta models: Llama 3.1 70B, Llama 3.1 8B, Llama 4 Maverick, Muse Spark, Muse Spark 1.1, Muse Spark 1.2