Clinical Benchmarks

Llama 4 Maverick

Released 5 Apr 20251M context400B A17BopenAlso written as fireworks/llama4-maverick-instruct-basic, llama4-maverick-instruct-basicCompare with other models

Clinical Benchmarks Index
39.0rank 101 of 148; 47.8 × 0.816 = 39.0, from 2 of 10 boards
Boards
2 of 13
Results
2
Latest measurement
Sep 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

Documentation and coding

  1. MedScribe (Vals AI)
    model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_tokens=30000
    54.22
    Rank 103 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  2. MedCode (Vals AI)
    model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_tokens=30000
    36.51
    Rank 74 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

Sources

Open a line for the quote and page.

  1. 103MedScribe (Vals AI) model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_to… 54.22
    Printed as 54.22%Official leaderboard, measured Sep 2026Configuration: model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 103 of 105 (Llama 4 Maverick), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["fireworks/llama4-maverick-instruct-basic"].
    103 | Llama 4 Maverick | 54.22%±1.87 | $0.22/$0.88 | 25.05s
    Every result from this document
  2. 74MedCode (Vals AI) model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_to… 36.51
    Printed as 36.51%Official leaderboard, measured Sep 2026Configuration: model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_tokens=30000
    Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 74 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["fireworks/llama4-maverick-instruct-basic"].
    74 | Llama 4 Maverick | 36.51%±1.99 | $0.22/$0.88 | 21.24s
    Every result from this document

Other Meta models: Llama 3.1 70B, Llama 3.1 8B, Llama 4 Scout, Muse Spark, Muse Spark 1.1, Muse Spark 1.2