Clinical Benchmarks

Grok 4.7

Released 21 Sep 2026$2 input, $6 output per million tokens500K contextproprietaryAlso written as grok/grok-4.7Compare with other models

Clinical Benchmarks Index
75.7rank 16 of 148; 75.7 × 1 = 75.7, from 3 of 10 boards
Boards
3 of 13
Results
3
Latest measurement
Sep 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

Clinical reasoning and knowledge

  1. HealthBench Professional
    SpaceXAI evaluation; xHigh effort; HealthBench Professional
    0.567
    Rank 16 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Sep 2026

Documentation and coding

  1. MedScribe (Vals AI)
    model ID grok/grok-4.7; temperature=1; top_p=0.95; reasoning_effort=xhigh
    89.38
    Rank 6 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  2. MedCode (Vals AI)
    model ID grok/grok-4.7; reasoning_effort=xhigh; temperature=1; top_p=0.95
    49.55
    Rank 19 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

Sources

Open a line for the quote and page.

  1. 16HealthBench Professional SpaceXAI evaluation; xHigh effort; HealthBench Professional 0.567
    Printed as 56.7%Vendor-reported, measured Sep 2026Configuration: SpaceXAI evaluation; xHigh effort; HealthBench Professional
    Introducing Grok 4.7 launch post, SpaceXAI, 21 Sep 2026. Model Improvements comparison table, Clinical reasoning / HealthBench Professional row; Grok 4.7 column; September 21, 2026.
    Grok 4.7 xHigh Grok 4.6 High GPT-5.6 Sol Max Fable 5.1 Max Clinical reasoningHealthBench Professional 56.7% 48.5% 60.5% 62.1%
    Every result from this document
  2. 6MedScribe (Vals AI) model ID grok/grok-4.7; temperature=1; top_p=0.95; reasoning_effort=xhigh 89.38
    Printed as 89.38%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4.7; temperature=1; top_p=0.95; reasoning_effort=xhigh
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 6 of 105 (Grok 4.7), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["grok/grok-4.7"].
    6 | Grok 4.7 | 89.38%±1.89 | $2/$6 | 2m33s
    Every result from this document
  3. 19MedCode (Vals AI) model ID grok/grok-4.7; reasoning_effort=xhigh; temperature=1; top_p=0.95 49.55
    Printed as 49.55%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4.7; reasoning_effort=xhigh; temperature=1; top_p=0.95
    Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 19 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["grok/grok-4.7"].
    19 | Grok 4.7 | 49.55%±2.17 | $2/$6 | 3m33s
    Every result from this document

Other SpaceX AI models: Grok 4, Grok 4.1 Fast Non-Reasoning, Grok 4.1 Fast (Reasoning), Grok 4.20, Grok 4.3, Grok 4.5, Grok 4.6, Grok 4 Fast (Non-Reasoning), Grok 4 Fast (Reasoning)