Clinical Benchmarks

SpaceX AI

SpaceX AI models with results here, each placed on its own board.

Models
10
Results
29
Boards
7 of 13

x.ai

Models

ModelKindReleasedBoards
Grok 4.7Model21 Sep 20263
Grok 4.6Model12 Aug 20263
Grok 4.5Model8 Jul 20262
Grok 4.20Model10 Mar 20264
Grok 4.1 Fast (Reasoning)Model19 Nov 20252
Grok 4.1 Fast Non-ReasoningModel19 Nov 20252
Grok 4 Fast (Non-Reasoning)Model19 Sep 20252
Grok 4 Fast (Reasoning)Model19 Sep 20252
Grok 4Model9 Jul 20252
Grok 4.3ModelNot stated4

Results

One line per result. Dark tick: the board leader.

Clinical reasoning and knowledge

  1. HealthBench Professional
    SpaceXAI evaluation; xHigh effort; HealthBench Professional
    0.567
    Rank 16 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Sep 2026
  2. HealthBench Professional
    SpaceXAI evaluation; High effort; HealthBench Professional
    0.485
    Rank 23 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Sep 2026
  3. MedXpertQA (MM)
    Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
    65.8
    Rank 13 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
    Independent run
    Measured Apr 2026
  4. 53.7
    Rank 8 of 8 here, 11 models on the boardLeader GPT-5.6 Sol 60.2
    Official leaderboard
    Measured Aug 2026

Documentation and coding

  1. MedScribe (Vals AI)
    model ID grok/grok-4.7; temperature=1; top_p=0.95; reasoning_effort=xhigh
    89.38
    Rank 6 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  2. MedScribe (Vals AI)
    model ID grok/grok-4.5; temperature=1; top_p=0.95; max_output_tokens=30000; reasoning_effort=high
    86.88
    Rank 16 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  3. 86.53
    Rank 19 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  4. MedScribe (Vals AI)
    model ID grok/grok-4-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000
    81.63
    Rank 50 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  5. MedScribe (Vals AI)
    model ID grok/grok-4-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000
    79.72
    Rank 57 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  6. MedScribe (Vals AI)
    model ID grok/grok-4-1-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000
    78.73
    Rank 60 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  7. MedScribe (Vals AI)
    model ID grok/grok-4-0709; temperature=1; top_p=0.95; max_output_tokens=30000
    78.15
    Rank 62 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  8. MedScribe (Vals AI)
    model ID grok/grok-4-1-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000
    77.46
    Rank 66 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  9. MedScribe (Vals AI)
    model ID grok/grok-4.3; temperature=1; top_p=0.95; max_output_tokens=30000
    74.40
    Rank 78 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  10. MedScribe (Vals AI)
    model ID grok/grok-4.20-0309-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000
    63.41
    Rank 99 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  11. MedCode (Vals AI)
    model ID grok/grok-4.7; reasoning_effort=xhigh; temperature=1; top_p=0.95
    49.55
    Rank 19 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  12. MedCode (Vals AI)
    model ID grok/grok-4.6; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000
    44.71
    Rank 37 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  13. MedCode (Vals AI)
    model ID grok/grok-4.5; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000
    43.29
    Rank 43 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  14. MedCode (Vals AI)
    model ID grok/grok-4-0709; temperature=1; top_p=0.95; max_output_tokens=30000
    38.08
    Rank 69 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  15. MedCode (Vals AI)
    model ID grok/grok-4.3; temperature=1; top_p=0.95; max_output_tokens=30000
    38.07
    Rank 70 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  16. MedCode (Vals AI)
    model ID grok/grok-4-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000
    37.38
    Rank 72 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  17. MedCode (Vals AI)
    model ID grok/grok-4.20-0309-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000
    32.16
    Rank 87 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  18. MedCode (Vals AI)
    model ID grok/grok-4-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000
    30.04
    Rank 93 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  19. MedCode (Vals AI)
    model ID grok/grok-4-1-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000
    28.35
    Rank 96 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  20. MedCode (Vals AI)
    model ID grok/grok-4-1-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000
    28.08
    Rank 97 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

EHR and workflow agents

  1. 5.3
    Rank 21 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
    Official leaderboard
    Measured May 2026
  2. CHI-Bench
    All Domains pass@1; openai-agents harness
    5.8
    Rank 37 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026
  3. CHI-Bench
    All Domains pass@1; hermes harness
    4.4
    Rank 39 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026
  4. CHI-Bench
    All Domains pass@1; deepagents harness
    2.2
    Rank 41 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026
  5. CHI-Bench
    All Domains pass@1; openclaw harness
    0.4
    Rank 42 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026