Clinical Benchmarks

Google

Google models with results here, each placed on its own board.

Models
26
Results
60
Boards
11 of 13

deepmind.google

Models

ModelKindReleasedBoards
Gemini 3.8 FlashModel2 Sep 20262
Gemini 3.7 FlashModel13 Aug 20262
Gemini 3.5 Flash LiteModel21 Jul 20262
Gemini 3.6 FlashModel21 Jul 20263
Gemma 4 12BModel3 Jun 20261
Gemini 3.5 FlashModel19 May 20263
Gemma 4 26B A4BModel2 Apr 20261
Gemma 4 31BModel2 Apr 20261
Gemma 4 E2BModel2 Apr 20261
Gemma 4 E4BModel2 Apr 20261
Gemini 3.1 Flash Lite PreviewModel3 Mar 20262
Gemini 3.1 ProModel19 Feb 202610
Gemini 3 FlashModel17 Dec 20253
Gemini 3 ProModel18 Nov 20252
Gemini 2.5 Flash Lite (9/25)Model25 Sep 20251
Gemini 2.5 Flash Preview (9/25)Model25 Sep 20251
Gemini 2.5 Flash LiteModel22 Jul 20252
Gemini 2.5 FlashModel17 Jun 20251
Gemini 2.5 ProModel17 Jun 20256
Gemma 3 12BModel12 Mar 20251
Gemma 3 27BModel12 Mar 20251
Gemini 2.0 FlashModel5 Feb 20251
Gemini 2.5 Flash (7/17)ModelNot stated1
Gemini 3 Pro (11/25)ModelNot stated1
MedGemma 27B TextModelNot stated1
MedGemma 4BModelNot stated1

Results

One line per result. Dark tick: the board leader.

Clinical reasoning and knowledge

  1. MedXpertQA (MM)
    Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
    81.3
    Rank 2 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
    Independent run
    Measured Apr 2026
  2. MedXpertQA (MM)
    Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors
    76.0
    Rank 7 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
    Independent run
    Measured Feb 2026
  3. MedXpertQA (MM)
    Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families
    61.3
    Rank 17 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
    Vendor-reported
    Measured Apr 2026
  4. MedXpertQA (MM)
    Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families
    58.1
    Rank 18 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
    Vendor-reported
    Measured Apr 2026
  5. MedXpertQA (MM)
    Google's Gemma 4 model card, Unified 12B; protocol not stated
    48.7
    Rank 19 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
    Vendor-reported
    Measured Apr 2026
  6. MedXpertQA (MM)
    Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families
    28.7
    Rank 21 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
    Vendor-reported
    Measured Apr 2026
  7. MedXpertQA (MM)
    Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families
    23.5
    Rank 22 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
    Vendor-reported
    Measured Apr 2026
  8. 59.3
    Rank 3 of 8 here, 11 models on the boardLeader GPT-5.6 Sol 60.2
    Official leaderboard
    Measured Aug 2026
  9. 58.9
    Rank 4 of 8 here, 11 models on the boardLeader GPT-5.6 Sol 60.2
    Official leaderboard
    Measured Aug 2026
  10. 0.652
    Rank 1 of 10 here, 11 models on the boardLeads this board
    Official leaderboard
    Measured May 2026
  11. 0.642
    Rank 2 of 10 here, 11 models on the boardLeader Gemini 3.1 Pro (Preview) 0.652
    Official leaderboard
    Measured May 2026
  12. 0.529
    Rank 6 of 10 here, 11 models on the boardLeader Gemini 3.1 Pro (Preview) 0.652
    Official leaderboard
    Measured May 2026
  13. 0.342
    Rank 10 of 10 here, 11 models on the boardLeader Gemini 3.1 Pro (Preview) 0.652
    Official leaderboard
    Measured May 2026

Documentation and coding

  1. MedScribe (Vals AI)
    model ID google/gemini-3.8-flash; temperature=1; max_output_tokens=65536; reasoning_effort=high
    84.50
    Rank 32 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  2. MedScribe (Vals AI)
    model ID google/gemini-3.7-flash; temperature=1; max_output_tokens=30000; reasoning_effort=high
    83.94
    Rank 37 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  3. MedScribe (Vals AI)
    model ID google/gemini-2.5-flash-thinking; temperature=1; max_output_tokens=30000
    82.98
    Rank 45 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  4. MedScribe (Vals AI)
    model ID google/gemini-2.5-flash; temperature=1; max_output_tokens=30000
    82.87
    Rank 47 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  5. MedScribe (Vals AI)
    model ID google/gemini-3.6-flash; temperature=1; max_output_tokens=30000; reasoning_effort=high
    79.66
    Rank 58 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  6. MedScribe (Vals AI)
    model ID google/gemini-2.5-flash-preview-09-2025-thinking; temperature=1; max_output_tokens=30000
    78.50
    Rank 61 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  7. MedScribe (Vals AI)
    model ID google/gemini-2.5-flash-preview-09-2025; temperature=1; max_output_tokens=30000
    77.95
    Rank 64 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  8. MedScribe (Vals AI)
    model ID google/gemini-3.5-flash; temperature=1; max_output_tokens=30000; reasoning_effort=high
    76.57
    Rank 71 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  9. MedScribe (Vals AI)
    model ID google/gemini-3.1-pro-preview; temperature=1; max_output_tokens=30000; reasoning_effort=high
    76.11
    Rank 73 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  10. MedScribe (Vals AI)
    model ID google/gemini-2.5-flash-lite-preview-09-2025; temperature=1; max_output_tokens=30000
    75.82
    Rank 75 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  11. MedScribe (Vals AI)
    model ID google/gemini-2.5-pro; temperature=1; max_output_tokens=30000
    73.55
    Rank 80 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  12. MedScribe (Vals AI)
    model ID google/gemini-2.5-flash-lite; temperature=1; max_output_tokens=30000
    72.83
    Rank 82 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  13. MedScribe (Vals AI)
    model ID google/gemini-3-pro-preview; temperature=1; max_output_tokens=30000; reasoning_effort=high
    72.04
    Rank 87 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  14. MedScribe (Vals AI)
    model ID google/gemini-3.5-flash-lite; temperature=1; max_output_tokens=30000; reasoning_effort=high
    70.89
    Rank 89 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  15. MedScribe (Vals AI)
    model ID google/gemini-3-flash-preview; temperature=1; max_output_tokens=30000; reasoning_effort=high
    69.92
    Rank 91 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  16. MedScribe (Vals AI)
    model ID google/gemini-2.5-flash-lite-preview-09-2025-thinking; temperature=1; max_output_tokens=30000
    66.88
    Rank 96 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  17. MedScribe (Vals AI)
    model ID google/gemini-3.1-flash-lite-preview; temperature=1; max_output_tokens=30000; reasoning_effort=high
    63.90
    Rank 98 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  18. 59.06
    Rank 2 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  19. 55.92
    Rank 4 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  20. 55.83
    Rank 5 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  21. MedCode (Vals AI)
    model ID google/gemini-3.7-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000
    53.39
    Rank 8 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  22. 53.15
    Rank 10 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  23. MedCode (Vals AI)
    model ID google/gemini-3-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000
    52.20
    Rank 13 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  24. MedCode (Vals AI)
    model ID google/gemini-2.5-pro; temperature=1; max_output_tokens=30000
    50.59
    Rank 15 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  25. MedCode (Vals AI)
    model ID google/gemini-3.8-flash; reasoning_effort=high; temperature=1; max_output_tokens=65536
    48.13
    Rank 28 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  26. MedCode (Vals AI)
    model ID google/gemini-3.1-flash-lite-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000
    47.60
    Rank 29 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  27. MedCode (Vals AI)
    model ID google/gemini-3.5-flash-lite; reasoning_effort=high; temperature=1; max_output_tokens=30000
    43.49
    Rank 41 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  28. MedCode (Vals AI)
    model ID google/gemini-2.5-flash-preview-09-2025; temperature=1; max_output_tokens=30000
    40.54
    Rank 60 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  29. MedCode (Vals AI)
    model ID google/gemini-2.5-flash-thinking; temperature=1; max_output_tokens=30000
    40.36
    Rank 62 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  30. MedCode (Vals AI)
    model ID google/gemini-2.5-flash-preview-09-2025-thinking; temperature=1; max_output_tokens=30000
    40.33
    Rank 63 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  31. MedCode (Vals AI)
    model ID google/gemini-2.5-flash; temperature=1; max_output_tokens=30000
    38.42
    Rank 68 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  32. MedCode (Vals AI)
    model ID google/gemini-2.5-flash-lite-preview-09-2025-thinking; temperature=1; max_output_tokens=30000
    34.19
    Rank 77 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  33. MedCode (Vals AI)
    model ID google/gemini-2.5-flash-lite; temperature=1; max_output_tokens=30000
    27.11
    Rank 98 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  34. MedCode (Vals AI)
    model ID google/gemini-2.5-flash-lite-preview-09-2025; temperature=1; max_output_tokens=30000
    27.08
    Rank 99 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

EHR and workflow agents

  1. 6.0
    Rank 20 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
    Official leaderboard
    Measured May 2026
  2. EHR-Complex
    validation configuration
    0.630
    Rank 2 of 18Leader GPT-5.4 (high reasoning) 0.650
    Official leaderboard
    Measured Jun 2026
  3. EHR-Complex
    Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    0.310
    Rank 15 of 18Leader GPT-5.4 (high reasoning) 0.650
    Official leaderboard
    Measured Jun 2026
  4. CHI-Bench
    All Domains pass@1; gemini-cli harness
    12.5
    Rank 28 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured Apr 2026
  5. HealthAdminBench
    screenshot-only, detailed prompting
    11.9
    Rank 6 of 7 here, 5 models on the boardLeader Claude Opus 4.6 (computer-use agent) 36.3
    Official leaderboard
    Measured Apr 2026

Safety

  1. MedPIC
    zero-shot; independent questions; exact option-set match
    80.7
    Rank 1 of 28Leads this board
    Independent run
    Measured Aug 2026
  2. MedPIC
    zero-shot; independent questions; exact option-set match
    61.2
    Rank 11 of 28Leader Gemini-3.1-Pro 80.7
    Independent run
    Measured Aug 2026
  3. MedPIC
    zero-shot; independent questions; exact option-set match
    54.4
    Rank 17 of 28Leader Gemini-3.1-Pro 80.7
    Independent run
    Measured Aug 2026
  4. MedPIC
    zero-shot; independent questions; exact option-set match
    40.3
    Rank 24 of 28Leader Gemini-3.1-Pro 80.7
    Independent run
    Measured Aug 2026
  5. MedPIC
    zero-shot; independent questions; exact option-set match
    39.6
    Rank 26 of 28Leader Gemini-3.1-Pro 80.7
    Independent run
    Measured Aug 2026
  6. MedPIC
    zero-shot; independent questions; exact option-set match
    39.6
    Rank 26 of 28Leader Gemini-3.1-Pro 80.7
    Independent run
    Measured Aug 2026
  7. 62.6
    Rank 12 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2
    Official leaderboard
    Measured Aug 2026
  8. 61.9
    Rank 13 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2
    Official leaderboard
    Measured Aug 2026