Google models with results here, each placed on its own board.
- Models
- 26
- Results
- 60
- Boards
- 11 of 13
Models
| Model | Kind | Released | Boards |
|---|---|---|---|
| Gemini 3.8 Flash | Model | 2 Sep 2026 | 2 |
| Gemini 3.7 Flash | Model | 13 Aug 2026 | 2 |
| Gemini 3.5 Flash Lite | Model | 21 Jul 2026 | 2 |
| Gemini 3.6 Flash | Model | 21 Jul 2026 | 3 |
| Gemma 4 12B | Model | 3 Jun 2026 | 1 |
| Gemini 3.5 Flash | Model | 19 May 2026 | 3 |
| Gemma 4 26B A4B | Model | 2 Apr 2026 | 1 |
| Gemma 4 31B | Model | 2 Apr 2026 | 1 |
| Gemma 4 E2B | Model | 2 Apr 2026 | 1 |
| Gemma 4 E4B | Model | 2 Apr 2026 | 1 |
| Gemini 3.1 Flash Lite Preview | Model | 3 Mar 2026 | 2 |
| Gemini 3.1 Pro | Model | 19 Feb 2026 | 10 |
| Gemini 3 Flash | Model | 17 Dec 2025 | 3 |
| Gemini 3 Pro | Model | 18 Nov 2025 | 2 |
| Gemini 2.5 Flash Lite (9/25) | Model | 25 Sep 2025 | 1 |
| Gemini 2.5 Flash Preview (9/25) | Model | 25 Sep 2025 | 1 |
| Gemini 2.5 Flash Lite | Model | 22 Jul 2025 | 2 |
| Gemini 2.5 Flash | Model | 17 Jun 2025 | 1 |
| Gemini 2.5 Pro | Model | 17 Jun 2025 | 6 |
| Gemma 3 12B | Model | 12 Mar 2025 | 1 |
| Gemma 3 27B | Model | 12 Mar 2025 | 1 |
| Gemini 2.0 Flash | Model | 5 Feb 2025 | 1 |
| Gemini 2.5 Flash (7/17) | Model | Not stated | 1 |
| Gemini 3 Pro (11/25) | Model | Not stated | 1 |
| MedGemma 27B Text | Model | Not stated | 1 |
| MedGemma 4B | Model | Not stated | 1 |
Results
One line per result. Dark tick: the board leader.
Clinical reasoning and knowledge
- MedXpertQA (MM)Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)81.3Rank 2 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
- MedXpertQA (MM)Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors76.0Rank 7 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
- MedXpertQA (MM)Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families61.3Rank 17 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
- MedXpertQA (MM)Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families58.1Rank 18 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
- 48.7Rank 19 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
- MedXpertQA (MM)Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families28.7Rank 21 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
- MedXpertQA (MM)Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families23.5Rank 22 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
- 59.3Rank 3 of 8 here, 11 models on the boardLeader GPT-5.6 Sol 60.2
- 58.9Rank 4 of 8 here, 11 models on the boardLeader GPT-5.6 Sol 60.2
- 0.652Rank 1 of 10 here, 11 models on the boardLeads this board
- 0.642Rank 2 of 10 here, 11 models on the boardLeader Gemini 3.1 Pro (Preview) 0.652
- 0.529Rank 6 of 10 here, 11 models on the boardLeader Gemini 3.1 Pro (Preview) 0.652
- 0.342Rank 10 of 10 here, 11 models on the boardLeader Gemini 3.1 Pro (Preview) 0.652
Documentation and coding
- MedScribe (Vals AI)model ID google/gemini-3.8-flash; temperature=1; max_output_tokens=65536; reasoning_effort=high84.50Rank 32 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID google/gemini-3.7-flash; temperature=1; max_output_tokens=30000; reasoning_effort=high83.94Rank 37 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID google/gemini-2.5-flash-thinking; temperature=1; max_output_tokens=3000082.98Rank 45 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID google/gemini-2.5-flash; temperature=1; max_output_tokens=3000082.87Rank 47 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID google/gemini-3.6-flash; temperature=1; max_output_tokens=30000; reasoning_effort=high79.66Rank 58 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID google/gemini-2.5-flash-preview-09-2025-thinking; temperature=1; max_output_tokens=3000078.50Rank 61 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID google/gemini-2.5-flash-preview-09-2025; temperature=1; max_output_tokens=3000077.95Rank 64 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID google/gemini-3.5-flash; temperature=1; max_output_tokens=30000; reasoning_effort=high76.57Rank 71 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID google/gemini-3.1-pro-preview; temperature=1; max_output_tokens=30000; reasoning_effort=high76.11Rank 73 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID google/gemini-2.5-flash-lite-preview-09-2025; temperature=1; max_output_tokens=3000075.82Rank 75 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID google/gemini-2.5-pro; temperature=1; max_output_tokens=3000073.55Rank 80 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID google/gemini-2.5-flash-lite; temperature=1; max_output_tokens=3000072.83Rank 82 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID google/gemini-3-pro-preview; temperature=1; max_output_tokens=30000; reasoning_effort=high72.04Rank 87 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID google/gemini-3.5-flash-lite; temperature=1; max_output_tokens=30000; reasoning_effort=high70.89Rank 89 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID google/gemini-3-flash-preview; temperature=1; max_output_tokens=30000; reasoning_effort=high69.92Rank 91 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID google/gemini-2.5-flash-lite-preview-09-2025-thinking; temperature=1; max_output_tokens=3000066.88Rank 96 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID google/gemini-3.1-flash-lite-preview; temperature=1; max_output_tokens=30000; reasoning_effort=high63.90Rank 98 of 105Leader Claude Opus 5.5 91.43
- 59.06Rank 2 of 103Leader Claude Opus 5 63.57
- 55.92Rank 4 of 103Leader Claude Opus 5 63.57
- 55.83Rank 5 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID google/gemini-3.7-flash; reasoning_effort=high; temperature=1; max_output_tokens=3000053.39Rank 8 of 103Leader Claude Opus 5 63.57
- 53.15Rank 10 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID google/gemini-3-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=3000052.20Rank 13 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID google/gemini-2.5-pro; temperature=1; max_output_tokens=3000050.59Rank 15 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID google/gemini-3.8-flash; reasoning_effort=high; temperature=1; max_output_tokens=6553648.13Rank 28 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID google/gemini-3.1-flash-lite-preview; reasoning_effort=high; temperature=1; max_output_tokens=3000047.60Rank 29 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID google/gemini-3.5-flash-lite; reasoning_effort=high; temperature=1; max_output_tokens=3000043.49Rank 41 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID google/gemini-2.5-flash-preview-09-2025; temperature=1; max_output_tokens=3000040.54Rank 60 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID google/gemini-2.5-flash-thinking; temperature=1; max_output_tokens=3000040.36Rank 62 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID google/gemini-2.5-flash-preview-09-2025-thinking; temperature=1; max_output_tokens=3000040.33Rank 63 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID google/gemini-2.5-flash; temperature=1; max_output_tokens=3000038.42Rank 68 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID google/gemini-2.5-flash-lite-preview-09-2025-thinking; temperature=1; max_output_tokens=3000034.19Rank 77 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID google/gemini-2.5-flash-lite; temperature=1; max_output_tokens=3000027.11Rank 98 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID google/gemini-2.5-flash-lite-preview-09-2025; temperature=1; max_output_tokens=3000027.08Rank 99 of 103Leader Claude Opus 5 63.57
EHR and workflow agents
- 6.0Rank 20 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
- 0.630Rank 2 of 18Leader GPT-5.4 (high reasoning) 0.650
- EHR-ComplexTable 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns0.310Rank 15 of 18Leader GPT-5.4 (high reasoning) 0.650
- 12.5Rank 28 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
- 11.9Rank 6 of 7 here, 5 models on the boardLeader Claude Opus 4.6 (computer-use agent) 36.3
Safety
- 80.7Rank 1 of 28Leads this board
- 61.2Rank 11 of 28Leader Gemini-3.1-Pro 80.7
- 54.4Rank 17 of 28Leader Gemini-3.1-Pro 80.7
- 40.3Rank 24 of 28Leader Gemini-3.1-Pro 80.7
- 39.6Rank 26 of 28Leader Gemini-3.1-Pro 80.7
- 39.6Rank 26 of 28Leader Gemini-3.1-Pro 80.7
- 62.6Rank 12 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2
- 61.9Rank 13 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2