Clinical Benchmarks

Gemini 3 Pro

Released 18 Nov 20251M contextproprietaryAlso written as gemini-3-pro-preview, Gemini 3 Pro (11/25), google/gemini-3-pro-preview, gemini-3-pro-11-25Compare with other models

Clinical Benchmarks Index
68.7rank 23 of 148; 84.2 × 0.816 = 68.7, from 2 of 10 boards
Boards
2 of 13
Results
2
Latest measurement
Sep 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

Clinical reasoning and knowledge

  1. MedXpertQA (MM)
    Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors
    76.0
    Rank 7 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
    Independent run
    Measured Feb 2026

Documentation and coding

  1. MedScribe (Vals AI)
    model ID google/gemini-3-pro-preview; temperature=1; max_output_tokens=30000; reasoning_effort=high
    72.04
    Rank 87 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026

Sources

Open a line for the quote and page.

  1. 87MedScribe (Vals AI) model ID google/gemini-3-pro-preview; temperature=1; max_output_tokens=30000; r… 72.04
    Printed as 72.04%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-3-pro-preview; temperature=1; max_output_tokens=30000; reasoning_effort=high
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 87 of 105 (Gemini 3 Pro (11/25)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-3-pro-preview"].
    87 | Gemini 3 Pro (11/25) | 72.04%±1.90 | $2/$12 | 43.39s
    Every result from this document
  2. 7MedXpertQA (MM) Qwen3.5 model card comparison; source-specific evaluation, not harmonized acros… 76.0
    Printed as 76.0Independent run, measured Feb 2026Configuration: Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors
    Qwen/Qwen3.5-397B-A17B model card model card, Alibaba / Qwen, 16 Feb 2026. Benchmark Results > Vision Language > Medical VQA; MedXpertQA-MM row, Gemini-3 Pro column.
    | GPT5.2 | Claude 4.5 Opus | Gemini-3 Pro | Qwen3-VL-235B-A22B | K2.5-1T-A32B | Qwen3.5-397B-A17B MedXpertQA-MM | 73.3 | 63.6 | 76.0 | 47.6 | 65.3 | 70.0
    Every result from this document

Other Google models: Gemini 2.0 Flash, Gemini 2.5 Flash, Gemini 2.5 Flash (7/17), Gemini 2.5 Flash Lite, Gemini 2.5 Flash Lite (9/25), Gemini 2.5 Flash Preview (9/25), Gemini 2.5 Pro, Gemini 3.1 Flash Lite Preview, Gemini 3.1 Pro, Gemini 3.5 Flash, Gemini 3.5 Flash Lite, Gemini 3.6 Flash, Gemini 3.7 Flash, Gemini 3.8 Flash, Gemini 3 Flash, Gemini 3 Pro (11/25), Gemma 3 12B, Gemma 3 27B, Gemma 4 12B, Gemma 4 26B A4B, Gemma 4 31B, Gemma 4 E2B, Gemma 4 E4B, MedGemma 27B Text, MedGemma 4B