Clinical Benchmarks

MedGemma 27B Text

128K context27BopenAlso written as MedGemma-27B-TextCompare with other models

Clinical Benchmarks Index
32.5rank 109 of 148; 56.4 × 0.577 = 32.5, from 1 of 10 boards
Boards
1 of 13
Results
76 on secondary measures
Latest measurement
Aug 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

Safety

  1. MedPIC
    zero-shot; independent questions; exact option-set match
    61.2
    Rank 11 of 28Leader Gemini-3.1-Pro 80.7
    Independent run
    Measured Aug 2026

Secondary measures

Safety

  1. MedPIC, Risk activation accuracy
    zero-shot; independent questions; exact option-set match
    76.1
    Rank 1 of 28Leads this board
    Independent run
    Measured Aug 2026
  2. MedPIC, Counterfactual accuracy (CF)
    zero-shot; independent questions; exact option-set match
    51.4
    Rank 10 of 28Leader Gemini-3.1-Pro 69.9
    Independent run
    Measured Aug 2026
  3. MedPIC, Guideline-following accuracy (GF)
    zero-shot; independent questions; exact option-set match
    67.6
    Rank 12 of 28Leader Gemini-3.1-Pro 87.7
    Independent run
    Measured Aug 2026
  4. MedPIC, Linked counterfactual pair accuracy
    zero-shot; independent questions; exact option-set match
    23.6
    Rank 12 of 28Leader Gemini-3.1-Pro 48.3
    Independent run
    Measured Aug 2026
  5. MedPIC, Risk deactivation accuracy
    zero-shot; independent questions; exact option-set match
    32.4
    Rank 16 of 28Leader Gemini-3.1-Pro 77.9
    Independent run
    Measured Aug 2026
  6. MedPIC, Published GF-minus-CF accuracy gap
    zero-shot; independent questions; exact option-set match
    16.2
    Descriptive measure, not ranked
    Independent run
    Measured Aug 2026

Sources

Open a line for the quote and page.

  1. 11MedPIC zero-shot; independent questions; exact option-set match 61.2
    Printed as 61.2Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row MedGemma-27B-Text, column Overall; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    MedGemma-27B-Text | 61.2 | 67.6 | 51.4 | 16.2 | 76.1 | 32.4 | 23.6
    Every result from this document
  2. 12MedPIC, Guideline-following accuracy (GF) zero-shot; independent questions; exact option-set match 67.6
    Printed as 67.6Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row MedGemma-27B-Text, column GF; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    MedGemma-27B-Text | 61.2 | 67.6 | 51.4 | 16.2 | 76.1 | 32.4 | 23.6
    Every result from this document
  3. 10MedPIC, Counterfactual accuracy (CF) zero-shot; independent questions; exact option-set match 51.4
    Printed as 51.4Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row MedGemma-27B-Text, column CF; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    MedGemma-27B-Text | 61.2 | 67.6 | 51.4 | 16.2 | 76.1 | 32.4 | 23.6
    Every result from this document
  4. 12MedPIC, Linked counterfactual pair accuracy zero-shot; independent questions; exact option-set match 23.6
    Printed as 23.6Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row MedGemma-27B-Text, column Pair; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    MedGemma-27B-Text | 61.2 | 67.6 | 51.4 | 16.2 | 76.1 | 32.4 | 23.6
    Every result from this document
  5. 1MedPIC, Risk activation accuracy zero-shot; independent questions; exact option-set match 76.1
    Printed as 76.1Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row MedGemma-27B-Text, column Activation; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    MedGemma-27B-Text | 61.2 | 67.6 | 51.4 | 16.2 | 76.1 | 32.4 | 23.6
    Every result from this document
  6. 16MedPIC, Risk deactivation accuracy zero-shot; independent questions; exact option-set match 32.4
    Printed as 32.4Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row MedGemma-27B-Text, column Deactivation; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    MedGemma-27B-Text | 61.2 | 67.6 | 51.4 | 16.2 | 76.1 | 32.4 | 23.6
    Every result from this document
  7. –MedPIC, Published GF-minus-CF accuracy gap zero-shot; independent questions; exact option-set match 16.2
    Printed as 16.2Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row MedGemma-27B-Text, column Δ_GF−CF; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    MedGemma-27B-Text | 61.2 | 67.6 | 51.4 | 16.2 | 76.1 | 32.4 | 23.6
    Every result from this document

Other Google models: Gemini 2.0 Flash, Gemini 2.5 Flash, Gemini 2.5 Flash (7/17), Gemini 2.5 Flash Lite, Gemini 2.5 Flash Lite (9/25), Gemini 2.5 Flash Preview (9/25), Gemini 2.5 Pro, Gemini 3.1 Flash Lite Preview, Gemini 3.1 Pro, Gemini 3.5 Flash, Gemini 3.5 Flash Lite, Gemini 3.6 Flash, Gemini 3.7 Flash, Gemini 3.8 Flash, Gemini 3 Flash, Gemini 3 Pro, Gemini 3 Pro (11/25), Gemma 3 12B, Gemma 3 27B, Gemma 4 12B, Gemma 4 26B A4B, Gemma 4 31B, Gemma 4 E2B, Gemma 4 E4B, MedGemma 4B