Clinical Benchmarks

Qwen3.5-9B

256K context9BopenCompare with other models

Clinical Benchmarks Index
35.9rank 105 of 148; 62.2 × 0.577 = 35.9, from 1 of 10 boards
Boards
1 of 13
Results
76 on secondary measures
Latest measurement
Aug 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

Safety

  1. MedPIC
    zero-shot; independent questions; exact option-set match
    63.8
    Rank 10 of 28Leader Gemini-3.1-Pro 80.7
    Independent run
    Measured Aug 2026

Secondary measures

Safety

  1. MedPIC, Risk deactivation accuracy
    zero-shot; independent questions; exact option-set match
    47.1
    Rank 9 of 28Leader Gemini-3.1-Pro 77.9
    Independent run
    Measured Aug 2026
  2. MedPIC, Guideline-following accuracy (GF)
    zero-shot; independent questions; exact option-set match
    73.9
    Rank 10 of 28Leader Gemini-3.1-Pro 87.7
    Independent run
    Measured Aug 2026
  3. MedPIC, Linked counterfactual pair accuracy
    zero-shot; independent questions; exact option-set match
    24.7
    Rank 10 of 28Leader Gemini-3.1-Pro 48.3
    Independent run
    Measured Aug 2026
  4. MedPIC, Counterfactual accuracy (CF)
    zero-shot; independent questions; exact option-set match
    48.1
    Rank 11 of 28Leader Gemini-3.1-Pro 69.9
    Independent run
    Measured Aug 2026
  5. MedPIC, Risk activation accuracy
    zero-shot; independent questions; exact option-set match
    54.9
    Rank 17 of 28Leader MedGemma-27B-Text 76.1
    Independent run
    Measured Aug 2026
  6. MedPIC, Published GF-minus-CF accuracy gap
    zero-shot; independent questions; exact option-set match
    25.9
    Descriptive measure, not ranked
    Independent run
    Measured Aug 2026

Sources

Open a line for the quote and page.

  1. 10MedPIC zero-shot; independent questions; exact option-set match 63.8
    Printed as 63.8Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row Qwen3.5-9B, column Overall; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    Qwen3.5-9B | 63.8 | 73.9 | 48.1 | 25.9 | 54.9 | 47.1 | 24.7
    Every result from this document
  2. 10MedPIC, Guideline-following accuracy (GF) zero-shot; independent questions; exact option-set match 73.9
    Printed as 73.9Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row Qwen3.5-9B, column GF; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    Qwen3.5-9B | 63.8 | 73.9 | 48.1 | 25.9 | 54.9 | 47.1 | 24.7
    Every result from this document
  3. 11MedPIC, Counterfactual accuracy (CF) zero-shot; independent questions; exact option-set match 48.1
    Printed as 48.1Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row Qwen3.5-9B, column CF; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    Qwen3.5-9B | 63.8 | 73.9 | 48.1 | 25.9 | 54.9 | 47.1 | 24.7
    Every result from this document
  4. 10MedPIC, Linked counterfactual pair accuracy zero-shot; independent questions; exact option-set match 24.7
    Printed as 24.7Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row Qwen3.5-9B, column Pair; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    Qwen3.5-9B | 63.8 | 73.9 | 48.1 | 25.9 | 54.9 | 47.1 | 24.7
    Every result from this document
  5. 17MedPIC, Risk activation accuracy zero-shot; independent questions; exact option-set match 54.9
    Printed as 54.9Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row Qwen3.5-9B, column Activation; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    Qwen3.5-9B | 63.8 | 73.9 | 48.1 | 25.9 | 54.9 | 47.1 | 24.7
    Every result from this document
  6. 9MedPIC, Risk deactivation accuracy zero-shot; independent questions; exact option-set match 47.1
    Printed as 47.1Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row Qwen3.5-9B, column Deactivation; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    Qwen3.5-9B | 63.8 | 73.9 | 48.1 | 25.9 | 54.9 | 47.1 | 24.7
    Every result from this document
  7. –MedPIC, Published GF-minus-CF accuracy gap zero-shot; independent questions; exact option-set match 25.9
    Printed as 25.9Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row Qwen3.5-9B, column Δ_GF−CF; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    Qwen3.5-9B | 63.8 | 73.9 | 48.1 | 25.9 | 54.9 | 47.1 | 24.7
    Every result from this document

Other Alibaba models: Lingshu-32B, Lingshu-7B, Qwen3-14B, Qwen3-235B-A22B-Instruct-2507, Qwen3-32B, Qwen3-4B, Qwen 3.5, Qwen3.5-27B, Qwen3.5-35B-A3B, Qwen3.5 397B A17B, Qwen 3.5 Flash, Qwen3.5-Plus, Qwen3.6-Max, Qwen3.6 Plus, Qwen 3.7 Max, Qwen3.7 Plus, Qwen 3.8 27B, Qwen3.8 Max, Qwen 3 Max Thinking, Qwen3-VL-235B-A22B, Qwen 3 VL Plus