Clinical Benchmarks

HuatuoGPT-o1-8B

8BopenCompare with other models

Clinical Benchmarks Index
10.7rank 133 of 148; 18.6 × 0.577 = 10.7, from 1 of 10 boards
Boards
1 of 13
Results
76 on secondary measures
Latest measurement
Aug 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

Safety

  1. MedPIC
    zero-shot; independent questions; exact option-set match
    44.3
    Rank 20 of 28Leader Gemini-3.1-Pro 80.7
    Independent run
    Measured Aug 2026

Secondary measures

Safety

  1. MedPIC, Linked counterfactual pair accuracy
    zero-shot; independent questions; exact option-set match
    13.5
    Rank 19 of 28Leader Gemini-3.1-Pro 48.3
    Independent run
    Measured Aug 2026
  2. MedPIC, Counterfactual accuracy (CF)
    zero-shot; independent questions; exact option-set match
    35.5
    Rank 20 of 28Leader Gemini-3.1-Pro 69.9
    Independent run
    Measured Aug 2026
  3. MedPIC, Guideline-following accuracy (GF)
    zero-shot; independent questions; exact option-set match
    50.0
    Rank 22 of 28Leader Gemini-3.1-Pro 87.7
    Independent run
    Measured Aug 2026
  4. MedPIC, Risk activation accuracy
    zero-shot; independent questions; exact option-set match
    49.3
    Rank 22 of 28Leader MedGemma-27B-Text 76.1
    Independent run
    Measured Aug 2026
  5. MedPIC, Risk deactivation accuracy
    zero-shot; independent questions; exact option-set match
    20.6
    Rank 23 of 28Leader Gemini-3.1-Pro 77.9
    Independent run
    Measured Aug 2026
  6. MedPIC, Published GF-minus-CF accuracy gap
    zero-shot; independent questions; exact option-set match
    14.5
    Descriptive measure, not ranked
    Independent run
    Measured Aug 2026

Sources

Open a line for the quote and page.

  1. 20MedPIC zero-shot; independent questions; exact option-set match 44.3
    Printed as 44.3Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row HuatuoGPT-o1-8B, column Overall; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    HuatuoGPT-o1-8B | 44.3 | 50.0 | 35.5 | 14.5 | 49.3 | 20.6 | 13.5
    Every result from this document
  2. 22MedPIC, Guideline-following accuracy (GF) zero-shot; independent questions; exact option-set match 50.0
    Printed as 50.0Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row HuatuoGPT-o1-8B, column GF; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    HuatuoGPT-o1-8B | 44.3 | 50.0 | 35.5 | 14.5 | 49.3 | 20.6 | 13.5
    Every result from this document
  3. 20MedPIC, Counterfactual accuracy (CF) zero-shot; independent questions; exact option-set match 35.5
    Printed as 35.5Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row HuatuoGPT-o1-8B, column CF; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    HuatuoGPT-o1-8B | 44.3 | 50.0 | 35.5 | 14.5 | 49.3 | 20.6 | 13.5
    Every result from this document
  4. 19MedPIC, Linked counterfactual pair accuracy zero-shot; independent questions; exact option-set match 13.5
    Printed as 13.5Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row HuatuoGPT-o1-8B, column Pair; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    HuatuoGPT-o1-8B | 44.3 | 50.0 | 35.5 | 14.5 | 49.3 | 20.6 | 13.5
    Every result from this document
  5. 22MedPIC, Risk activation accuracy zero-shot; independent questions; exact option-set match 49.3
    Printed as 49.3Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row HuatuoGPT-o1-8B, column Activation; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    HuatuoGPT-o1-8B | 44.3 | 50.0 | 35.5 | 14.5 | 49.3 | 20.6 | 13.5
    Every result from this document
  6. 23MedPIC, Risk deactivation accuracy zero-shot; independent questions; exact option-set match 20.6
    Printed as 20.6Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row HuatuoGPT-o1-8B, column Deactivation; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    HuatuoGPT-o1-8B | 44.3 | 50.0 | 35.5 | 14.5 | 49.3 | 20.6 | 13.5
    Every result from this document
  7. –MedPIC, Published GF-minus-CF accuracy gap zero-shot; independent questions; exact option-set match 14.5
    Printed as 14.5Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row HuatuoGPT-o1-8B, column Δ_GF−CF; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    HuatuoGPT-o1-8B | 44.3 | 50.0 | 35.5 | 14.5 | 49.3 | 20.6 | 13.5
    Every result from this document

Other FreedomIntelligence models: HuatuoGPT-o1-70B