HuatuoGPT-o1-70B
70BopenCompare with other models
- Clinical Benchmarks Index
- 24.8rank 122 of 148; 43.0 × 0.577 = 24.8, from 1 of 10 boards
- Boards
- 1 of 13
- Results
- 76 on secondary measures
- Latest measurement
- Aug 2026
Results
Each line is placed on its own board. Dark tick: the board leader.
Safety
- MedPICzero-shot; independent questions; exact option-set match55.2Rank 16 of 28Leader Gemini-3.1-Pro 80.7
Secondary measures
Safety
- 63.4Rank 8 of 28Leader MedGemma-27B-Text 76.1
- 63.7Rank 14 of 28Leader Gemini-3.1-Pro 87.7
- 42.1Rank 17 of 28Leader Gemini-3.1-Pro 69.9
- 14.6Rank 17 of 28Leader Gemini-3.1-Pro 48.3
- 23.5Rank 22 of 28Leader Gemini-3.1-Pro 77.9
- 21.7Descriptive measure, not ranked
Sources
Open a line for the quote and page.
16MedPIC zero-shot; independent questions; exact option-set match 55.2
Printed as 55.2Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set matchEvaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row HuatuoGPT-o1-70B, column Overall; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | PairHuatuoGPT-o1-70B | 55.2 | 63.7 | 42.1 | 21.7 | 63.4 | 23.5 | 14.6
Every result from this document14MedPIC, Guideline-following accuracy (GF) zero-shot; independent questions; exact option-set match 63.7
Printed as 63.7Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set matchEvaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row HuatuoGPT-o1-70B, column GF; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | PairHuatuoGPT-o1-70B | 55.2 | 63.7 | 42.1 | 21.7 | 63.4 | 23.5 | 14.6
Every result from this document17MedPIC, Counterfactual accuracy (CF) zero-shot; independent questions; exact option-set match 42.1
Printed as 42.1Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set matchEvaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row HuatuoGPT-o1-70B, column CF; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | PairHuatuoGPT-o1-70B | 55.2 | 63.7 | 42.1 | 21.7 | 63.4 | 23.5 | 14.6
Every result from this document17MedPIC, Linked counterfactual pair accuracy zero-shot; independent questions; exact option-set match 14.6
Printed as 14.6Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set matchEvaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row HuatuoGPT-o1-70B, column Pair; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | PairHuatuoGPT-o1-70B | 55.2 | 63.7 | 42.1 | 21.7 | 63.4 | 23.5 | 14.6
Every result from this document8MedPIC, Risk activation accuracy zero-shot; independent questions; exact option-set match 63.4
Printed as 63.4Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set matchEvaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row HuatuoGPT-o1-70B, column Activation; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | PairHuatuoGPT-o1-70B | 55.2 | 63.7 | 42.1 | 21.7 | 63.4 | 23.5 | 14.6
Every result from this document22MedPIC, Risk deactivation accuracy zero-shot; independent questions; exact option-set match 23.5
Printed as 23.5Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set matchEvaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row HuatuoGPT-o1-70B, column Deactivation; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | PairHuatuoGPT-o1-70B | 55.2 | 63.7 | 42.1 | 21.7 | 63.4 | 23.5 | 14.6
Every result from this document–MedPIC, Published GF-minus-CF accuracy gap zero-shot; independent questions; exact option-set match 21.7
Printed as 21.7Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set matchEvaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row HuatuoGPT-o1-70B, column Δ_GF−CF; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | PairHuatuoGPT-o1-70B | 55.2 | 63.7 | 42.1 | 21.7 | 63.4 | 23.5 | 14.6
Every result from this document
Other FreedomIntelligence models: HuatuoGPT-o1-8B