Llama 3.1 8B
Released 23 Jul 2024128K context8BopenAlso written as Llama-3.1-8BCompare with other models
- Clinical Benchmarks Index
- 5.1rank 140 of 148; 8.9 × 0.577 = 5.1, from 1 of 10 boards
- Boards
- 1 of 13
- Results
- 76 on secondary measures
- Latest measurement
- Aug 2026
Results
Each line is placed on its own board. Dark tick: the board leader.
Safety
- MedPICzero-shot; independent questions; exact option-set match40.0Rank 25 of 28Leader Gemini-3.1-Pro 80.7
Secondary measures
Safety
- 54.9Rank 17 of 28Leader MedGemma-27B-Text 76.1
- 33.3Rank 21 of 28Leader Gemini-3.1-Pro 69.9
- 5.6Rank 24 of 28Leader Gemini-3.1-Pro 48.3
- 10.3Rank 25 of 28Leader Gemini-3.1-Pro 77.9
- 44.4Rank 26 of 28Leader Gemini-3.1-Pro 87.7
- 11.0Descriptive measure, not ranked
Sources
Open a line for the quote and page.
25MedPIC zero-shot; independent questions; exact option-set match 40.0
Printed as 40.0Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set matchEvaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row Llama-3.1-8B, column Overall; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | PairLlama-3.1-8B | 40.0 | 44.4 | 33.3 | 11.0 | 54.9 | 10.3 | 5.6
Every result from this document26MedPIC, Guideline-following accuracy (GF) zero-shot; independent questions; exact option-set match 44.4
Printed as 44.4Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set matchEvaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row Llama-3.1-8B, column GF; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | PairLlama-3.1-8B | 40.0 | 44.4 | 33.3 | 11.0 | 54.9 | 10.3 | 5.6
Every result from this document21MedPIC, Counterfactual accuracy (CF) zero-shot; independent questions; exact option-set match 33.3
Printed as 33.3Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set matchEvaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row Llama-3.1-8B, column CF; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | PairLlama-3.1-8B | 40.0 | 44.4 | 33.3 | 11.0 | 54.9 | 10.3 | 5.6
Every result from this document24MedPIC, Linked counterfactual pair accuracy zero-shot; independent questions; exact option-set match 5.6
Printed as 5.6Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set matchEvaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row Llama-3.1-8B, column Pair; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | PairLlama-3.1-8B | 40.0 | 44.4 | 33.3 | 11.0 | 54.9 | 10.3 | 5.6
Every result from this document17MedPIC, Risk activation accuracy zero-shot; independent questions; exact option-set match 54.9
Printed as 54.9Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set matchEvaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row Llama-3.1-8B, column Activation; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | PairLlama-3.1-8B | 40.0 | 44.4 | 33.3 | 11.0 | 54.9 | 10.3 | 5.6
Every result from this document25MedPIC, Risk deactivation accuracy zero-shot; independent questions; exact option-set match 10.3
Printed as 10.3Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set matchEvaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row Llama-3.1-8B, column Deactivation; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | PairLlama-3.1-8B | 40.0 | 44.4 | 33.3 | 11.0 | 54.9 | 10.3 | 5.6
Every result from this document–MedPIC, Published GF-minus-CF accuracy gap zero-shot; independent questions; exact option-set match 11.0
Printed as 11.0Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set matchEvaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row Llama-3.1-8B, column Δ_GF−CF; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | PairLlama-3.1-8B | 40.0 | 44.4 | 33.3 | 11.0 | 54.9 | 10.3 | 5.6
Every result from this document
Other Meta models: Llama 3.1 70B, Llama 4 Maverick, Llama 4 Scout, Muse Spark, Muse Spark 1.1, Muse Spark 1.2