DeepSeek-V3.1
Released 21 Aug 2025128K context671B A37BopenCompare with other models
- Clinical Benchmarks Index
- 47.1rank 70 of 148; 81.6 × 0.577 = 47.1, from 1 of 10 boards
- Boards
- 1 of 13
- Results
- 1
- Latest measurement
- Jun 2026
Results
Each line is placed on its own board. Dark tick: the board leader.
EHR and workflow agents
- EHR-ComplexTable 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns0.560Rank 7 of 18Leader GPT-5.4 (high reasoning) 0.650
Sources
Open a line for the quote and page.
7EHR-Complex Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.560
Printed as 0.56Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turnsEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, DeepSeek-V3.1 row, final Avg. column (PDF p. 6).Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. DeepSeek-V3.1 | 0.81 | 0.33 | 0.7 | 0.27 | 0.65 | 0.24 | 0.89 | 0.68 | 0.82 | 0.39 | 0.72 | 0.21 | 0.56
Every result from this document
Other DeepSeek models: DeepSeek R1, DeepSeek-V3.2-Exp, DeepSeek V4.1 Flash, DeepSeek V4 Flash, DeepSeek V4 Flash 0731, DeepSeek V4 Pro