Clinical Benchmarks

DeepSeek V4 Pro

Released 24 Apr 2026$1.32 input, $3.96 output per million tokens1M context1.6T total / 49B activemitAlso written as DS-V4-Pro Max, DeepSeek-V4-Pro, deepseek/deepseek-v4-pro-0813, deepseek/deepseek-v4-pro, DeepSeek V4 Pro 0813, deepseek-v4, deepseek-v4-pro, DeepSeek-V4-pro, deepseek-v4-pro-0813, DeepSeek V4, DeepSeek-V4-Pro-MaxCompare with other models

Clinical Benchmarks Index
60.0rank 44 of 148; 60.0 × 1 = 60.0, from 4 of 10 boards
Boards
5 of 13
Results
166 on secondary measures
Latest measurement
Sep 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

Documentation and coding

  1. MedScribe (Vals AI)
    model ID deepseek/deepseek-v4-pro-0813; max_output_tokens=30000; reasoning_effort=max
    80.17
    Rank 55 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  2. MedScribe (Vals AI)
    model ID deepseek/deepseek-v4-pro; max_output_tokens=128000; reasoning_effort=max
    75.14
    Rank 77 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  3. MedCode (Vals AI)
    model ID deepseek/deepseek-v4-pro-0813; reasoning_effort=max; max_output_tokens=30000
    42.47
    Rank 47 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  4. MedCode (Vals AI)
    model ID deepseek/deepseek-v4-pro; reasoning_effort=max; max_output_tokens=128000
    40.45
    Rank 61 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

EHR and workflow agents

  1. PhysicianBench
    Pass@1 over 3 runs; Pass^3 6.0; shared FHIR tool harness; up to 100 turns; high reasoning when supported
    18.7
    Rank 15 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
    Official leaderboard
    Measured May 2026
  2. CHI-Bench
    Listed as openai-agents + deepseek-v4-pro
    All Domains pass@1; openai-agents harness
    14.2
    Rank 24 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026
  3. CHI-Bench
    Listed as hermes + deepseek-v4-pro
    All Domains pass@1; hermes harness
    13.8
    Rank 25 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026
  4. CHI-Bench
    Listed as openclaw + deepseek-v4-pro
    All Domains pass@1; openclaw harness
    11.1
    Rank 29 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026
  5. CHI-Bench
    Listed as deepagents + deepseek-v4-pro
    All Domains pass@1; deepagents harness
    10.7
    Rank 31 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026

Safety

  1. MedPIC
    zero-shot; independent questions; exact option-set match
    71.7
    Rank 4 of 28Leader Gemini-3.1-Pro 80.7
    Independent run
    Measured Aug 2026

Secondary measures

Safety

  1. MedPIC, Linked counterfactual pair accuracy
    zero-shot; independent questions; exact option-set match
    32.6
    Rank 3 of 28Leader Gemini-3.1-Pro 48.3
    Independent run
    Measured Aug 2026
  2. MedPIC, Guideline-following accuracy (GF)
    zero-shot; independent questions; exact option-set match
    81.3
    Rank 4 of 28Leader Gemini-3.1-Pro 87.7
    Independent run
    Measured Aug 2026
  3. MedPIC, Counterfactual accuracy (CF)
    zero-shot; independent questions; exact option-set match
    56.8
    Rank 4 of 28Leader Gemini-3.1-Pro 69.9
    Independent run
    Measured Aug 2026
  4. MedPIC, Risk activation accuracy
    zero-shot; independent questions; exact option-set match
    67.6
    Rank 5 of 28Leader MedGemma-27B-Text 76.1
    Independent run
    Measured Aug 2026
  5. MedPIC, Risk deactivation accuracy
    zero-shot; independent questions; exact option-set match
    50.0
    Rank 7 of 28Leader Gemini-3.1-Pro 77.9
    Independent run
    Measured Aug 2026
  6. MedPIC, Published GF-minus-CF accuracy gap
    zero-shot; independent questions; exact option-set match
    24.5
    Descriptive measure, not ranked
    Independent run
    Measured Aug 2026

Sources

Open a line for the quote and page.

  1. 55MedScribe (Vals AI) model ID deepseek/deepseek-v4-pro-0813; max_output_tokens=30000; reasoning_effo… 80.17
    Printed as 80.17%Official leaderboard, measured Sep 2026Configuration: model ID deepseek/deepseek-v4-pro-0813; max_output_tokens=30000; reasoning_effort=max
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 55 of 105 (DeepSeek V4 Pro 0813), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["deepseek/deepseek-v4-pro-0813"].
    55 | DeepSeek V4 Pro 0813 | 80.17%±2.00 | $1.32/$3.96 | 2m36s
    Every result from this document
  2. 77MedScribe (Vals AI) model ID deepseek/deepseek-v4-pro; max_output_tokens=128000; reasoning_effort=m… 75.14
    Printed as 75.14%Official leaderboard, measured Sep 2026Configuration: model ID deepseek/deepseek-v4-pro; max_output_tokens=128000; reasoning_effort=max
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 77 of 105 (DeepSeek V4), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["deepseek/deepseek-v4-pro"].
    77 | DeepSeek V4 | 75.14%±2.00 | $1.32/$3.96 | 5m46s
    Every result from this document
  3. 47MedCode (Vals AI) model ID deepseek/deepseek-v4-pro-0813; reasoning_effort=max; max_output_tokens… 42.47
    Printed as 42.47%Official leaderboard, measured Sep 2026Configuration: model ID deepseek/deepseek-v4-pro-0813; reasoning_effort=max; max_output_tokens=30000
    Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 47 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["deepseek/deepseek-v4-pro-0813"].
    47 | DeepSeek V4 Pro 0813 | 42.47%±2.16 | $1.32/$3.96 | 3m01s
    Every result from this document
  4. 61MedCode (Vals AI) model ID deepseek/deepseek-v4-pro; reasoning_effort=max; max_output_tokens=1280… 40.45
    Printed as 40.45%Official leaderboard, measured Sep 2026Configuration: model ID deepseek/deepseek-v4-pro; reasoning_effort=max; max_output_tokens=128000
    Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 61 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["deepseek/deepseek-v4-pro"].
    61 | DeepSeek V4 | 40.45%±2.12 | $1.32/$3.96 | 6m23s
    Every result from this document
  5. 4MedPIC zero-shot; independent questions; exact option-set match 71.7
    Printed as 71.7Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row DeepSeek-V4-Pro, column Overall; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    DeepSeek-V4-Pro | 71.7 | 81.3 | 56.8 | 24.5 | 67.6 | 50.0 | 32.6
    Every result from this document
  6. 15PhysicianBench Pass@1 over 3 runs; Pass^3 6.0; shared FHIR tool harness; up to 100 turns; high… 18.7
    Printed as 18.7 ± 2.9Official leaderboard, measured May 2026Configuration: Pass@1 over 3 runs; Pass^3 6.0; shared FHIR tool harness; up to 100 turns; high reasoning when supported
    PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026. arXiv HTML 2605.02240v1, Section 5.2, Table 2; DeepSeek V4-Pro row, Pass@1 column (PDF p. 8).
    Model | Pass@1 | Pass@3 | Pass^3 | #Turns DeepSeek V4-Pro | 18.7 ± 2.9 | 27.9 | 6.0 | 35.3
    Every result from this document
  7. 24CHI-Bench All Domains pass@1; openai-agents harness 14.2
    Printed as 14.2%Official leaderboard, measured May 2026Configuration: All Domains pass@1; openai-agents harness
    CHI-Bench leaderboard (actAVA) official leaderboard, actAVA, 12 Aug 2026. CHI-Bench v1.0.0, All Domains, rank 24, Agent/Model and Accuracy columns; PA/UM/CM follow; board last updated 2026-08-12; board Date column 2026-05-01; accessed 2026-09-30. Linked submission provenance: run 2026-05-03T08:35:38Z to 2026-05-04T06:59:15Z; submitted_at 2026-05-04T06:59:15Z.
    24 | openai-agents | deepseek-v4-pro | Open-source | 14.2% | 10.7% | 28.0% | 4.0% | 2026-05-01
    Every result from this document
  8. 25CHI-Bench All Domains pass@1; hermes harness 13.8
    Printed as 13.8%Official leaderboard, measured May 2026Configuration: All Domains pass@1; hermes harness
    CHI-Bench leaderboard (actAVA) official leaderboard, actAVA, 12 Aug 2026. CHI-Bench v1.0.0, All Domains, rank 25, Agent/Model and Accuracy columns; PA/UM/CM follow; board last updated 2026-08-12; board Date column 2026-05-01; accessed 2026-09-30. Linked submission provenance: run 2026-05-03T06:28:45Z to 2026-05-04T05:51:44Z; submitted_at 2026-05-04T05:51:44Z.
    25 | hermes | deepseek-v4-pro | Open-source | 13.8% | 8.0% | 25.3% | 8.0% | 2026-05-01
    Every result from this document
  9. 29CHI-Bench All Domains pass@1; openclaw harness 11.1
    Printed as 11.1%Official leaderboard, measured May 2026Configuration: All Domains pass@1; openclaw harness
    CHI-Bench leaderboard (actAVA) official leaderboard, actAVA, 12 Aug 2026. CHI-Bench v1.0.0, All Domains, rank 29, Agent/Model and Accuracy columns; PA/UM/CM follow; board last updated 2026-08-12; board Date column 2026-05-01; accessed 2026-09-30. Linked submission provenance: run 2026-05-02T09:01:51Z to 2026-05-04T03:00:01Z; submitted_at 2026-05-04T03:00:01Z.
    29 | openclaw | deepseek-v4-pro | Open-source | 11.1% | 14.7% | 12.0% | 6.7% | 2026-05-01
    Every result from this document
  10. 31CHI-Bench All Domains pass@1; deepagents harness 10.7
    Printed as 10.7%Official leaderboard, measured May 2026Configuration: All Domains pass@1; deepagents harness
    CHI-Bench leaderboard (actAVA) official leaderboard, actAVA, 12 Aug 2026. CHI-Bench v1.0.0, All Domains, rank 31, Agent/Model and Accuracy columns; PA/UM/CM follow; board last updated 2026-08-12; board Date column 2026-05-01; accessed 2026-09-30. Linked submission provenance: run 2026-05-03T07:35:56Z to 2026-05-05T01:34:20Z; submitted_at 2026-05-05T01:34:20Z.
    31 | deepagents | deepseek-v4-pro | Open-source | 10.7% | 14.7% | 10.7% | 6.7% | 2026-05-01
    Every result from this document
  11. 4MedPIC, Guideline-following accuracy (GF) zero-shot; independent questions; exact option-set match 81.3
    Printed as 81.3Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row DeepSeek-V4-Pro, column GF; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    DeepSeek-V4-Pro | 71.7 | 81.3 | 56.8 | 24.5 | 67.6 | 50.0 | 32.6
    Every result from this document
  12. 4MedPIC, Counterfactual accuracy (CF) zero-shot; independent questions; exact option-set match 56.8
    Printed as 56.8Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row DeepSeek-V4-Pro, column CF; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    DeepSeek-V4-Pro | 71.7 | 81.3 | 56.8 | 24.5 | 67.6 | 50.0 | 32.6
    Every result from this document
  13. 3MedPIC, Linked counterfactual pair accuracy zero-shot; independent questions; exact option-set match 32.6
    Printed as 32.6Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row DeepSeek-V4-Pro, column Pair; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    DeepSeek-V4-Pro | 71.7 | 81.3 | 56.8 | 24.5 | 67.6 | 50.0 | 32.6
    Every result from this document
  14. 5MedPIC, Risk activation accuracy zero-shot; independent questions; exact option-set match 67.6
    Printed as 67.6Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row DeepSeek-V4-Pro, column Activation; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    DeepSeek-V4-Pro | 71.7 | 81.3 | 56.8 | 24.5 | 67.6 | 50.0 | 32.6
    Every result from this document
  15. 7MedPIC, Risk deactivation accuracy zero-shot; independent questions; exact option-set match 50.0
    Printed as 50.0Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row DeepSeek-V4-Pro, column Deactivation; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    DeepSeek-V4-Pro | 71.7 | 81.3 | 56.8 | 24.5 | 67.6 | 50.0 | 32.6
    Every result from this document
  16. –MedPIC, Published GF-minus-CF accuracy gap zero-shot; independent questions; exact option-set match 24.5
    Printed as 24.5Independent run, measured Aug 2026Configuration: zero-shot; independent questions; exact option-set match
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning paper, Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang (The Hong Kong Polytechnic University; InfiX.ai; Sun Yat-sen University), 4 Aug 2026. arXiv:2608.03028v1, Section 4, Table 2 (#S4.T2), row DeepSeek-V4-Pro, column Δ_GF−CF; row order: Model | Overall | GF | CF | Δ_GF−CF | Activation | Deactivation | Pair
    DeepSeek-V4-Pro | 71.7 | 81.3 | 56.8 | 24.5 | 67.6 | 50.0 | 32.6
    Every result from this document

Other DeepSeek models: DeepSeek R1, DeepSeek-V3.1, DeepSeek-V3.2-Exp, DeepSeek V4.1 Flash, DeepSeek V4 Flash, DeepSeek V4 Flash 0731