Clinical Benchmarks

DeepSeek V4.1 Flash

Released 10 Sep 2026$0.3 input, $1.2 output per million tokens1M contextopenAlso written as deepseek/deepseek-v4.1-flashCompare with other models

Clinical Benchmarks Index
58.0rank 48 of 148; 71.1 × 0.816 = 58.0, from 2 of 10 boards
Boards
2 of 13
Results
2
Latest measurement
Sep 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

Documentation and coding

  1. MedScribe (Vals AI)
    model ID deepseek/deepseek-v4.1-flash; max_output_tokens=384000; reasoning_effort=high
    85.50
    Rank 24 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  2. MedCode (Vals AI)
    model ID deepseek/deepseek-v4.1-flash; reasoning_effort=high; max_output_tokens=384000
    41.17
    Rank 54 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

Sources

Open a line for the quote and page.

  1. 24MedScribe (Vals AI) model ID deepseek/deepseek-v4.1-flash; max_output_tokens=384000; reasoning_effo… 85.50
    Printed as 85.50%Official leaderboard, measured Sep 2026Configuration: model ID deepseek/deepseek-v4.1-flash; max_output_tokens=384000; reasoning_effort=high
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 24 of 105 (DeepSeek V4.1 Flash), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["deepseek/deepseek-v4.1-flash"].
    24 | DeepSeek V4.1 Flash | 85.50%±1.92 | $0.3/$1.2 | 53.83s
    Every result from this document
  2. 54MedCode (Vals AI) model ID deepseek/deepseek-v4.1-flash; reasoning_effort=high; max_output_tokens… 41.17
    Printed as 41.17%Official leaderboard, measured Sep 2026Configuration: model ID deepseek/deepseek-v4.1-flash; reasoning_effort=high; max_output_tokens=384000
    Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 54 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["deepseek/deepseek-v4.1-flash"].
    54 | DeepSeek V4.1 Flash | 41.17%±2.04 | $0.3/$1.2 | 35.13s
    Every result from this document

Other DeepSeek models: DeepSeek R1, DeepSeek-V3.1, DeepSeek-V3.2-Exp, DeepSeek V4 Flash, DeepSeek V4 Flash 0731, DeepSeek V4 Pro