Clinical Benchmarks

MedScribe (Vals AI)

Quality and compliance of SOAP notes generated from clinical visits, scored against documentation rubrics.

Vals AIOfficial page
More about this boardLess

Clinical documentation support: quality of SOAP notes generated from clinical visits, scored against rubrics for documentation quality and compliance.

Vals AI self-runs; 105 models on the board, last updated September 29, 2026. Top scores cluster near 90, so the leaders sit close to the ceiling.

Published by Vals AI (dataset with Protege), released Feb 2026. 100 rubric-scored SOAP-note cases. Percentage accuracy 0-100, higher better.

Rows on the board
105
Board last updated
29 Sep 2026
Rows
105
Models
94 of 105on the board
Labs
19
Last measured
Sep 2026
Leader
91.43Claude Opus 5.5

Ranking

Percentage accuracy 0-100, higher better

  • Anthropic
  • OpenAI
  • Meta
  • Moonshot AI
  • Other labs
All official leaderboard
  1. 1Claude Opus 5.5model ID anthropic/claude-opus-5-591.43
  2. 2Claude Fable 5.191.29
  3. 3Claude Sonnet 5.5model ID anthropic/claude-sonnet-5-591.10
  4. 4Claude Opus 590.98
  5. 5Muse Spark 1.290.06
  6. 6Grok 4.7model ID grok/grok-489.38
  7. 7GLM 5.3 Flashmodel ID zai/glm-588.94
  8. 8Muse Spark 1.188.89
  9. 9GLM 5.3model ID zai/glm-588.81
  10. 10Claude Fable 588.52
  11. 11MiMo V2.6 Promodel ID xiaomi/mimo-v288.31
  12. 12GPT 5.188.09
  13. 13Kimi K387.96
  14. 14GPT-6 Astra87.91
  15. 15MiniMax-M387.25
  16. 16Grok 4.5model ID grok/grok-486.88
  17. 17GPT 5.586.87
  18. 18Claude Opus 4.6 (Nonthinking)86.74
  19. 19Grok 4.686.53
  20. 20GPT-6.1 Solmodel ID openai/gpt-686.45

20 of 105 rows

Rows and sources

Open a row for the quote, the page and the document.

  1. 1Claude Opus 5.5 model ID anthropic/claude-opus-5-5; temperature=1; max_output_tokens=128000; co… 91.43
    Printed as 91.43%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-5-5; temperature=1; max_output_tokens=128000; compute_effort=max
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 1 of 105 (Claude Opus 5.5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-opus-5-5"].
    1 | Claude Opus 5.5 | 91.43%±1.93 | $4/$20 | 7m22s
    Every result from this document
  2. 2Claude Fable 5.1 91.29
    Printed as 91.29%Official leaderboard, measured Sep 2026
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 2 of 105 (Claude Fable 5.1), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-fable-5-1"].
    2 | Claude Fable 5.1 | 91.29%±1.95 | $10/$50 | 3m08s
    Every result from this document
  3. 3Claude Sonnet 5.5 model ID anthropic/claude-sonnet-5-5; temperature=1; max_output_tokens=128000;… 91.10
    Printed as 91.10%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-sonnet-5-5; temperature=1; max_output_tokens=128000; compute_effort=max
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 3 of 105 (Claude Sonnet 5.5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-sonnet-5-5"].
    3 | Claude Sonnet 5.5 | 91.10%±1.96 | $2/$10 | 5m07s
    Every result from this document
  4. 4Claude Opus 5 90.98
    Printed as 90.98%Official leaderboard, measured Sep 2026
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 4 of 105 (Claude Opus 5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-opus-5"].
    4 | Claude Opus 5 | 90.98%±1.92 | $5/$25 | 76.56s
    Every result from this document
  5. 5Muse Spark 1.2 90.06
    Printed as 90.06%Official leaderboard, measured Sep 2026
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 5 of 105 (Muse Spark 1.2), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["meta/muse_spark_1_2"].
    5 | Muse Spark 1.2 | 90.06%±1.96 | $1.25/$4.25 | 61.46s
    Every result from this document
  6. 6Grok 4.7 model ID grok/grok-4.7; temperature=1; top_p=0.95; reasoning_effort=xhigh 89.38
    Printed as 89.38%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4.7; temperature=1; top_p=0.95; reasoning_effort=xhigh
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 6 of 105 (Grok 4.7), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["grok/grok-4.7"].
    6 | Grok 4.7 | 89.38%±1.89 | $2/$6 | 2m33s
    Every result from this document
  7. 7GLM 5.3 Flash model ID zai/glm-5.3-flash; temperature=1; top_p=0.95; max_output_tokens=30000;… 88.94
    Printed as 88.94%Official leaderboard, measured Sep 2026Configuration: model ID zai/glm-5.3-flash; temperature=1; top_p=0.95; max_output_tokens=30000; reasoning_effort=max
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 7 of 105 (GLM 5.3 Flash), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["zai/glm-5.3-flash"].
    7 | GLM 5.3 Flash | 88.94%±1.91 | $0.075/$0.25 | 86.98s
    Every result from this document
  8. 8Muse Spark 1.1 88.89
    Printed as 88.89%Official leaderboard, measured Sep 2026
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 8 of 105 (Muse Spark 1.1), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["meta/muse_spark_1_1"].
    8 | Muse Spark 1.1 | 88.89%±1.95 | $1.25/$4.25 | 63.34s
    Every result from this document
  9. 9GLM 5.3 model ID zai/glm-5.3; temperature=1; top_p=0.95; max_output_tokens=30000; reaso… 88.81
    Printed as 88.81%Official leaderboard, measured Sep 2026Configuration: model ID zai/glm-5.3; temperature=1; top_p=0.95; max_output_tokens=30000; reasoning_effort=max
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 9 of 105 (GLM 5.3), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["zai/glm-5.3"].
    9 | GLM 5.3 | 88.81%±2.00 | $1.4/$4.4 | 2m02s
    Every result from this document
  10. 10Claude Fable 5 88.52
    Printed as 88.52%Official leaderboard, measured Sep 2026
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 10 of 105 (Claude Fable 5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-fable-5"].
    10 | Claude Fable 5 | 88.52%±1.95 | $10/$50 | 119.47s
    Every result from this document
  11. 11MiMo V2.6 Pro model ID xiaomi/mimo-v2.6-pro; temperature=1; top_p=0.95; max_output_tokens=128… 88.31
    Printed as 88.31%Official leaderboard, measured Sep 2026Configuration: model ID xiaomi/mimo-v2.6-pro; temperature=1; top_p=0.95; max_output_tokens=128000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 11 of 105 (MiMo V2.6 Pro), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["xiaomi/mimo-v2.6-pro"].
    11 | MiMo V2.6 Pro | 88.31%±1.94 | $0.435/$0.87 | 2m53s
    Every result from this document
  12. 12GPT 5.1 88.09
    Printed as 88.09%Official leaderboard, measured Sep 2026
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 12 of 105 (GPT 5.1), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-5.1-2025-11-13"].
    12 | GPT 5.1 | 88.09%±1.94 | $1.25/$10 | 77.98s
    Every result from this document
  13. 13Kimi K3 87.96
    Printed as 87.96%Official leaderboard, measured Sep 2026
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 13 of 105 (Kimi K3), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["kimi/kimi-k3"].
    13 | Kimi K3 | 87.96%±1.89 | $3/$15 | 2m17s
    Every result from this document
  14. 14GPT-6 Astra 87.91
    Printed as 87.91%Official leaderboard, measured Sep 2026
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 14 of 105 (GPT-6 Astra), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-6-astra"].
    14 | GPT-6 Astra | 87.91%±1.94 | $10/$50 | 2m34s
    Every result from this document
  15. 15MiniMax-M3 87.25
    Printed as 87.25%Official leaderboard, measured Sep 2026
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 15 of 105 (MiniMax-M3), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["minimax/MiniMax-M3"].
    15 | MiniMax-M3 | 87.25%±1.96 | $0.6/$2.4 | 2m04s
    Every result from this document
  16. 16Grok 4.5 model ID grok/grok-4.5; temperature=1; top_p=0.95; max_output_tokens=30000; rea… 86.88
    Printed as 86.88%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4.5; temperature=1; top_p=0.95; max_output_tokens=30000; reasoning_effort=high
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 16 of 105 (Grok 4.5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["grok/grok-4.5"].
    16 | Grok 4.5 | 86.88%±1.94 | $2/$6 | 11m45s
    Every result from this document
  17. 17GPT 5.5 86.87
    Printed as 86.87%Official leaderboard, measured Sep 2026
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 17 of 105 (GPT 5.5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-5.5"].
    17 | GPT 5.5 | 86.87%±1.93 | $5/$30 | 2m13s
    Every result from this document
  18. 18Claude Opus 4.6 (Nonthinking) 86.74
    Printed as 86.74%Official leaderboard, measured Sep 2026
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 18 of 105 (Claude Opus 4.6 (Nonthinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-opus-4-6"].
    18 | Claude Opus 4.6 (Nonthinking) | 86.74%±1.94 | $5/$25 | 54.32s
    Every result from this document
  19. 19Grok 4.6 86.53
    Printed as 86.53%Official leaderboard, measured Sep 2026
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 19 of 105 (Grok 4.6), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["grok/grok-4.6"].
    19 | Grok 4.6 | 86.53%±1.96 | $2/$6 | 70.94s
    Every result from this document
  20. 20GPT-6.1 Sol model ID openai/gpt-6.1-sol; max_output_tokens=128000; reasoning_effort=max 86.45
    Printed as 86.45%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-6.1-sol; max_output_tokens=128000; reasoning_effort=max
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 20 of 105 (GPT-6.1 Sol), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-6.1-sol"].
    20 | GPT-6.1 Sol | 86.45%±1.90 | $2/$10 | 2m50s
    Every result from this document
  21. 21Claude Opus 4.6 (Thinking) model ID anthropic/claude-opus-4-6-thinking; temperature=1; max_output_tokens=3… 86.13
    Printed as 86.13%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-6-thinking; temperature=1; max_output_tokens=30000; compute_effort=max
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 21 of 105 (Claude Opus 4.6 (Thinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-opus-4-6-thinking"].
    21 | Claude Opus 4.6 (Thinking) | 86.13%±1.94 | $5/$25 | 2m08s
    Every result from this document
  22. 22Muse Spark model ID meta/muse_spark; temperature=1; max_output_tokens=30000 85.90
    Printed as 85.90%Official leaderboard, measured Sep 2026Configuration: model ID meta/muse_spark; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 22 of 105 (Muse Spark), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["meta/muse_spark"].
    22 | Muse Spark | 85.90%±1.85 | N/A | 3m11s
    Every result from this document
  23. 23Claude Opus 4.8 model ID anthropic/claude-opus-4-8; temperature=1; max_output_tokens=30000; com… 85.75
    Printed as 85.75%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-8; temperature=1; max_output_tokens=30000; compute_effort=max
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 23 of 105 (Claude Opus 4.8), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-opus-4-8"].
    23 | Claude Opus 4.8 | 85.75%±1.93 | $5/$25 | 82.36s
    Every result from this document
  24. 24DeepSeek V4.1 Flash model ID deepseek/deepseek-v4.1-flash; max_output_tokens=384000; reasoning_effo… 85.50
    Printed as 85.50%Official leaderboard, measured Sep 2026Configuration: model ID deepseek/deepseek-v4.1-flash; max_output_tokens=384000; reasoning_effort=high
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 24 of 105 (DeepSeek V4.1 Flash), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["deepseek/deepseek-v4.1-flash"].
    24 | DeepSeek V4.1 Flash | 85.50%±1.92 | $0.3/$1.2 | 53.83s
    Every result from this document
  25. 25Inkling model ID thinkingmachines/inkling; temperature=1; top_p=1; max_output_tokens=30… 85.41
    Printed as 85.41%Official leaderboard, measured Sep 2026Configuration: model ID thinkingmachines/inkling; temperature=1; top_p=1; max_output_tokens=30000; reasoning_effort=0.99
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 25 of 105 (Inkling), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["thinkingmachines/inkling"].
    25 | Inkling | 85.41%±1.84 | $1/$4.05 | 4m05s
    Every result from this document
  26. 26Claude Opus 4.5 (Thinking) model ID anthropic/claude-opus-4-5-20251101-thinking; temperature=1; max_output… 85.32
    Printed as 85.32%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-5-20251101-thinking; temperature=1; max_output_tokens=30000; compute_effort=high
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 26 of 105 (Claude Opus 4.5 (Thinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-opus-4-5-20251101-thinking"].
    26 | Claude Opus 4.5 (Thinking) | 85.32%±1.90 | $5/$25 | 72.11s
    Every result from this document
  27. 27MiMo V2.6 Flash model ID xiaomi/mimo-v2.6-flash; temperature=1; top_p=0.95; max_output_tokens=1… 85.28
    Printed as 85.28%Official leaderboard, measured Sep 2026Configuration: model ID xiaomi/mimo-v2.6-flash; temperature=1; top_p=0.95; max_output_tokens=128000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 27 of 105 (MiMo V2.6 Flash), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["xiaomi/mimo-v2.6-flash"].
    27 | MiMo V2.6 Flash | 85.28%±1.99 | $0.14/$0.28 | 74.48s
    Every result from this document
  28. 28Claude Haiku 4.5 (Thinking) model ID anthropic/claude-haiku-4-5-20251001-thinking; temperature=1; max_outpu… 85.23
    Printed as 85.23%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-haiku-4-5-20251001-thinking; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 29 of 105 (Claude Haiku 4.5 (Thinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-haiku-4-5-20251001-thinking"].
    29 | Claude Haiku 4.5 (Thinking) | 85.23%±1.90 | $1/$5 | 66.20s
    Every result from this document
  29. 28GPT-5.6 Sol model ID openai/gpt-5.6-sol; max_output_tokens=30000; reasoning_effort=max 85.23
    Printed as 85.23%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5.6-sol; max_output_tokens=30000; reasoning_effort=max
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 28 of 105 (GPT-5.6 Sol), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-5.6-sol"].
    28 | GPT-5.6 Sol | 85.23%±1.97 | $4/$20 | 94.40s
    Every result from this document
  30. 30Qwen 3.8 Max model ID alibaba/qwen3.8-max; temperature=1; max_output_tokens=30000 84.95
    Printed as 84.95%Official leaderboard, measured Sep 2026Configuration: model ID alibaba/qwen3.8-max; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 30 of 105 (Qwen 3.8 Max), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["alibaba/qwen3.8-max"].
    30 | Qwen 3.8 Max | 84.95%±2.00 | $2/$6 | 4m11s
    Every result from this document
  31. 31Claude Sonnet 4.5 (Nonthinking) model ID anthropic/claude-sonnet-4-5-20250929; temperature=1; max_output_tokens… 84.52
    Printed as 84.52%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-sonnet-4-5-20250929; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 31 of 105 (Claude Sonnet 4.5 (Nonthinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-sonnet-4-5-20250929"].
    31 | Claude Sonnet 4.5 (Nonthinking) | 84.52%±1.93 | $3/$15 | 44.42s
    Every result from this document
  32. 32Gemini 3.8 Flash model ID google/gemini-3.8-flash; temperature=1; max_output_tokens=65536; reaso… 84.50
    Printed as 84.50%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-3.8-flash; temperature=1; max_output_tokens=65536; reasoning_effort=high
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 32 of 105 (Gemini 3.8 Flash), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-3.8-flash"].
    32 | Gemini 3.8 Flash | 84.50%±1.94 | $1.5/$7.5 | 21.14s
    Every result from this document
  33. 33GPT 5.2 model ID openai/gpt-5.2-2025-12-11; max_output_tokens=30000; reasoning_effort=x… 84.39
    Printed as 84.39%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5.2-2025-12-11; max_output_tokens=30000; reasoning_effort=xhigh
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 34 of 105 (GPT 5.2), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-5.2-2025-12-11"].
    34 | GPT 5.2 | 84.39%±1.86 | $1.75/$14 | 2m06s
    Every result from this document
  34. 33GPT-5.6 Luna model ID openai/gpt-5.6-luna; max_output_tokens=30000; reasoning_effort=max 84.39
    Printed as 84.39%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5.6-luna; max_output_tokens=30000; reasoning_effort=max
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 33 of 105 (GPT-5.6 Luna), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-5.6-luna"].
    33 | GPT-5.6 Luna | 84.39%±2.58 | $0.2/$1.2 | 116.89s
    Every result from this document
  35. 35Inkling Small model ID thinkingmachines/inkling-small; temperature=1; top_p=1; max_output_tok… 84.11
    Printed as 84.11%Official leaderboard, measured Sep 2026Configuration: model ID thinkingmachines/inkling-small; temperature=1; top_p=1; max_output_tokens=30000; reasoning_effort=0.99
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 35 of 105 (Inkling Small), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["thinkingmachines/inkling-small"].
    35 | Inkling Small | 84.11%±1.87 | $0.3/$1.2 | 3m50s
    Every result from this document
  36. 36Claude Sonnet 4.5 (Thinking) model ID anthropic/claude-sonnet-4-5-20250929-thinking; temperature=1; max_outp… 84.10
    Printed as 84.10%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-sonnet-4-5-20250929-thinking; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 36 of 105 (Claude Sonnet 4.5 (Thinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-sonnet-4-5-20250929-thinking"].
    36 | Claude Sonnet 4.5 (Thinking) | 84.10%±1.87 | $3/$15 | 67.92s
    Every result from this document
  37. 37Gemini 3.7 Flash model ID google/gemini-3.7-flash; temperature=1; max_output_tokens=30000; reaso… 83.94
    Printed as 83.94%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-3.7-flash; temperature=1; max_output_tokens=30000; reasoning_effort=high
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 37 of 105 (Gemini 3.7 Flash), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-3.7-flash"].
    37 | Gemini 3.7 Flash | 83.94%±2.00 | $1.5/$7.5 | 16.82s
    Every result from this document
  38. 38Qwen 3.8 27B model ID alibaba/qwen3.8-27b; temperature=1; top_p=0.95; max_output_tokens=3000… 83.85
    Printed as 83.85%Official leaderboard, measured Sep 2026Configuration: model ID alibaba/qwen3.8-27b; temperature=1; top_p=0.95; max_output_tokens=30000; reasoning_effort=xhigh
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 38 of 105 (Qwen 3.8 27B), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["alibaba/qwen3.8-27b"].
    38 | Qwen 3.8 27B | 83.85%±1.98 | $0.5/$3 | 89.70s
    Every result from this document
  39. 39MiMo V2.5 Pro model ID xiaomi/mimo-v2.5-pro; temperature=1; top_p=0.95; max_output_tokens=300… 83.73
    Printed as 83.73%Official leaderboard, measured Sep 2026Configuration: model ID xiaomi/mimo-v2.5-pro; temperature=1; top_p=0.95; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 39 of 105 (MiMo V2.5 Pro), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["xiaomi/mimo-v2.5-pro"].
    39 | MiMo V2.5 Pro | 83.73%±2.06 | $0.435/$0.87 | 90.71s
    Every result from this document
  40. 40GPT-6 Luna model ID openai/gpt-6-luna; max_output_tokens=128000; reasoning_effort=max 83.71
    Printed as 83.71%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-6-luna; max_output_tokens=128000; reasoning_effort=max
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 40 of 105 (GPT-6 Luna), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-6-luna"].
    40 | GPT-6 Luna | 83.71%±1.95 | $0.1/$0.5 | 2m50s
    Every result from this document
  41. 41GPT 5 model ID openai/gpt-5-2025-08-07; max_output_tokens=30000; reasoning_effort=high 83.65
    Printed as 83.65%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5-2025-08-07; max_output_tokens=30000; reasoning_effort=high
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 41 of 105 (GPT 5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-5-2025-08-07"].
    41 | GPT 5 | 83.65%±1.94 | $1.25/$10 | 3m05s
    Every result from this document
  42. 42Hy4 Preview model ID tencent/hy4-preview; temperature=1; top_p=1; max_output_tokens=64000 83.60
    Printed as 83.60%Official leaderboard, measured Sep 2026Configuration: model ID tencent/hy4-preview; temperature=1; top_p=1; max_output_tokens=64000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 42 of 105 (Hy4 Preview), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["tencent/hy4-preview"].
    42 | Hy4 Preview | 83.60%±2.06 | $0.834/$2.501 | 5m50s
    Every result from this document
  43. 43GLM 5.2 model ID zai/glm-5.2; temperature=1; max_output_tokens=30000 83.53
    Printed as 83.53%Official leaderboard, measured Sep 2026Configuration: model ID zai/glm-5.2; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 43 of 105 (GLM 5.2), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["zai/glm-5.2"].
    43 | GLM 5.2 | 83.53%±2.00 | $1.4/$4.4 | 2m18s
    Every result from this document
  44. 44Claude Opus 4.5 (Nonthinking) model ID anthropic/claude-opus-4-5-20251101; temperature=1; max_output_tokens=3… 83.25
    Printed as 83.25%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-5-20251101; temperature=1; max_output_tokens=30000; compute_effort=high
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 44 of 105 (Claude Opus 4.5 (Nonthinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-opus-4-5-20251101"].
    44 | Claude Opus 4.5 (Nonthinking) | 83.25%±1.93 | $5/$25 | 43.30s
    Every result from this document
  45. 45Gemini 2.5 Flash (7/17) (Thinking) model ID google/gemini-2.5-flash-thinking; temperature=1; max_output_tokens=300… 82.98
    Printed as 82.98%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-2.5-flash-thinking; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 45 of 105 (Gemini 2.5 Flash (7/17) (Thinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-2.5-flash-thinking"].
    45 | Gemini 2.5 Flash (7/17) (Thinking) | 82.98%±1.91 | $0.3/$2.5 | 22.53s
    Every result from this document
  46. 46Claude Opus 4.7 model ID anthropic/claude-opus-4-7; temperature=1; max_output_tokens=30000; com… 82.95
    Printed as 82.95%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-7; temperature=1; max_output_tokens=30000; compute_effort=max
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 46 of 105 (Claude Opus 4.7), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-opus-4-7"].
    46 | Claude Opus 4.7 | 82.95%±1.98 | $5/$25 | 67.50s
    Every result from this document
  47. 47Gemini 2.5 Flash (7/17) (Nonthinking) model ID google/gemini-2.5-flash; temperature=1; max_output_tokens=30000 82.87
    Printed as 82.87%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-2.5-flash; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 47 of 105 (Gemini 2.5 Flash (7/17) (Nonthinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-2.5-flash"].
    47 | Gemini 2.5 Flash (7/17) (Nonthinking) | 82.87%±1.91 | $0.3/$2.5 | 22.79s
    Every result from this document
  48. 47GPT-5.6 Terra model ID openai/gpt-5.6-terra; max_output_tokens=30000; reasoning_effort=xhigh 82.87
    Printed as 82.87%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5.6-terra; max_output_tokens=30000; reasoning_effort=xhigh
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 48 of 105 (GPT-5.6 Terra), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-5.6-terra"].
    48 | GPT-5.6 Terra | 82.87%±1.95 | $2/$12 | 35.53s
    Every result from this document
  49. 49GPT-6 Sol model ID openai/gpt-6-sol; max_output_tokens=128000; reasoning_effort=max 82.03
    Printed as 82.03%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-6-sol; max_output_tokens=128000; reasoning_effort=max
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 49 of 105 (GPT-6 Sol), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-6-sol"].
    49 | GPT-6 Sol | 82.03%±1.94 | $2/$10 | 81.14s
    Every result from this document
  50. 50Grok 4 Fast (Reasoning) model ID grok/grok-4-fast-reasoning; temperature=1; top_p=0.95; max_output_toke… 81.63
    Printed as 81.63%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 50 of 105 (Grok 4 Fast (Reasoning)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["grok/grok-4-fast-reasoning"].
    50 | Grok 4 Fast (Reasoning) | 81.63%±2.14 | $0.2/$0.5 | 12.42s
    Every result from this document
  51. 51Ling 3.0 Flash model ID ant/ling-3.0-flash-2607; temperature=1; top_p=0.95; max_output_tokens=… 80.90
    Printed as 80.90%Official leaderboard, measured Sep 2026Configuration: model ID ant/ling-3.0-flash-2607; temperature=1; top_p=0.95; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 51 of 105 (Ling 3.0 Flash), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["ant/ling-3.0-flash-2607"].
    51 | Ling 3.0 Flash | 80.90%±2.04 | $0.075/$0.22 | 12.12s
    Every result from this document
  52. 52MiniMax-M2.1 model ID minimax/MiniMax-M2.1; temperature=1; top_p=0.95; max_output_tokens=300… 80.78
    Printed as 80.78%Official leaderboard, measured Sep 2026Configuration: model ID minimax/MiniMax-M2.1; temperature=1; top_p=0.95; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 52 of 105 (MiniMax-M2.1), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["minimax/MiniMax-M2.1"].
    52 | MiniMax-M2.1 | 80.78%±1.83 | $0.3/$1.2 | 53.16s
    Every result from this document
  53. 53GPT 5 Mini model ID openai/gpt-5-mini-2025-08-07; max_output_tokens=30000; reasoning_effor… 80.58
    Printed as 80.58%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5-mini-2025-08-07; max_output_tokens=30000; reasoning_effort=high
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 53 of 105 (GPT 5 Mini), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-5-mini-2025-08-07"].
    53 | GPT 5 Mini | 80.58%±1.92 | $0.25/$2 | 4m36s
    Every result from this document
  54. 54DeepSeek V4 Flash 0731 model ID deepseek/deepseek-v4-flash-0731; max_output_tokens=30000; reasoning_ef… 80.36
    Printed as 80.36%Official leaderboard, measured Sep 2026Configuration: model ID deepseek/deepseek-v4-flash-0731; max_output_tokens=30000; reasoning_effort=high
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 54 of 105 (DeepSeek V4 Flash 0731), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["deepseek/deepseek-v4-flash-0731"].
    54 | DeepSeek V4 Flash 0731 | 80.36%±1.97 | $0.44/$1.32 | 77.55s
    Every result from this document
  55. 55DeepSeek V4 Pro 0813 model ID deepseek/deepseek-v4-pro-0813; max_output_tokens=30000; reasoning_effo… 80.17
    Printed as 80.17%Official leaderboard, measured Sep 2026Configuration: model ID deepseek/deepseek-v4-pro-0813; max_output_tokens=30000; reasoning_effort=max
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 55 of 105 (DeepSeek V4 Pro 0813), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["deepseek/deepseek-v4-pro-0813"].
    55 | DeepSeek V4 Pro 0813 | 80.17%±2.00 | $1.32/$3.96 | 2m36s
    Every result from this document
  56. 56MiniMax-M2.7 model ID minimax/MiniMax-M2.7; temperature=1; top_p=0.95; max_output_tokens=300… 79.87
    Printed as 79.87%Official leaderboard, measured Sep 2026Configuration: model ID minimax/MiniMax-M2.7; temperature=1; top_p=0.95; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 56 of 105 (MiniMax-M2.7), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["minimax/MiniMax-M2.7"].
    56 | MiniMax-M2.7 | 79.87%±1.86 | $0.3/$1.2 | 27.25s
    Every result from this document
  57. 57Grok 4 Fast (Non-Reasoning) model ID grok/grok-4-fast-non-reasoning; temperature=1; top_p=0.95; max_output_… 79.72
    Printed as 79.72%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 57 of 105 (Grok 4 Fast (Non-Reasoning)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["grok/grok-4-fast-non-reasoning"].
    57 | Grok 4 Fast (Non-Reasoning) | 79.72%±1.87 | $0.2/$0.5 | 8.48s
    Every result from this document
  58. 58Gemini 3.6 Flash model ID google/gemini-3.6-flash; temperature=1; max_output_tokens=30000; reaso… 79.66
    Printed as 79.66%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-3.6-flash; temperature=1; max_output_tokens=30000; reasoning_effort=high
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 58 of 105 (Gemini 3.6 Flash), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-3.6-flash"].
    58 | Gemini 3.6 Flash | 79.66%±1.86 | $1.5/$7.5 | 33.55s
    Every result from this document
  59. 59Qwen 3.7 Max model ID alibaba/qwen3.7-max; temperature=1; max_output_tokens=30000 79.40
    Printed as 79.40%Official leaderboard, measured Sep 2026Configuration: model ID alibaba/qwen3.7-max; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 59 of 105 (Qwen 3.7 Max), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["alibaba/qwen3.7-max"].
    59 | Qwen 3.7 Max | 79.40%±1.91 | $2.5/$7.5 | 108.05s
    Every result from this document
  60. 60Grok 4.1 Fast (Reasoning) model ID grok/grok-4-1-fast-reasoning; temperature=1; top_p=0.95; max_output_to… 78.73
    Printed as 78.73%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4-1-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 60 of 105 (Grok 4.1 Fast (Reasoning)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["grok/grok-4-1-fast-reasoning"].
    60 | Grok 4.1 Fast (Reasoning) | 78.73%±1.87 | $0.2/$0.5 | 39.29s
    Every result from this document
  61. 61Gemini 2.5 Flash Preview (9/25) (Thinking) model ID google/gemini-2.5-flash-preview-09-2025-thinking; temperature=1; max_o… 78.50
    Printed as 78.50%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-2.5-flash-preview-09-2025-thinking; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 61 of 105 (Gemini 2.5 Flash Preview (9/25) (Thinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-2.5-flash-preview-09-2025-thinking"].
    61 | Gemini 2.5 Flash Preview (9/25) (Thinking) | 78.50%±1.99 | $0.3/$2.5 | 31.39s
    Every result from this document
  62. 62Grok 4 model ID grok/grok-4-0709; temperature=1; top_p=0.95; max_output_tokens=30000 78.15
    Printed as 78.15%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4-0709; temperature=1; top_p=0.95; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 62 of 105 (Grok 4), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["grok/grok-4-0709"].
    62 | Grok 4 | 78.15%±2.08 | $3/$15 | 74.05s
    Every result from this document
  63. 62Kimi K2.6 model ID kimi/kimi-k2.6; top_p=0.95; max_output_tokens=30000 78.15
    Printed as 78.15%Official leaderboard, measured Sep 2026Configuration: model ID kimi/kimi-k2.6; top_p=0.95; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 63 of 105 (Kimi K2.6), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["kimi/kimi-k2.6"].
    63 | Kimi K2.6 | 78.15%±1.79 | $0.95/$4 | 8m14s
    Every result from this document
  64. 64Gemini 2.5 Flash Preview (9/25) (Nonthinking) model ID google/gemini-2.5-flash-preview-09-2025; temperature=1; max_output_tok… 77.95
    Printed as 77.95%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-2.5-flash-preview-09-2025; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 64 of 105 (Gemini 2.5 Flash Preview (9/25) (Nonthinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-2.5-flash-preview-09-2025"].
    64 | Gemini 2.5 Flash Preview (9/25) (Nonthinking) | 77.95%±1.92 | $0.3/$2.5 | 22.89s
    Every result from this document
  65. 65GPT 5.4 (xhigh) model ID openai/gpt-5.4-2026-03-05; max_output_tokens=30000; reasoning_effort=x… 77.55
    Printed as 77.55%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5.4-2026-03-05; max_output_tokens=30000; reasoning_effort=xhigh
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 65 of 105 (GPT 5.4 (xhigh)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-5.4-2026-03-05"].
    65 | GPT 5.4 (xhigh) | 77.55%±3.32 | $2.5/$15 | 4m43s
    Every result from this document
  66. 66Grok 4.1 Fast Non-Reasoning model ID grok/grok-4-1-fast-non-reasoning; temperature=1; top_p=0.95; max_outpu… 77.46
    Printed as 77.46%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4-1-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 66 of 105 (Grok 4.1 Fast Non-Reasoning), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["grok/grok-4-1-fast-non-reasoning"].
    66 | Grok 4.1 Fast Non-Reasoning | 77.46%±2.04 | $0.2/$0.5 | 19.86s
    Every result from this document
  67. 67Qwen 3 VL Plus model ID alibaba/qwen3-vl-plus-2025-09-23; temperature=1; max_output_tokens=300… 77.13
    Printed as 77.13%Official leaderboard, measured Sep 2026Configuration: model ID alibaba/qwen3-vl-plus-2025-09-23; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 67 of 105 (Qwen 3 VL Plus), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["alibaba/qwen3-vl-plus-2025-09-23"].
    67 | Qwen 3 VL Plus | 77.13%±1.92 | $0.2/$1.6 | 71.35s
    Every result from this document
  68. 68GPT 5.4 Nano model ID openai/gpt-5.4-nano-2026-03-17; max_output_tokens=30000; reasoning_eff… 77.09
    Printed as 77.09%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5.4-nano-2026-03-17; max_output_tokens=30000; reasoning_effort=high
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 68 of 105 (GPT 5.4 Nano), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-5.4-nano-2026-03-17"].
    68 | GPT 5.4 Nano | 77.09%±1.89 | $0.2/$1.25 | 20.53s
    Every result from this document
  69. 69Qwen 3.6 Plus model ID alibaba/qwen3.6-plus; temperature=1; max_output_tokens=30000 76.96
    Printed as 76.96%Official leaderboard, measured Sep 2026Configuration: model ID alibaba/qwen3.6-plus; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 69 of 105 (Qwen 3.6 Plus), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["alibaba/qwen3.6-plus"].
    69 | Qwen 3.6 Plus | 76.96%±1.92 | $0.5/$3 | 2m54s
    Every result from this document
  70. 70o3 model ID openai/o3-2025-04-16; max_output_tokens=30000; reasoning_effort=high 76.65
    Printed as 76.65%Official leaderboard, measured Sep 2026Configuration: model ID openai/o3-2025-04-16; max_output_tokens=30000; reasoning_effort=high
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 70 of 105 (o3), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/o3-2025-04-16"].
    70 | o3 | 76.65%±1.87 | $2/$8 | 46.39s
    Every result from this document
  71. 71Gemini 3.5 Flash model ID google/gemini-3.5-flash; temperature=1; max_output_tokens=30000; reaso… 76.57
    Printed as 76.57%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-3.5-flash; temperature=1; max_output_tokens=30000; reasoning_effort=high
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 71 of 105 (Gemini 3.5 Flash), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-3.5-flash"].
    71 | Gemini 3.5 Flash | 76.57%±1.92 | $1.5/$9 | 57.08s
    Every result from this document
  72. 72Kimi K2.5 model ID kimi/kimi-k2.5-thinking; top_p=0.95; max_output_tokens=30000 76.44
    Printed as 76.44%Official leaderboard, measured Sep 2026Configuration: model ID kimi/kimi-k2.5-thinking; top_p=0.95; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 72 of 105 (Kimi K2.5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["kimi/kimi-k2.5-thinking"].
    72 | Kimi K2.5 | 76.44%±1.99 | $0.6/$3 | 2m27s
    Every result from this document
  73. 73Gemini 3.1 Pro Preview (02/26) model ID google/gemini-3.1-pro-preview; temperature=1; max_output_tokens=30000;… 76.11
    Printed as 76.11%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-3.1-pro-preview; temperature=1; max_output_tokens=30000; reasoning_effort=high
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 73 of 105 (Gemini 3.1 Pro Preview (02/26)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-3.1-pro-preview"].
    73 | Gemini 3.1 Pro Preview (02/26) | 76.11%±1.92 | $2/$12 | 69.14s
    Every result from this document
  74. 74Claude Sonnet 5 model ID anthropic/claude-sonnet-5; temperature=1; max_output_tokens=30000; com… 76.05
    Printed as 76.05%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-sonnet-5; temperature=1; max_output_tokens=30000; compute_effort=max
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 74 of 105 (Claude Sonnet 5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-sonnet-5"].
    74 | Claude Sonnet 5 | 76.05%±3.05 | $2/$10 | 4m12s
    Every result from this document
  75. 75Gemini 2.5 Flash Lite (9/25) (Nonthinking) model ID google/gemini-2.5-flash-lite-preview-09-2025; temperature=1; max_outpu… 75.82
    Printed as 75.82%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-2.5-flash-lite-preview-09-2025; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 75 of 105 (Gemini 2.5 Flash Lite (9/25) (Nonthinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-2.5-flash-lite-preview-09-2025"].
    75 | Gemini 2.5 Flash Lite (9/25) (Nonthinking) | 75.82%±1.85 | $0.1/$0.4 | 4.49s
    Every result from this document
  76. 76Ling 3.0 Flash Fin model ID ant/ling-3.0-flash-af-rc3; temperature=1; top_p=0.95; max_output_token… 75.59
    Printed as 75.59%Official leaderboard, measured Sep 2026Configuration: model ID ant/ling-3.0-flash-af-rc3; temperature=1; top_p=0.95; max_output_tokens=131072
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 76 of 105 (Ling 3.0 Flash Fin), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["ant/ling-3.0-flash-af-rc3"].
    76 | Ling 3.0 Flash Fin | 75.59%±2.03 | $0.06/$0.18 | 44.81s
    Every result from this document
  77. 77DeepSeek V4 model ID deepseek/deepseek-v4-pro; max_output_tokens=128000; reasoning_effort=m… 75.14
    Printed as 75.14%Official leaderboard, measured Sep 2026Configuration: model ID deepseek/deepseek-v4-pro; max_output_tokens=128000; reasoning_effort=max
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 77 of 105 (DeepSeek V4), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["deepseek/deepseek-v4-pro"].
    77 | DeepSeek V4 | 75.14%±2.00 | $1.32/$3.96 | 5m46s
    Every result from this document
  78. 78Grok 4.3 model ID grok/grok-4.3; temperature=1; top_p=0.95; max_output_tokens=30000 74.40
    Printed as 74.40%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4.3; temperature=1; top_p=0.95; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 78 of 105 (Grok 4.3), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["grok/grok-4.3"].
    78 | Grok 4.3 | 74.40%±2.02 | $1.25/$2.5 | 100.46s
    Every result from this document
  79. 79Claude Opus 4.1 (Thinking) model ID anthropic/claude-opus-4-1-20250805-thinking; temperature=1; max_output… 73.90
    Printed as 73.90%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-1-20250805-thinking; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 79 of 105 (Claude Opus 4.1 (Thinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-opus-4-1-20250805-thinking"].
    79 | Claude Opus 4.1 (Thinking) | 73.90%±1.97 | $15/$75 | 57.40s
    Every result from this document
  80. 80Gemini 2.5 Pro model ID google/gemini-2.5-pro; temperature=1; max_output_tokens=30000 73.55
    Printed as 73.55%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-2.5-pro; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 80 of 105 (Gemini 2.5 Pro), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-2.5-pro"].
    80 | Gemini 2.5 Pro | 73.55%±1.91 | $1.25/$10 | 35.91s
    Every result from this document
  81. 81GPT 5 Nano model ID openai/gpt-5-nano-2025-08-07; max_output_tokens=30000; reasoning_effor… 72.86
    Printed as 72.86%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5-nano-2025-08-07; max_output_tokens=30000; reasoning_effort=high
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 81 of 105 (GPT 5 Nano), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-5-nano-2025-08-07"].
    81 | GPT 5 Nano | 72.86%±1.89 | $0.05/$0.4 | 112.91s
    Every result from this document
  82. 82Gemini 2.5 Flash Lite (Nonthinking) model ID google/gemini-2.5-flash-lite; temperature=1; max_output_tokens=30000 72.83
    Printed as 72.83%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-2.5-flash-lite; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 82 of 105 (Gemini 2.5 Flash Lite (Nonthinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-2.5-flash-lite"].
    82 | Gemini 2.5 Flash Lite (Nonthinking) | 72.83%±1.98 | $0.1/$0.4 | 4.85s
    Every result from this document
  83. 83Qwen 3 Max Thinking model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=30000 72.71
    Printed as 72.71%Official leaderboard, measured Sep 2026Configuration: model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 83 of 105 (Qwen 3 Max Thinking), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["alibaba/qwen3-max-2026-01-23"].
    83 | Qwen 3 Max Thinking | 72.71%±1.91 | $1.2/$6 | 6m02s
    Every result from this document
  84. 84Claude Sonnet 4 (Nonthinking) model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=3… 72.41
    Printed as 72.41%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 84 of 105 (Claude Sonnet 4 (Nonthinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-sonnet-4-20250514"].
    84 | Claude Sonnet 4 (Nonthinking) | 72.41%±1.93 | $3/$15 | 25.67s
    Every result from this document
  85. 85GLM 5.1 model ID zai/glm-5.1; temperature=1; top_p=0.95; max_output_tokens=30000 72.27
    Printed as 72.27%Official leaderboard, measured Sep 2026Configuration: model ID zai/glm-5.1; temperature=1; top_p=0.95; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 85 of 105 (GLM 5.1), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["zai/glm-5.1"].
    85 | GLM 5.1 | 72.27%±2.06 | $1/$3.2 | 95.70s
    Every result from this document
  86. 86MiMo V2.5 model ID xiaomi/mimo-v2.5; temperature=1; top_p=0.95; max_output_tokens=30000 72.15
    Printed as 72.15%Official leaderboard, measured Sep 2026Configuration: model ID xiaomi/mimo-v2.5; temperature=1; top_p=0.95; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 86 of 105 (MiMo V2.5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["xiaomi/mimo-v2.5"].
    86 | MiMo V2.5 | 72.15%±1.85 | $0.14/$0.28 | 20.26s
    Every result from this document
  87. 87Gemini 3 Pro (11/25) model ID google/gemini-3-pro-preview; temperature=1; max_output_tokens=30000; r… 72.04
    Printed as 72.04%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-3-pro-preview; temperature=1; max_output_tokens=30000; reasoning_effort=high
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 87 of 105 (Gemini 3 Pro (11/25)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-3-pro-preview"].
    87 | Gemini 3 Pro (11/25) | 72.04%±1.90 | $2/$12 | 43.39s
    Every result from this document
  88. 88Claude Opus 4.1 (Nonthinking) model ID anthropic/claude-opus-4-1-20250805; temperature=1; max_output_tokens=3… 71.75
    Printed as 71.75%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-1-20250805; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 88 of 105 (Claude Opus 4.1 (Nonthinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-opus-4-1-20250805"].
    88 | Claude Opus 4.1 (Nonthinking) | 71.75%±2.02 | $15/$75 | 38.04s
    Every result from this document
  89. 89Gemini 3.5 Flash Lite model ID google/gemini-3.5-flash-lite; temperature=1; max_output_tokens=30000;… 70.89
    Printed as 70.89%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-3.5-flash-lite; temperature=1; max_output_tokens=30000; reasoning_effort=high
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 89 of 105 (Gemini 3.5 Flash Lite), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-3.5-flash-lite"].
    89 | Gemini 3.5 Flash Lite | 70.89%±2.03 | $0.3/$2.5 | 18.99s
    Every result from this document
  90. 90Qwen 3.5 Flash model ID alibaba/qwen3.5-flash; temperature=1; max_output_tokens=30000 70.62
    Printed as 70.62%Official leaderboard, measured Sep 2026Configuration: model ID alibaba/qwen3.5-flash; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 90 of 105 (Qwen 3.5 Flash), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["alibaba/qwen3.5-flash"].
    90 | Qwen 3.5 Flash | 70.62%±2.09 | $0.1/$0.4 | 80.41s
    Every result from this document
  91. 91Gemini 3 Flash (12/25) model ID google/gemini-3-flash-preview; temperature=1; max_output_tokens=30000;… 69.92
    Printed as 69.92%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-3-flash-preview; temperature=1; max_output_tokens=30000; reasoning_effort=high
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 91 of 105 (Gemini 3 Flash (12/25)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-3-flash-preview"].
    91 | Gemini 3 Flash (12/25) | 69.92%±1.90 | $0.5/$3 | 23.49s
    Every result from this document
  92. 92Claude Sonnet 4 (Thinking) model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=30000 69.35
    Printed as 69.35%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 92 of 105 (Claude Sonnet 4 (Thinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-sonnet-4-20250514-thinking"].
    92 | Claude Sonnet 4 (Thinking) | 69.35%±2.21 | $3/$15 | 39.57s
    Every result from this document
  93. 93o4 Mini model ID openai/o4-mini-2025-04-16; max_output_tokens=30000; reasoning_effort=h… 69.14
    Printed as 69.14%Official leaderboard, measured Sep 2026Configuration: model ID openai/o4-mini-2025-04-16; max_output_tokens=30000; reasoning_effort=high
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 93 of 105 (o4 Mini), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/o4-mini-2025-04-16"].
    93 | o4 Mini | 69.14%±1.96 | $1.1/$4.4 | 81.96s
    Every result from this document
  94. 94GLM 4.7 model ID zai/glm-4.7; temperature=1; top_p=1; max_output_tokens=30000 68.63
    Printed as 68.63%Official leaderboard, measured Sep 2026Configuration: model ID zai/glm-4.7; temperature=1; top_p=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 94 of 105 (GLM 4.7), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["zai/glm-4.7"].
    94 | GLM 4.7 | 68.63%±2.12 | $0.6/$2.2 | 2m47s
    Every result from this document
  95. 95Mistral Medium 3.5 model ID mistralai/mistral-medium-3.5; temperature=1; top_p=0.95; max_output_to… 67.73
    Printed as 67.73%Official leaderboard, measured Sep 2026Configuration: model ID mistralai/mistral-medium-3.5; temperature=1; top_p=0.95; max_output_tokens=30000; reasoning_effort=high
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 95 of 105 (Mistral Medium 3.5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["mistralai/mistral-medium-3.5"].
    95 | Mistral Medium 3.5 | 67.73%±2.01 | $1.5/$7.5 | 104.05s
    Every result from this document
  96. 96Gemini 2.5 Flash Lite (9/25) (Thinking) model ID google/gemini-2.5-flash-lite-preview-09-2025-thinking; temperature=1;… 66.88
    Printed as 66.88%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-2.5-flash-lite-preview-09-2025-thinking; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 96 of 105 (Gemini 2.5 Flash Lite (9/25) (Thinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-2.5-flash-lite-preview-09-2025-thinking"].
    96 | Gemini 2.5 Flash Lite (9/25) (Thinking) | 66.88%±1.92 | $0.1/$0.4 | 11.52s
    Every result from this document
  97. 97Laguna M.1 model ID poolside/laguna-m.1; temperature=1; max_output_tokens=30000 65.91
    Printed as 65.91%Official leaderboard, measured Sep 2026Configuration: model ID poolside/laguna-m.1; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 97 of 105 (Laguna M.1), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["poolside/laguna-m.1"].
    97 | Laguna M.1 | 65.91%±2.01 | N/A | 111.82s
    Every result from this document
  98. 98Gemini 3.1 Flash Lite Preview model ID google/gemini-3.1-flash-lite-preview; temperature=1; max_output_tokens… 63.90
    Printed as 63.90%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-3.1-flash-lite-preview; temperature=1; max_output_tokens=30000; reasoning_effort=high
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 98 of 105 (Gemini 3.1 Flash Lite Preview), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-3.1-flash-lite-preview"].
    98 | Gemini 3.1 Flash Lite Preview | 63.90%±1.82 | $0.25/$1.5 | 16.79s
    Every result from this document
  99. 99Grok 4.20 (Reasoning) model ID grok/grok-4.20-0309-reasoning; temperature=1; top_p=0.95; max_output_t… 63.41
    Printed as 63.41%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4.20-0309-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 99 of 105 (Grok 4.20 (Reasoning)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["grok/grok-4.20-0309-reasoning"].
    99 | Grok 4.20 (Reasoning) | 63.41%±2.10 | $2/$6 | 18.60s
    Every result from this document
  100. 100Laguna XS.2 model ID poolside/laguna-xs.2; temperature=1; max_output_tokens=30000 61.43
    Printed as 61.43%Official leaderboard, measured Sep 2026Configuration: model ID poolside/laguna-xs.2; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 100 of 105 (Laguna XS.2), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["poolside/laguna-xs.2"].
    100 | Laguna XS.2 | 61.43%±2.35 | N/A | 68.82s
    Every result from this document
  101. 101Command A+ model ID cohere/command-a-plus-05-2026; temperature=1; top_p=0.95; max_output_t… 55.68
    Printed as 55.68%Official leaderboard, measured Sep 2026Configuration: model ID cohere/command-a-plus-05-2026; temperature=1; top_p=0.95; max_output_tokens=64000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 101 of 105 (Command A+), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["cohere/command-a-plus-05-2026"].
    101 | Command A+ | 55.68%±3.65 | N/A | 3m23s
    Every result from this document
  102. 102Mercury 2.5 model ID inception/mercury-2.5; temperature=1; max_output_tokens=65536; reasoni… 55.09
    Printed as 55.09%Official leaderboard, measured Sep 2026Configuration: model ID inception/mercury-2.5; temperature=1; max_output_tokens=65536; reasoning_effort=high
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 102 of 105 (Mercury 2.5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["inception/mercury-2.5"].
    102 | Mercury 2.5 | 55.09%±2.10 | $0.2/$0.75 | 9.99s
    Every result from this document
  103. 103Llama 4 Maverick model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_to… 54.22
    Printed as 54.22%Official leaderboard, measured Sep 2026Configuration: model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 103 of 105 (Llama 4 Maverick), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["fireworks/llama4-maverick-instruct-basic"].
    103 | Llama 4 Maverick | 54.22%±1.87 | $0.22/$0.88 | 25.05s
    Every result from this document
  104. 104Llama 4 Scout model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max… 50.59
    Printed as 50.59%Official leaderboard, measured Sep 2026Configuration: model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 104 of 105 (Llama 4 Scout), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["together/meta-llama/Llama-4-Scout-17B-16E-Instruct"].
    104 | Llama 4 Scout | 50.59%±1.90 | $0.18/$0.59 | 11.32s
    Every result from this document
  105. 105Nemotron 3.5 Lightning model ID fireworks/nemotron-lightning-3p5-30b-a3b; temperature=1; top_p=0.95; m… 4.27
    Printed as 4.27%Official leaderboard, measured Sep 2026Configuration: model ID fireworks/nemotron-lightning-3p5-30b-a3b; temperature=1; top_p=0.95; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 105 of 105 (Nemotron 3.5 Lightning), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["fireworks/nemotron-lightning-3p5-30b-a3b"].
    105 | Nemotron 3.5 Lightning | 4.27%±0.49 | $0.05/$0.2 | 105.52s
    Every result from this document

Documents

1