MedScribe (Vals AI)
Quality and compliance of SOAP notes generated from clinical visits, scored against documentation rubrics.
Vals AIOfficial page
More about this boardLess
Clinical documentation support: quality of SOAP notes generated from clinical visits, scored against rubrics for documentation quality and compliance.
Vals AI self-runs; 105 models on the board, last updated September 29, 2026. Top scores cluster near 90, so the leaders sit close to the ceiling.
Published by Vals AI (dataset with Protege), released Feb 2026. 100 rubric-scored SOAP-note cases. Percentage accuracy 0-100, higher better.
- Rows on the board
- 105
- Board last updated
- 29 Sep 2026
- Rows
- 105
- Models
- 94 of 105on the board
- Labs
- 19
- Last measured
- Sep 2026
- Leader
- 91.43Claude Opus 5.5
Ranking
Percentage accuracy 0-100, higher better
- Anthropic
- OpenAI
- Meta
- Moonshot AI
- Other labs
All official leaderboard
20 of 105 rows
Rows and sources
Open a row for the quote, the page and the document.
1Claude Opus 5.5 model ID anthropic/claude-opus-5-5; temperature=1; max_output_tokens=128000; co… 91.43
Printed as 91.43%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-5-5; temperature=1; max_output_tokens=128000; compute_effort=maxVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 1 of 105 (Claude Opus 5.5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-opus-5-5"].1 | Claude Opus 5.5 | 91.43%±1.93 | $4/$20 | 7m22s
Every result from this document2Claude Fable 5.1 91.29
Printed as 91.29%Official leaderboard, measured Sep 2026Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 2 of 105 (Claude Fable 5.1), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-fable-5-1"].2 | Claude Fable 5.1 | 91.29%±1.95 | $10/$50 | 3m08s
Every result from this document3Claude Sonnet 5.5 model ID anthropic/claude-sonnet-5-5; temperature=1; max_output_tokens=128000;… 91.10
Printed as 91.10%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-sonnet-5-5; temperature=1; max_output_tokens=128000; compute_effort=maxVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 3 of 105 (Claude Sonnet 5.5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-sonnet-5-5"].3 | Claude Sonnet 5.5 | 91.10%±1.96 | $2/$10 | 5m07s
Every result from this document4Claude Opus 5 90.98
Printed as 90.98%Official leaderboard, measured Sep 2026Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 4 of 105 (Claude Opus 5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-opus-5"].4 | Claude Opus 5 | 90.98%±1.92 | $5/$25 | 76.56s
Every result from this document5Muse Spark 1.2 90.06
Printed as 90.06%Official leaderboard, measured Sep 2026Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 5 of 105 (Muse Spark 1.2), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["meta/muse_spark_1_2"].5 | Muse Spark 1.2 | 90.06%±1.96 | $1.25/$4.25 | 61.46s
Every result from this document6Grok 4.7 model ID grok/grok-4.7; temperature=1; top_p=0.95; reasoning_effort=xhigh 89.38
Printed as 89.38%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4.7; temperature=1; top_p=0.95; reasoning_effort=xhighVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 6 of 105 (Grok 4.7), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["grok/grok-4.7"].6 | Grok 4.7 | 89.38%±1.89 | $2/$6 | 2m33s
Every result from this document7GLM 5.3 Flash model ID zai/glm-5.3-flash; temperature=1; top_p=0.95; max_output_tokens=30000;… 88.94
Printed as 88.94%Official leaderboard, measured Sep 2026Configuration: model ID zai/glm-5.3-flash; temperature=1; top_p=0.95; max_output_tokens=30000; reasoning_effort=maxVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 7 of 105 (GLM 5.3 Flash), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["zai/glm-5.3-flash"].7 | GLM 5.3 Flash | 88.94%±1.91 | $0.075/$0.25 | 86.98s
Every result from this document8Muse Spark 1.1 88.89
Printed as 88.89%Official leaderboard, measured Sep 2026Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 8 of 105 (Muse Spark 1.1), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["meta/muse_spark_1_1"].8 | Muse Spark 1.1 | 88.89%±1.95 | $1.25/$4.25 | 63.34s
Every result from this document9GLM 5.3 model ID zai/glm-5.3; temperature=1; top_p=0.95; max_output_tokens=30000; reaso… 88.81
Printed as 88.81%Official leaderboard, measured Sep 2026Configuration: model ID zai/glm-5.3; temperature=1; top_p=0.95; max_output_tokens=30000; reasoning_effort=maxVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 9 of 105 (GLM 5.3), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["zai/glm-5.3"].9 | GLM 5.3 | 88.81%±2.00 | $1.4/$4.4 | 2m02s
Every result from this document10Claude Fable 5 88.52
Printed as 88.52%Official leaderboard, measured Sep 2026Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 10 of 105 (Claude Fable 5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-fable-5"].10 | Claude Fable 5 | 88.52%±1.95 | $10/$50 | 119.47s
Every result from this document11MiMo V2.6 Pro model ID xiaomi/mimo-v2.6-pro; temperature=1; top_p=0.95; max_output_tokens=128… 88.31
Printed as 88.31%Official leaderboard, measured Sep 2026Configuration: model ID xiaomi/mimo-v2.6-pro; temperature=1; top_p=0.95; max_output_tokens=128000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 11 of 105 (MiMo V2.6 Pro), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["xiaomi/mimo-v2.6-pro"].11 | MiMo V2.6 Pro | 88.31%±1.94 | $0.435/$0.87 | 2m53s
Every result from this document12GPT 5.1 88.09
Printed as 88.09%Official leaderboard, measured Sep 2026Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 12 of 105 (GPT 5.1), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-5.1-2025-11-13"].12 | GPT 5.1 | 88.09%±1.94 | $1.25/$10 | 77.98s
Every result from this document13Kimi K3 87.96
Printed as 87.96%Official leaderboard, measured Sep 2026Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 13 of 105 (Kimi K3), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["kimi/kimi-k3"].13 | Kimi K3 | 87.96%±1.89 | $3/$15 | 2m17s
Every result from this document14GPT-6 Astra 87.91
Printed as 87.91%Official leaderboard, measured Sep 2026Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 14 of 105 (GPT-6 Astra), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-6-astra"].14 | GPT-6 Astra | 87.91%±1.94 | $10/$50 | 2m34s
Every result from this document15MiniMax-M3 87.25
Printed as 87.25%Official leaderboard, measured Sep 2026Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 15 of 105 (MiniMax-M3), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["minimax/MiniMax-M3"].15 | MiniMax-M3 | 87.25%±1.96 | $0.6/$2.4 | 2m04s
Every result from this document16Grok 4.5 model ID grok/grok-4.5; temperature=1; top_p=0.95; max_output_tokens=30000; rea… 86.88
Printed as 86.88%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4.5; temperature=1; top_p=0.95; max_output_tokens=30000; reasoning_effort=highVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 16 of 105 (Grok 4.5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["grok/grok-4.5"].16 | Grok 4.5 | 86.88%±1.94 | $2/$6 | 11m45s
Every result from this document17GPT 5.5 86.87
Printed as 86.87%Official leaderboard, measured Sep 2026Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 17 of 105 (GPT 5.5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-5.5"].17 | GPT 5.5 | 86.87%±1.93 | $5/$30 | 2m13s
Every result from this document18Claude Opus 4.6 (Nonthinking) 86.74
Printed as 86.74%Official leaderboard, measured Sep 2026Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 18 of 105 (Claude Opus 4.6 (Nonthinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-opus-4-6"].18 | Claude Opus 4.6 (Nonthinking) | 86.74%±1.94 | $5/$25 | 54.32s
Every result from this document19Grok 4.6 86.53
Printed as 86.53%Official leaderboard, measured Sep 2026Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 19 of 105 (Grok 4.6), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["grok/grok-4.6"].19 | Grok 4.6 | 86.53%±1.96 | $2/$6 | 70.94s
Every result from this document20GPT-6.1 Sol model ID openai/gpt-6.1-sol; max_output_tokens=128000; reasoning_effort=max 86.45
Printed as 86.45%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-6.1-sol; max_output_tokens=128000; reasoning_effort=maxVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 20 of 105 (GPT-6.1 Sol), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-6.1-sol"].20 | GPT-6.1 Sol | 86.45%±1.90 | $2/$10 | 2m50s
Every result from this document21Claude Opus 4.6 (Thinking) model ID anthropic/claude-opus-4-6-thinking; temperature=1; max_output_tokens=3… 86.13
Printed as 86.13%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-6-thinking; temperature=1; max_output_tokens=30000; compute_effort=maxVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 21 of 105 (Claude Opus 4.6 (Thinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-opus-4-6-thinking"].21 | Claude Opus 4.6 (Thinking) | 86.13%±1.94 | $5/$25 | 2m08s
Every result from this document22Muse Spark model ID meta/muse_spark; temperature=1; max_output_tokens=30000 85.90
Printed as 85.90%Official leaderboard, measured Sep 2026Configuration: model ID meta/muse_spark; temperature=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 22 of 105 (Muse Spark), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["meta/muse_spark"].22 | Muse Spark | 85.90%±1.85 | N/A | 3m11s
Every result from this document23Claude Opus 4.8 model ID anthropic/claude-opus-4-8; temperature=1; max_output_tokens=30000; com… 85.75
Printed as 85.75%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-8; temperature=1; max_output_tokens=30000; compute_effort=maxVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 23 of 105 (Claude Opus 4.8), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-opus-4-8"].23 | Claude Opus 4.8 | 85.75%±1.93 | $5/$25 | 82.36s
Every result from this document24DeepSeek V4.1 Flash model ID deepseek/deepseek-v4.1-flash; max_output_tokens=384000; reasoning_effo… 85.50
Printed as 85.50%Official leaderboard, measured Sep 2026Configuration: model ID deepseek/deepseek-v4.1-flash; max_output_tokens=384000; reasoning_effort=highVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 24 of 105 (DeepSeek V4.1 Flash), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["deepseek/deepseek-v4.1-flash"].24 | DeepSeek V4.1 Flash | 85.50%±1.92 | $0.3/$1.2 | 53.83s
Every result from this document25Inkling model ID thinkingmachines/inkling; temperature=1; top_p=1; max_output_tokens=30… 85.41
Printed as 85.41%Official leaderboard, measured Sep 2026Configuration: model ID thinkingmachines/inkling; temperature=1; top_p=1; max_output_tokens=30000; reasoning_effort=0.99Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 25 of 105 (Inkling), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["thinkingmachines/inkling"].25 | Inkling | 85.41%±1.84 | $1/$4.05 | 4m05s
Every result from this document26Claude Opus 4.5 (Thinking) model ID anthropic/claude-opus-4-5-20251101-thinking; temperature=1; max_output… 85.32
Printed as 85.32%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-5-20251101-thinking; temperature=1; max_output_tokens=30000; compute_effort=highVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 26 of 105 (Claude Opus 4.5 (Thinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-opus-4-5-20251101-thinking"].26 | Claude Opus 4.5 (Thinking) | 85.32%±1.90 | $5/$25 | 72.11s
Every result from this document27MiMo V2.6 Flash model ID xiaomi/mimo-v2.6-flash; temperature=1; top_p=0.95; max_output_tokens=1… 85.28
Printed as 85.28%Official leaderboard, measured Sep 2026Configuration: model ID xiaomi/mimo-v2.6-flash; temperature=1; top_p=0.95; max_output_tokens=128000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 27 of 105 (MiMo V2.6 Flash), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["xiaomi/mimo-v2.6-flash"].27 | MiMo V2.6 Flash | 85.28%±1.99 | $0.14/$0.28 | 74.48s
Every result from this document28Claude Haiku 4.5 (Thinking) model ID anthropic/claude-haiku-4-5-20251001-thinking; temperature=1; max_outpu… 85.23
Printed as 85.23%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-haiku-4-5-20251001-thinking; temperature=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 29 of 105 (Claude Haiku 4.5 (Thinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-haiku-4-5-20251001-thinking"].29 | Claude Haiku 4.5 (Thinking) | 85.23%±1.90 | $1/$5 | 66.20s
Every result from this document28GPT-5.6 Sol model ID openai/gpt-5.6-sol; max_output_tokens=30000; reasoning_effort=max 85.23
Printed as 85.23%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5.6-sol; max_output_tokens=30000; reasoning_effort=maxVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 28 of 105 (GPT-5.6 Sol), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-5.6-sol"].28 | GPT-5.6 Sol | 85.23%±1.97 | $4/$20 | 94.40s
Every result from this document30Qwen 3.8 Max model ID alibaba/qwen3.8-max; temperature=1; max_output_tokens=30000 84.95
Printed as 84.95%Official leaderboard, measured Sep 2026Configuration: model ID alibaba/qwen3.8-max; temperature=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 30 of 105 (Qwen 3.8 Max), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["alibaba/qwen3.8-max"].30 | Qwen 3.8 Max | 84.95%±2.00 | $2/$6 | 4m11s
Every result from this document31Claude Sonnet 4.5 (Nonthinking) model ID anthropic/claude-sonnet-4-5-20250929; temperature=1; max_output_tokens… 84.52
Printed as 84.52%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-sonnet-4-5-20250929; temperature=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 31 of 105 (Claude Sonnet 4.5 (Nonthinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-sonnet-4-5-20250929"].31 | Claude Sonnet 4.5 (Nonthinking) | 84.52%±1.93 | $3/$15 | 44.42s
Every result from this document32Gemini 3.8 Flash model ID google/gemini-3.8-flash; temperature=1; max_output_tokens=65536; reaso… 84.50
Printed as 84.50%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-3.8-flash; temperature=1; max_output_tokens=65536; reasoning_effort=highVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 32 of 105 (Gemini 3.8 Flash), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-3.8-flash"].32 | Gemini 3.8 Flash | 84.50%±1.94 | $1.5/$7.5 | 21.14s
Every result from this document33GPT 5.2 model ID openai/gpt-5.2-2025-12-11; max_output_tokens=30000; reasoning_effort=x… 84.39
Printed as 84.39%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5.2-2025-12-11; max_output_tokens=30000; reasoning_effort=xhighVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 34 of 105 (GPT 5.2), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-5.2-2025-12-11"].34 | GPT 5.2 | 84.39%±1.86 | $1.75/$14 | 2m06s
Every result from this document33GPT-5.6 Luna model ID openai/gpt-5.6-luna; max_output_tokens=30000; reasoning_effort=max 84.39
Printed as 84.39%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5.6-luna; max_output_tokens=30000; reasoning_effort=maxVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 33 of 105 (GPT-5.6 Luna), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-5.6-luna"].33 | GPT-5.6 Luna | 84.39%±2.58 | $0.2/$1.2 | 116.89s
Every result from this document35Inkling Small model ID thinkingmachines/inkling-small; temperature=1; top_p=1; max_output_tok… 84.11
Printed as 84.11%Official leaderboard, measured Sep 2026Configuration: model ID thinkingmachines/inkling-small; temperature=1; top_p=1; max_output_tokens=30000; reasoning_effort=0.99Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 35 of 105 (Inkling Small), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["thinkingmachines/inkling-small"].35 | Inkling Small | 84.11%±1.87 | $0.3/$1.2 | 3m50s
Every result from this document36Claude Sonnet 4.5 (Thinking) model ID anthropic/claude-sonnet-4-5-20250929-thinking; temperature=1; max_outp… 84.10
Printed as 84.10%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-sonnet-4-5-20250929-thinking; temperature=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 36 of 105 (Claude Sonnet 4.5 (Thinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-sonnet-4-5-20250929-thinking"].36 | Claude Sonnet 4.5 (Thinking) | 84.10%±1.87 | $3/$15 | 67.92s
Every result from this document37Gemini 3.7 Flash model ID google/gemini-3.7-flash; temperature=1; max_output_tokens=30000; reaso… 83.94
Printed as 83.94%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-3.7-flash; temperature=1; max_output_tokens=30000; reasoning_effort=highVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 37 of 105 (Gemini 3.7 Flash), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-3.7-flash"].37 | Gemini 3.7 Flash | 83.94%±2.00 | $1.5/$7.5 | 16.82s
Every result from this document38Qwen 3.8 27B model ID alibaba/qwen3.8-27b; temperature=1; top_p=0.95; max_output_tokens=3000… 83.85
Printed as 83.85%Official leaderboard, measured Sep 2026Configuration: model ID alibaba/qwen3.8-27b; temperature=1; top_p=0.95; max_output_tokens=30000; reasoning_effort=xhighVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 38 of 105 (Qwen 3.8 27B), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["alibaba/qwen3.8-27b"].38 | Qwen 3.8 27B | 83.85%±1.98 | $0.5/$3 | 89.70s
Every result from this document39MiMo V2.5 Pro model ID xiaomi/mimo-v2.5-pro; temperature=1; top_p=0.95; max_output_tokens=300… 83.73
Printed as 83.73%Official leaderboard, measured Sep 2026Configuration: model ID xiaomi/mimo-v2.5-pro; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 39 of 105 (MiMo V2.5 Pro), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["xiaomi/mimo-v2.5-pro"].39 | MiMo V2.5 Pro | 83.73%±2.06 | $0.435/$0.87 | 90.71s
Every result from this document40GPT-6 Luna model ID openai/gpt-6-luna; max_output_tokens=128000; reasoning_effort=max 83.71
Printed as 83.71%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-6-luna; max_output_tokens=128000; reasoning_effort=maxVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 40 of 105 (GPT-6 Luna), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-6-luna"].40 | GPT-6 Luna | 83.71%±1.95 | $0.1/$0.5 | 2m50s
Every result from this document41GPT 5 model ID openai/gpt-5-2025-08-07; max_output_tokens=30000; reasoning_effort=high 83.65
Printed as 83.65%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5-2025-08-07; max_output_tokens=30000; reasoning_effort=highVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 41 of 105 (GPT 5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-5-2025-08-07"].41 | GPT 5 | 83.65%±1.94 | $1.25/$10 | 3m05s
Every result from this document42Hy4 Preview model ID tencent/hy4-preview; temperature=1; top_p=1; max_output_tokens=64000 83.60
Printed as 83.60%Official leaderboard, measured Sep 2026Configuration: model ID tencent/hy4-preview; temperature=1; top_p=1; max_output_tokens=64000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 42 of 105 (Hy4 Preview), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["tencent/hy4-preview"].42 | Hy4 Preview | 83.60%±2.06 | $0.834/$2.501 | 5m50s
Every result from this document43GLM 5.2 model ID zai/glm-5.2; temperature=1; max_output_tokens=30000 83.53
Printed as 83.53%Official leaderboard, measured Sep 2026Configuration: model ID zai/glm-5.2; temperature=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 43 of 105 (GLM 5.2), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["zai/glm-5.2"].43 | GLM 5.2 | 83.53%±2.00 | $1.4/$4.4 | 2m18s
Every result from this document44Claude Opus 4.5 (Nonthinking) model ID anthropic/claude-opus-4-5-20251101; temperature=1; max_output_tokens=3… 83.25
Printed as 83.25%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-5-20251101; temperature=1; max_output_tokens=30000; compute_effort=highVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 44 of 105 (Claude Opus 4.5 (Nonthinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-opus-4-5-20251101"].44 | Claude Opus 4.5 (Nonthinking) | 83.25%±1.93 | $5/$25 | 43.30s
Every result from this document45Gemini 2.5 Flash (7/17) (Thinking) model ID google/gemini-2.5-flash-thinking; temperature=1; max_output_tokens=300… 82.98
Printed as 82.98%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-2.5-flash-thinking; temperature=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 45 of 105 (Gemini 2.5 Flash (7/17) (Thinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-2.5-flash-thinking"].45 | Gemini 2.5 Flash (7/17) (Thinking) | 82.98%±1.91 | $0.3/$2.5 | 22.53s
Every result from this document46Claude Opus 4.7 model ID anthropic/claude-opus-4-7; temperature=1; max_output_tokens=30000; com… 82.95
Printed as 82.95%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-7; temperature=1; max_output_tokens=30000; compute_effort=maxVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 46 of 105 (Claude Opus 4.7), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-opus-4-7"].46 | Claude Opus 4.7 | 82.95%±1.98 | $5/$25 | 67.50s
Every result from this document47Gemini 2.5 Flash (7/17) (Nonthinking) model ID google/gemini-2.5-flash; temperature=1; max_output_tokens=30000 82.87
Printed as 82.87%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-2.5-flash; temperature=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 47 of 105 (Gemini 2.5 Flash (7/17) (Nonthinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-2.5-flash"].47 | Gemini 2.5 Flash (7/17) (Nonthinking) | 82.87%±1.91 | $0.3/$2.5 | 22.79s
Every result from this document47GPT-5.6 Terra model ID openai/gpt-5.6-terra; max_output_tokens=30000; reasoning_effort=xhigh 82.87
Printed as 82.87%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5.6-terra; max_output_tokens=30000; reasoning_effort=xhighVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 48 of 105 (GPT-5.6 Terra), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-5.6-terra"].48 | GPT-5.6 Terra | 82.87%±1.95 | $2/$12 | 35.53s
Every result from this document49GPT-6 Sol model ID openai/gpt-6-sol; max_output_tokens=128000; reasoning_effort=max 82.03
Printed as 82.03%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-6-sol; max_output_tokens=128000; reasoning_effort=maxVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 49 of 105 (GPT-6 Sol), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-6-sol"].49 | GPT-6 Sol | 82.03%±1.94 | $2/$10 | 81.14s
Every result from this document50Grok 4 Fast (Reasoning) model ID grok/grok-4-fast-reasoning; temperature=1; top_p=0.95; max_output_toke… 81.63
Printed as 81.63%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 50 of 105 (Grok 4 Fast (Reasoning)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["grok/grok-4-fast-reasoning"].50 | Grok 4 Fast (Reasoning) | 81.63%±2.14 | $0.2/$0.5 | 12.42s
Every result from this document51Ling 3.0 Flash model ID ant/ling-3.0-flash-2607; temperature=1; top_p=0.95; max_output_tokens=… 80.90
Printed as 80.90%Official leaderboard, measured Sep 2026Configuration: model ID ant/ling-3.0-flash-2607; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 51 of 105 (Ling 3.0 Flash), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["ant/ling-3.0-flash-2607"].51 | Ling 3.0 Flash | 80.90%±2.04 | $0.075/$0.22 | 12.12s
Every result from this document52MiniMax-M2.1 model ID minimax/MiniMax-M2.1; temperature=1; top_p=0.95; max_output_tokens=300… 80.78
Printed as 80.78%Official leaderboard, measured Sep 2026Configuration: model ID minimax/MiniMax-M2.1; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 52 of 105 (MiniMax-M2.1), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["minimax/MiniMax-M2.1"].52 | MiniMax-M2.1 | 80.78%±1.83 | $0.3/$1.2 | 53.16s
Every result from this document53GPT 5 Mini model ID openai/gpt-5-mini-2025-08-07; max_output_tokens=30000; reasoning_effor… 80.58
Printed as 80.58%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5-mini-2025-08-07; max_output_tokens=30000; reasoning_effort=highVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 53 of 105 (GPT 5 Mini), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-5-mini-2025-08-07"].53 | GPT 5 Mini | 80.58%±1.92 | $0.25/$2 | 4m36s
Every result from this document54DeepSeek V4 Flash 0731 model ID deepseek/deepseek-v4-flash-0731; max_output_tokens=30000; reasoning_ef… 80.36
Printed as 80.36%Official leaderboard, measured Sep 2026Configuration: model ID deepseek/deepseek-v4-flash-0731; max_output_tokens=30000; reasoning_effort=highVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 54 of 105 (DeepSeek V4 Flash 0731), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["deepseek/deepseek-v4-flash-0731"].54 | DeepSeek V4 Flash 0731 | 80.36%±1.97 | $0.44/$1.32 | 77.55s
Every result from this document55DeepSeek V4 Pro 0813 model ID deepseek/deepseek-v4-pro-0813; max_output_tokens=30000; reasoning_effo… 80.17
Printed as 80.17%Official leaderboard, measured Sep 2026Configuration: model ID deepseek/deepseek-v4-pro-0813; max_output_tokens=30000; reasoning_effort=maxVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 55 of 105 (DeepSeek V4 Pro 0813), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["deepseek/deepseek-v4-pro-0813"].55 | DeepSeek V4 Pro 0813 | 80.17%±2.00 | $1.32/$3.96 | 2m36s
Every result from this document56MiniMax-M2.7 model ID minimax/MiniMax-M2.7; temperature=1; top_p=0.95; max_output_tokens=300… 79.87
Printed as 79.87%Official leaderboard, measured Sep 2026Configuration: model ID minimax/MiniMax-M2.7; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 56 of 105 (MiniMax-M2.7), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["minimax/MiniMax-M2.7"].56 | MiniMax-M2.7 | 79.87%±1.86 | $0.3/$1.2 | 27.25s
Every result from this document57Grok 4 Fast (Non-Reasoning) model ID grok/grok-4-fast-non-reasoning; temperature=1; top_p=0.95; max_output_… 79.72
Printed as 79.72%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 57 of 105 (Grok 4 Fast (Non-Reasoning)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["grok/grok-4-fast-non-reasoning"].57 | Grok 4 Fast (Non-Reasoning) | 79.72%±1.87 | $0.2/$0.5 | 8.48s
Every result from this document58Gemini 3.6 Flash model ID google/gemini-3.6-flash; temperature=1; max_output_tokens=30000; reaso… 79.66
Printed as 79.66%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-3.6-flash; temperature=1; max_output_tokens=30000; reasoning_effort=highVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 58 of 105 (Gemini 3.6 Flash), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-3.6-flash"].58 | Gemini 3.6 Flash | 79.66%±1.86 | $1.5/$7.5 | 33.55s
Every result from this document59Qwen 3.7 Max model ID alibaba/qwen3.7-max; temperature=1; max_output_tokens=30000 79.40
Printed as 79.40%Official leaderboard, measured Sep 2026Configuration: model ID alibaba/qwen3.7-max; temperature=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 59 of 105 (Qwen 3.7 Max), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["alibaba/qwen3.7-max"].59 | Qwen 3.7 Max | 79.40%±1.91 | $2.5/$7.5 | 108.05s
Every result from this document60Grok 4.1 Fast (Reasoning) model ID grok/grok-4-1-fast-reasoning; temperature=1; top_p=0.95; max_output_to… 78.73
Printed as 78.73%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4-1-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 60 of 105 (Grok 4.1 Fast (Reasoning)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["grok/grok-4-1-fast-reasoning"].60 | Grok 4.1 Fast (Reasoning) | 78.73%±1.87 | $0.2/$0.5 | 39.29s
Every result from this document61Gemini 2.5 Flash Preview (9/25) (Thinking) model ID google/gemini-2.5-flash-preview-09-2025-thinking; temperature=1; max_o… 78.50
Printed as 78.50%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-2.5-flash-preview-09-2025-thinking; temperature=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 61 of 105 (Gemini 2.5 Flash Preview (9/25) (Thinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-2.5-flash-preview-09-2025-thinking"].61 | Gemini 2.5 Flash Preview (9/25) (Thinking) | 78.50%±1.99 | $0.3/$2.5 | 31.39s
Every result from this document62Grok 4 model ID grok/grok-4-0709; temperature=1; top_p=0.95; max_output_tokens=30000 78.15
Printed as 78.15%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4-0709; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 62 of 105 (Grok 4), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["grok/grok-4-0709"].62 | Grok 4 | 78.15%±2.08 | $3/$15 | 74.05s
Every result from this document62Kimi K2.6 model ID kimi/kimi-k2.6; top_p=0.95; max_output_tokens=30000 78.15
Printed as 78.15%Official leaderboard, measured Sep 2026Configuration: model ID kimi/kimi-k2.6; top_p=0.95; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 63 of 105 (Kimi K2.6), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["kimi/kimi-k2.6"].63 | Kimi K2.6 | 78.15%±1.79 | $0.95/$4 | 8m14s
Every result from this document64Gemini 2.5 Flash Preview (9/25) (Nonthinking) model ID google/gemini-2.5-flash-preview-09-2025; temperature=1; max_output_tok… 77.95
Printed as 77.95%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-2.5-flash-preview-09-2025; temperature=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 64 of 105 (Gemini 2.5 Flash Preview (9/25) (Nonthinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-2.5-flash-preview-09-2025"].64 | Gemini 2.5 Flash Preview (9/25) (Nonthinking) | 77.95%±1.92 | $0.3/$2.5 | 22.89s
Every result from this document65GPT 5.4 (xhigh) model ID openai/gpt-5.4-2026-03-05; max_output_tokens=30000; reasoning_effort=x… 77.55
Printed as 77.55%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5.4-2026-03-05; max_output_tokens=30000; reasoning_effort=xhighVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 65 of 105 (GPT 5.4 (xhigh)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-5.4-2026-03-05"].65 | GPT 5.4 (xhigh) | 77.55%±3.32 | $2.5/$15 | 4m43s
Every result from this document66Grok 4.1 Fast Non-Reasoning model ID grok/grok-4-1-fast-non-reasoning; temperature=1; top_p=0.95; max_outpu… 77.46
Printed as 77.46%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4-1-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 66 of 105 (Grok 4.1 Fast Non-Reasoning), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["grok/grok-4-1-fast-non-reasoning"].66 | Grok 4.1 Fast Non-Reasoning | 77.46%±2.04 | $0.2/$0.5 | 19.86s
Every result from this document67Qwen 3 VL Plus model ID alibaba/qwen3-vl-plus-2025-09-23; temperature=1; max_output_tokens=300… 77.13
Printed as 77.13%Official leaderboard, measured Sep 2026Configuration: model ID alibaba/qwen3-vl-plus-2025-09-23; temperature=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 67 of 105 (Qwen 3 VL Plus), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["alibaba/qwen3-vl-plus-2025-09-23"].67 | Qwen 3 VL Plus | 77.13%±1.92 | $0.2/$1.6 | 71.35s
Every result from this document68GPT 5.4 Nano model ID openai/gpt-5.4-nano-2026-03-17; max_output_tokens=30000; reasoning_eff… 77.09
Printed as 77.09%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5.4-nano-2026-03-17; max_output_tokens=30000; reasoning_effort=highVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 68 of 105 (GPT 5.4 Nano), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-5.4-nano-2026-03-17"].68 | GPT 5.4 Nano | 77.09%±1.89 | $0.2/$1.25 | 20.53s
Every result from this document69Qwen 3.6 Plus model ID alibaba/qwen3.6-plus; temperature=1; max_output_tokens=30000 76.96
Printed as 76.96%Official leaderboard, measured Sep 2026Configuration: model ID alibaba/qwen3.6-plus; temperature=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 69 of 105 (Qwen 3.6 Plus), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["alibaba/qwen3.6-plus"].69 | Qwen 3.6 Plus | 76.96%±1.92 | $0.5/$3 | 2m54s
Every result from this document70o3 model ID openai/o3-2025-04-16; max_output_tokens=30000; reasoning_effort=high 76.65
Printed as 76.65%Official leaderboard, measured Sep 2026Configuration: model ID openai/o3-2025-04-16; max_output_tokens=30000; reasoning_effort=highVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 70 of 105 (o3), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/o3-2025-04-16"].70 | o3 | 76.65%±1.87 | $2/$8 | 46.39s
Every result from this document71Gemini 3.5 Flash model ID google/gemini-3.5-flash; temperature=1; max_output_tokens=30000; reaso… 76.57
Printed as 76.57%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-3.5-flash; temperature=1; max_output_tokens=30000; reasoning_effort=highVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 71 of 105 (Gemini 3.5 Flash), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-3.5-flash"].71 | Gemini 3.5 Flash | 76.57%±1.92 | $1.5/$9 | 57.08s
Every result from this document72Kimi K2.5 model ID kimi/kimi-k2.5-thinking; top_p=0.95; max_output_tokens=30000 76.44
Printed as 76.44%Official leaderboard, measured Sep 2026Configuration: model ID kimi/kimi-k2.5-thinking; top_p=0.95; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 72 of 105 (Kimi K2.5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["kimi/kimi-k2.5-thinking"].72 | Kimi K2.5 | 76.44%±1.99 | $0.6/$3 | 2m27s
Every result from this document73Gemini 3.1 Pro Preview (02/26) model ID google/gemini-3.1-pro-preview; temperature=1; max_output_tokens=30000;… 76.11
Printed as 76.11%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-3.1-pro-preview; temperature=1; max_output_tokens=30000; reasoning_effort=highVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 73 of 105 (Gemini 3.1 Pro Preview (02/26)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-3.1-pro-preview"].73 | Gemini 3.1 Pro Preview (02/26) | 76.11%±1.92 | $2/$12 | 69.14s
Every result from this document74Claude Sonnet 5 model ID anthropic/claude-sonnet-5; temperature=1; max_output_tokens=30000; com… 76.05
Printed as 76.05%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-sonnet-5; temperature=1; max_output_tokens=30000; compute_effort=maxVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 74 of 105 (Claude Sonnet 5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-sonnet-5"].74 | Claude Sonnet 5 | 76.05%±3.05 | $2/$10 | 4m12s
Every result from this document75Gemini 2.5 Flash Lite (9/25) (Nonthinking) model ID google/gemini-2.5-flash-lite-preview-09-2025; temperature=1; max_outpu… 75.82
Printed as 75.82%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-2.5-flash-lite-preview-09-2025; temperature=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 75 of 105 (Gemini 2.5 Flash Lite (9/25) (Nonthinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-2.5-flash-lite-preview-09-2025"].75 | Gemini 2.5 Flash Lite (9/25) (Nonthinking) | 75.82%±1.85 | $0.1/$0.4 | 4.49s
Every result from this document76Ling 3.0 Flash Fin model ID ant/ling-3.0-flash-af-rc3; temperature=1; top_p=0.95; max_output_token… 75.59
Printed as 75.59%Official leaderboard, measured Sep 2026Configuration: model ID ant/ling-3.0-flash-af-rc3; temperature=1; top_p=0.95; max_output_tokens=131072Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 76 of 105 (Ling 3.0 Flash Fin), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["ant/ling-3.0-flash-af-rc3"].76 | Ling 3.0 Flash Fin | 75.59%±2.03 | $0.06/$0.18 | 44.81s
Every result from this document77DeepSeek V4 model ID deepseek/deepseek-v4-pro; max_output_tokens=128000; reasoning_effort=m… 75.14
Printed as 75.14%Official leaderboard, measured Sep 2026Configuration: model ID deepseek/deepseek-v4-pro; max_output_tokens=128000; reasoning_effort=maxVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 77 of 105 (DeepSeek V4), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["deepseek/deepseek-v4-pro"].77 | DeepSeek V4 | 75.14%±2.00 | $1.32/$3.96 | 5m46s
Every result from this document78Grok 4.3 model ID grok/grok-4.3; temperature=1; top_p=0.95; max_output_tokens=30000 74.40
Printed as 74.40%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4.3; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 78 of 105 (Grok 4.3), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["grok/grok-4.3"].78 | Grok 4.3 | 74.40%±2.02 | $1.25/$2.5 | 100.46s
Every result from this document79Claude Opus 4.1 (Thinking) model ID anthropic/claude-opus-4-1-20250805-thinking; temperature=1; max_output… 73.90
Printed as 73.90%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-1-20250805-thinking; temperature=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 79 of 105 (Claude Opus 4.1 (Thinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-opus-4-1-20250805-thinking"].79 | Claude Opus 4.1 (Thinking) | 73.90%±1.97 | $15/$75 | 57.40s
Every result from this document80Gemini 2.5 Pro model ID google/gemini-2.5-pro; temperature=1; max_output_tokens=30000 73.55
Printed as 73.55%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-2.5-pro; temperature=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 80 of 105 (Gemini 2.5 Pro), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-2.5-pro"].80 | Gemini 2.5 Pro | 73.55%±1.91 | $1.25/$10 | 35.91s
Every result from this document81GPT 5 Nano model ID openai/gpt-5-nano-2025-08-07; max_output_tokens=30000; reasoning_effor… 72.86
Printed as 72.86%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5-nano-2025-08-07; max_output_tokens=30000; reasoning_effort=highVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 81 of 105 (GPT 5 Nano), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-5-nano-2025-08-07"].81 | GPT 5 Nano | 72.86%±1.89 | $0.05/$0.4 | 112.91s
Every result from this document82Gemini 2.5 Flash Lite (Nonthinking) model ID google/gemini-2.5-flash-lite; temperature=1; max_output_tokens=30000 72.83
Printed as 72.83%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-2.5-flash-lite; temperature=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 82 of 105 (Gemini 2.5 Flash Lite (Nonthinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-2.5-flash-lite"].82 | Gemini 2.5 Flash Lite (Nonthinking) | 72.83%±1.98 | $0.1/$0.4 | 4.85s
Every result from this document83Qwen 3 Max Thinking model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=30000 72.71
Printed as 72.71%Official leaderboard, measured Sep 2026Configuration: model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 83 of 105 (Qwen 3 Max Thinking), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["alibaba/qwen3-max-2026-01-23"].83 | Qwen 3 Max Thinking | 72.71%±1.91 | $1.2/$6 | 6m02s
Every result from this document84Claude Sonnet 4 (Nonthinking) model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=3… 72.41
Printed as 72.41%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 84 of 105 (Claude Sonnet 4 (Nonthinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-sonnet-4-20250514"].84 | Claude Sonnet 4 (Nonthinking) | 72.41%±1.93 | $3/$15 | 25.67s
Every result from this document85GLM 5.1 model ID zai/glm-5.1; temperature=1; top_p=0.95; max_output_tokens=30000 72.27
Printed as 72.27%Official leaderboard, measured Sep 2026Configuration: model ID zai/glm-5.1; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 85 of 105 (GLM 5.1), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["zai/glm-5.1"].85 | GLM 5.1 | 72.27%±2.06 | $1/$3.2 | 95.70s
Every result from this document86MiMo V2.5 model ID xiaomi/mimo-v2.5; temperature=1; top_p=0.95; max_output_tokens=30000 72.15
Printed as 72.15%Official leaderboard, measured Sep 2026Configuration: model ID xiaomi/mimo-v2.5; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 86 of 105 (MiMo V2.5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["xiaomi/mimo-v2.5"].86 | MiMo V2.5 | 72.15%±1.85 | $0.14/$0.28 | 20.26s
Every result from this document87Gemini 3 Pro (11/25) model ID google/gemini-3-pro-preview; temperature=1; max_output_tokens=30000; r… 72.04
Printed as 72.04%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-3-pro-preview; temperature=1; max_output_tokens=30000; reasoning_effort=highVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 87 of 105 (Gemini 3 Pro (11/25)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-3-pro-preview"].87 | Gemini 3 Pro (11/25) | 72.04%±1.90 | $2/$12 | 43.39s
Every result from this document88Claude Opus 4.1 (Nonthinking) model ID anthropic/claude-opus-4-1-20250805; temperature=1; max_output_tokens=3… 71.75
Printed as 71.75%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-1-20250805; temperature=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 88 of 105 (Claude Opus 4.1 (Nonthinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-opus-4-1-20250805"].88 | Claude Opus 4.1 (Nonthinking) | 71.75%±2.02 | $15/$75 | 38.04s
Every result from this document89Gemini 3.5 Flash Lite model ID google/gemini-3.5-flash-lite; temperature=1; max_output_tokens=30000;… 70.89
Printed as 70.89%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-3.5-flash-lite; temperature=1; max_output_tokens=30000; reasoning_effort=highVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 89 of 105 (Gemini 3.5 Flash Lite), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-3.5-flash-lite"].89 | Gemini 3.5 Flash Lite | 70.89%±2.03 | $0.3/$2.5 | 18.99s
Every result from this document90Qwen 3.5 Flash model ID alibaba/qwen3.5-flash; temperature=1; max_output_tokens=30000 70.62
Printed as 70.62%Official leaderboard, measured Sep 2026Configuration: model ID alibaba/qwen3.5-flash; temperature=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 90 of 105 (Qwen 3.5 Flash), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["alibaba/qwen3.5-flash"].90 | Qwen 3.5 Flash | 70.62%±2.09 | $0.1/$0.4 | 80.41s
Every result from this document91Gemini 3 Flash (12/25) model ID google/gemini-3-flash-preview; temperature=1; max_output_tokens=30000;… 69.92
Printed as 69.92%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-3-flash-preview; temperature=1; max_output_tokens=30000; reasoning_effort=highVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 91 of 105 (Gemini 3 Flash (12/25)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-3-flash-preview"].91 | Gemini 3 Flash (12/25) | 69.92%±1.90 | $0.5/$3 | 23.49s
Every result from this document92Claude Sonnet 4 (Thinking) model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=30000 69.35
Printed as 69.35%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 92 of 105 (Claude Sonnet 4 (Thinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-sonnet-4-20250514-thinking"].92 | Claude Sonnet 4 (Thinking) | 69.35%±2.21 | $3/$15 | 39.57s
Every result from this document93o4 Mini model ID openai/o4-mini-2025-04-16; max_output_tokens=30000; reasoning_effort=h… 69.14
Printed as 69.14%Official leaderboard, measured Sep 2026Configuration: model ID openai/o4-mini-2025-04-16; max_output_tokens=30000; reasoning_effort=highVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 93 of 105 (o4 Mini), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/o4-mini-2025-04-16"].93 | o4 Mini | 69.14%±1.96 | $1.1/$4.4 | 81.96s
Every result from this document94GLM 4.7 model ID zai/glm-4.7; temperature=1; top_p=1; max_output_tokens=30000 68.63
Printed as 68.63%Official leaderboard, measured Sep 2026Configuration: model ID zai/glm-4.7; temperature=1; top_p=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 94 of 105 (GLM 4.7), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["zai/glm-4.7"].94 | GLM 4.7 | 68.63%±2.12 | $0.6/$2.2 | 2m47s
Every result from this document95Mistral Medium 3.5 model ID mistralai/mistral-medium-3.5; temperature=1; top_p=0.95; max_output_to… 67.73
Printed as 67.73%Official leaderboard, measured Sep 2026Configuration: model ID mistralai/mistral-medium-3.5; temperature=1; top_p=0.95; max_output_tokens=30000; reasoning_effort=highVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 95 of 105 (Mistral Medium 3.5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["mistralai/mistral-medium-3.5"].95 | Mistral Medium 3.5 | 67.73%±2.01 | $1.5/$7.5 | 104.05s
Every result from this document96Gemini 2.5 Flash Lite (9/25) (Thinking) model ID google/gemini-2.5-flash-lite-preview-09-2025-thinking; temperature=1;… 66.88
Printed as 66.88%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-2.5-flash-lite-preview-09-2025-thinking; temperature=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 96 of 105 (Gemini 2.5 Flash Lite (9/25) (Thinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-2.5-flash-lite-preview-09-2025-thinking"].96 | Gemini 2.5 Flash Lite (9/25) (Thinking) | 66.88%±1.92 | $0.1/$0.4 | 11.52s
Every result from this document97Laguna M.1 model ID poolside/laguna-m.1; temperature=1; max_output_tokens=30000 65.91
Printed as 65.91%Official leaderboard, measured Sep 2026Configuration: model ID poolside/laguna-m.1; temperature=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 97 of 105 (Laguna M.1), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["poolside/laguna-m.1"].97 | Laguna M.1 | 65.91%±2.01 | N/A | 111.82s
Every result from this document98Gemini 3.1 Flash Lite Preview model ID google/gemini-3.1-flash-lite-preview; temperature=1; max_output_tokens… 63.90
Printed as 63.90%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-3.1-flash-lite-preview; temperature=1; max_output_tokens=30000; reasoning_effort=highVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 98 of 105 (Gemini 3.1 Flash Lite Preview), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["google/gemini-3.1-flash-lite-preview"].98 | Gemini 3.1 Flash Lite Preview | 63.90%±1.82 | $0.25/$1.5 | 16.79s
Every result from this document99Grok 4.20 (Reasoning) model ID grok/grok-4.20-0309-reasoning; temperature=1; top_p=0.95; max_output_t… 63.41
Printed as 63.41%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4.20-0309-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 99 of 105 (Grok 4.20 (Reasoning)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["grok/grok-4.20-0309-reasoning"].99 | Grok 4.20 (Reasoning) | 63.41%±2.10 | $2/$6 | 18.60s
Every result from this document100Laguna XS.2 model ID poolside/laguna-xs.2; temperature=1; max_output_tokens=30000 61.43
Printed as 61.43%Official leaderboard, measured Sep 2026Configuration: model ID poolside/laguna-xs.2; temperature=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 100 of 105 (Laguna XS.2), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["poolside/laguna-xs.2"].100 | Laguna XS.2 | 61.43%±2.35 | N/A | 68.82s
Every result from this document101Command A+ model ID cohere/command-a-plus-05-2026; temperature=1; top_p=0.95; max_output_t… 55.68
Printed as 55.68%Official leaderboard, measured Sep 2026Configuration: model ID cohere/command-a-plus-05-2026; temperature=1; top_p=0.95; max_output_tokens=64000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 101 of 105 (Command A+), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["cohere/command-a-plus-05-2026"].101 | Command A+ | 55.68%±3.65 | N/A | 3m23s
Every result from this document102Mercury 2.5 model ID inception/mercury-2.5; temperature=1; max_output_tokens=65536; reasoni… 55.09
Printed as 55.09%Official leaderboard, measured Sep 2026Configuration: model ID inception/mercury-2.5; temperature=1; max_output_tokens=65536; reasoning_effort=highVals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 102 of 105 (Mercury 2.5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["inception/mercury-2.5"].102 | Mercury 2.5 | 55.09%±2.10 | $0.2/$0.75 | 9.99s
Every result from this document103Llama 4 Maverick model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_to… 54.22
Printed as 54.22%Official leaderboard, measured Sep 2026Configuration: model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 103 of 105 (Llama 4 Maverick), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["fireworks/llama4-maverick-instruct-basic"].103 | Llama 4 Maverick | 54.22%±1.87 | $0.22/$0.88 | 25.05s
Every result from this document104Llama 4 Scout model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max… 50.59
Printed as 50.59%Official leaderboard, measured Sep 2026Configuration: model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 104 of 105 (Llama 4 Scout), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["together/meta-llama/Llama-4-Scout-17B-16E-Instruct"].104 | Llama 4 Scout | 50.59%±1.90 | $0.18/$0.59 | 11.32s
Every result from this document105Nemotron 3.5 Lightning model ID fireworks/nemotron-lightning-3p5-30b-a3b; temperature=1; top_p=0.95; m… 4.27
Printed as 4.27%Official leaderboard, measured Sep 2026Configuration: model ID fireworks/nemotron-lightning-3p5-30b-a3b; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 105 of 105 (Nemotron 3.5 Lightning), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["fireworks/nemotron-lightning-3p5-30b-a3b"].105 | Nemotron 3.5 Lightning | 4.27%±0.49 | $0.05/$0.2 | 105.52s
Every result from this document
Documents
1
- Vals AI MedScribe leaderboardofficial leaderboard, Vals AI, 3 Sep 2026Results it supports