MedCode (Vals AI)
ICD-10-CM coding of whole hospital stays from discharge summaries and notes, checked against professional coders.
More about this boardLess
ICD-10-CM coding of whole hospital stays from discharge summaries and notes, checked against professional coders.
ICD-10-CM diagnosis coding for entire hospital stays: models assign primary and secondary codes from discharge summaries plus progress/consult notes; ground truth double-annotated by certified professional coders.
Vals AI runs every model itself; 103 models on the board, last updated September 29, 2026. Vals AI pairs this board with MedScribe, and its writeup notes that coding accuracy lags documentation quality.
Published by Vals AI (dataset with Protege), released Feb 2026. 2,755 patient records. Percentage accuracy 0-100, higher better.
- Rows on the board
- 103
- Board last updated
- 29 Sep 2026
- Rows
- 103
- Models
- 94 of 103on the board
- Labs
- 19
- Last measured
- Sep 2026
- Leader
- 63.57Claude Opus 5
Ranking
Percentage accuracy 0-100, higher better
- Anthropic
- OpenAI
- Meta
- Other labs
20 of 103 rows
Rows and sources
Open a row for the quote, the page and the document.
1Claude Opus 5 63.57
Printed as 63.57%Official leaderboard, measured Sep 2026Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 1 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-opus-5"].1 | Claude Opus 5 | 63.57%±1.99 | $5/$25 | 24.50s
Every result from this document2Gemini 3.1 Pro Preview (02/26) 59.06
Printed as 59.06%Official leaderboard, measured Sep 2026Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 2 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["google/gemini-3.1-pro-preview"].2 | Gemini 3.1 Pro Preview (02/26) | 59.06%±2.00 | $2/$12 | 38.52s
Every result from this document3Claude Fable 5 56.07
Printed as 56.07%Official leaderboard, measured Sep 2026Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 3 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-fable-5"].3 | Claude Fable 5 | 56.07%±2.20 | $10/$50 | 91.44s
Every result from this document4Gemini 3 Flash (12/25) 55.92
Printed as 55.92%Official leaderboard, measured Sep 2026Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 4 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["google/gemini-3-flash-preview"].4 | Gemini 3 Flash (12/25) | 55.92%±2.11 | $0.5/$3 | 44.15s
Every result from this document5Gemini 3.5 Flash 55.83
Printed as 55.83%Official leaderboard, measured Sep 2026Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 5 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["google/gemini-3.5-flash"].5 | Gemini 3.5 Flash | 55.83%±2.11 | $1.5/$9 | 25.29s
Every result from this document6Claude Opus 4.7 54.86
Printed as 54.86%Official leaderboard, measured Sep 2026Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 6 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-opus-4-7"].6 | Claude Opus 4.7 | 54.86%±2.21 | $5/$25 | 54.25s
Every result from this document7Claude Fable 5.1 53.51
Printed as 53.51%Official leaderboard, measured Sep 2026Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 7 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-fable-5-1"].7 | Claude Fable 5.1 | 53.51%±2.17 | $10/$50 | 3m35s
Every result from this document8Gemini 3.7 Flash model ID google/gemini-3.7-flash; reasoning_effort=high; temperature=1; max_out… 53.39
Printed as 53.39%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-3.7-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 8 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["google/gemini-3.7-flash"].8 | Gemini 3.7 Flash | 53.39%±2.12 | $1.5/$7.5 | 9.16s
Every result from this document9Claude Opus 4.8 53.22
Printed as 53.22%Official leaderboard, measured Sep 2026Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 9 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-opus-4-8"].9 | Claude Opus 4.8 | 53.22%±2.17 | $5/$25 | 105.92s
Every result from this document10Gemini 3.6 Flash 53.15
Printed as 53.15%Official leaderboard, measured Sep 2026Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 10 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["google/gemini-3.6-flash"].10 | Gemini 3.6 Flash | 53.15%±2.16 | $1.5/$7.5 | 16.68s
Every result from this document11Claude Sonnet 5.5 model ID anthropic/claude-sonnet-5-5; compute_effort=max; temperature=1; max_ou… 52.92
Printed as 52.92%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-sonnet-5-5; compute_effort=max; temperature=1; max_output_tokens=128000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 11 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-sonnet-5-5"].11 | Claude Sonnet 5.5 | 52.92%±2.12 | $2/$10 | 3m59s
Every result from this document12GPT 5.1 52.73
Printed as 52.73%Official leaderboard, measured Sep 2026Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 12 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["openai/gpt-5.1-2025-11-13"].12 | GPT 5.1 | 52.73%±2.15 | $1.25/$10 | 54.55s
Every result from this document13Gemini 3 Pro (11/25) model ID google/gemini-3-pro-preview; reasoning_effort=high; temperature=1; max… 52.20
Printed as 52.20%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-3-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 13 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["google/gemini-3-pro-preview"].13 | Gemini 3 Pro (11/25) | 52.20%±2.07 | $2/$12 | 55.84s
Every result from this document14Muse Spark model ID meta/muse_spark; temperature=1; max_output_tokens=30000 51.31
Printed as 51.31%Official leaderboard, measured Sep 2026Configuration: model ID meta/muse_spark; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 14 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["meta/muse_spark"].14 | Muse Spark | 51.31%±2.24 | N/A | 2m03s
Every result from this document15Gemini 2.5 Pro model ID google/gemini-2.5-pro; temperature=1; max_output_tokens=30000 50.59
Printed as 50.59%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-2.5-pro; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 15 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["google/gemini-2.5-pro"].15 | Gemini 2.5 Pro | 50.59%±2.11 | $1.25/$10 | 27.41s
Every result from this document16Claude Opus 5.5 model ID anthropic/claude-opus-5-5; compute_effort=max; temperature=1; max_outp… 49.80
Printed as 49.80%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-5-5; compute_effort=max; temperature=1; max_output_tokens=128000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 16 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-opus-5-5"].16 | Claude Opus 5.5 | 49.80%±2.27 | $4/$20 | 4m07s
Every result from this document17GPT 5.2 model ID openai/gpt-5.2-2025-12-11; reasoning_effort=xhigh; max_output_tokens=3… 49.75
Printed as 49.75%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5.2-2025-12-11; reasoning_effort=xhigh; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 17 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["openai/gpt-5.2-2025-12-11"].17 | GPT 5.2 | 49.75%±2.26 | $1.75/$14 | 2m37s
Every result from this document18GPT 5 model ID openai/gpt-5-2025-08-07; reasoning_effort=high; max_output_tokens=30000 49.63
Printed as 49.63%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5-2025-08-07; reasoning_effort=high; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 18 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["openai/gpt-5-2025-08-07"].18 | GPT 5 | 49.63%±2.10 | $1.25/$10 | 58.27s
Every result from this document19Grok 4.7 model ID grok/grok-4.7; reasoning_effort=xhigh; temperature=1; top_p=0.95 49.55
Printed as 49.55%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4.7; reasoning_effort=xhigh; temperature=1; top_p=0.95Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 19 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["grok/grok-4.7"].19 | Grok 4.7 | 49.55%±2.17 | $2/$6 | 3m33s
Every result from this document20Muse Spark 1.2 model ID meta/muse_spark_1_2; reasoning_effort=xhigh; temperature=1; max_output… 49.35
Printed as 49.35%Official leaderboard, measured Sep 2026Configuration: model ID meta/muse_spark_1_2; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 20 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["meta/muse_spark_1_2"].20 | Muse Spark 1.2 | 49.35%±2.19 | $1.25/$4.25 | 60.17s
Every result from this document21Claude Opus 4.5 (Thinking) model ID anthropic/claude-opus-4-5-20251101-thinking; compute_effort=high; temp… 49.16
Printed as 49.16%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-5-20251101-thinking; compute_effort=high; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 21 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-opus-4-5-20251101-thinking"].21 | Claude Opus 4.5 (Thinking) | 49.16%±2.01 | $5/$25 | 60.84s
Every result from this document22Claude Opus 4.6 (Thinking) model ID anthropic/claude-opus-4-6-thinking; compute_effort=max; temperature=1;… 49.13
Printed as 49.13%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-6-thinking; compute_effort=max; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 22 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-opus-4-6-thinking"].22 | Claude Opus 4.6 (Thinking) | 49.13%±2.08 | $5/$25 | 2m36s
Every result from this document23GPT 5.5 model ID openai/gpt-5.5; reasoning_effort=xhigh; temperature=1; max_output_toke… 49.10
Printed as 49.10%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5.5; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 23 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["openai/gpt-5.5"].23 | GPT 5.5 | 49.10%±2.19 | $5/$30 | 2m40s
Every result from this document24Kimi K3 model ID kimi/kimi-k3; temperature=1; max_output_tokens=30000 48.88
Printed as 48.88%Official leaderboard, measured Sep 2026Configuration: model ID kimi/kimi-k3; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 24 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["kimi/kimi-k3"].24 | Kimi K3 | 48.88%±2.19 | $3/$15 | 116.18s
Every result from this document25GPT-6.1 Sol model ID openai/gpt-6.1-sol; reasoning_effort=max; max_output_tokens=128000 48.84
Printed as 48.84%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-6.1-sol; reasoning_effort=max; max_output_tokens=128000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 25 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["openai/gpt-6.1-sol"].25 | GPT-6.1 Sol | 48.84%±2.12 | $2/$10 | 115.34s
Every result from this document26GPT-6 Astra 48.49
Printed as 48.49%Official leaderboard, measured Sep 2026Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 26 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["openai/gpt-6-astra"].26 | GPT-6 Astra | 48.49%±2.13 | $10/$50 | 102.86s
Every result from this document27Claude Opus 4.6 (Nonthinking) model ID anthropic/claude-opus-4-6; compute_effort=max; temperature=1; max_outp… 48.24
Printed as 48.24%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-6; compute_effort=max; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 27 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-opus-4-6"].27 | Claude Opus 4.6 (Nonthinking) | 48.24%±2.05 | $5/$25 | 4.06s
Every result from this document28Gemini 3.8 Flash model ID google/gemini-3.8-flash; reasoning_effort=high; temperature=1; max_out… 48.13
Printed as 48.13%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-3.8-flash; reasoning_effort=high; temperature=1; max_output_tokens=65536Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 28 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["google/gemini-3.8-flash"].28 | Gemini 3.8 Flash | 48.13%±2.18 | $1.5/$7.5 | 42.89s
Every result from this document29Gemini 3.1 Flash Lite Preview model ID google/gemini-3.1-flash-lite-preview; reasoning_effort=high; temperatu… 47.60
Printed as 47.60%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-3.1-flash-lite-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 29 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["google/gemini-3.1-flash-lite-preview"].29 | Gemini 3.1 Flash Lite Preview | 47.60%±2.07 | $0.25/$1.5 | 8.55s
Every result from this document30Claude Sonnet 5 model ID anthropic/claude-sonnet-5; compute_effort=max; temperature=1; max_outp… 47.54
Printed as 47.54%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-sonnet-5; compute_effort=max; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 30 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-sonnet-5"].30 | Claude Sonnet 5 | 47.54%±2.27 | $2/$10 | 2m15s
Every result from this document31o3 model ID openai/o3-2025-04-16; reasoning_effort=high; max_output_tokens=30000 47.29
Printed as 47.29%Official leaderboard, measured Sep 2026Configuration: model ID openai/o3-2025-04-16; reasoning_effort=high; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 31 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["openai/o3-2025-04-16"].31 | o3 | 47.29%±2.16 | $2/$8 | 17.68s
Every result from this document32Claude Opus 4.1 (Thinking) model ID anthropic/claude-opus-4-1-20250805-thinking; temperature=1; max_output… 47.23
Printed as 47.23%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-1-20250805-thinking; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 32 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-opus-4-1-20250805-thinking"].32 | Claude Opus 4.1 (Thinking) | 47.23%±2.07 | $15/$75 | 33.26s
Every result from this document33GPT-6 Sol model ID openai/gpt-6-sol; reasoning_effort=max; max_output_tokens=128000 47.07
Printed as 47.07%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-6-sol; reasoning_effort=max; max_output_tokens=128000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 33 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["openai/gpt-6-sol"].33 | GPT-6 Sol | 47.07%±2.12 | $2/$10 | 62.71s
Every result from this document34MiniMax-M3 model ID minimax/MiniMax-M3; temperature=1; top_p=0.95; max_output_tokens=30000 46.29
Printed as 46.29%Official leaderboard, measured Sep 2026Configuration: model ID minimax/MiniMax-M3; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 34 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["minimax/MiniMax-M3"].34 | MiniMax-M3 | 46.29%±2.10 | $0.6/$2.4 | 63.12s
Every result from this document35Claude Opus 4.5 (Nonthinking) model ID anthropic/claude-opus-4-5-20251101; compute_effort=high; temperature=1… 45.17
Printed as 45.17%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-5-20251101; compute_effort=high; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 35 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-opus-4-5-20251101"].35 | Claude Opus 4.5 (Nonthinking) | 45.17%±1.89 | $5/$25 | 5.02s
Every result from this document36MiMo V2.6 Pro model ID xiaomi/mimo-v2.6-pro; temperature=1; top_p=0.95; max_output_tokens=128… 44.97
Printed as 44.97%Official leaderboard, measured Sep 2026Configuration: model ID xiaomi/mimo-v2.6-pro; temperature=1; top_p=0.95; max_output_tokens=128000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 36 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["xiaomi/mimo-v2.6-pro"].36 | MiMo V2.6 Pro | 44.97%±2.10 | $0.435/$0.87 | 2m57s
Every result from this document37Grok 4.6 model ID grok/grok-4.6; reasoning_effort=high; temperature=1; top_p=0.95; max_o… 44.71
Printed as 44.71%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4.6; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 37 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["grok/grok-4.6"].37 | Grok 4.6 | 44.71%±2.26 | $2/$6 | 2m07s
Every result from this document38GPT-6 Luna model ID openai/gpt-6-luna; reasoning_effort=max; max_output_tokens=128000 44.69
Printed as 44.69%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-6-luna; reasoning_effort=max; max_output_tokens=128000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 38 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["openai/gpt-6-luna"].38 | GPT-6 Luna | 44.69%±2.30 | $0.1/$0.5 | 108.77s
Every result from this document39Claude Sonnet 4.5 (Thinking) model ID anthropic/claude-sonnet-4-5-20250929-thinking; temperature=1; max_outp… 44.13
Printed as 44.13%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-sonnet-4-5-20250929-thinking; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 39 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-sonnet-4-5-20250929-thinking"].39 | Claude Sonnet 4.5 (Thinking) | 44.13%±2.00 | $3/$15 | 74.33s
Every result from this document40GPT-5.6 Sol model ID openai/gpt-5.6-sol; reasoning_effort=max; max_output_tokens=30000 43.97
Printed as 43.97%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5.6-sol; reasoning_effort=max; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 40 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["openai/gpt-5.6-sol"].40 | GPT-5.6 Sol | 43.97%±2.26 | $4/$20 | 96.59s
Every result from this document41Gemini 3.5 Flash Lite model ID google/gemini-3.5-flash-lite; reasoning_effort=high; temperature=1; ma… 43.49
Printed as 43.49%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-3.5-flash-lite; reasoning_effort=high; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 41 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["google/gemini-3.5-flash-lite"].41 | Gemini 3.5 Flash Lite | 43.49%±1.95 | $0.3/$2.5 | 6.02s
Every result from this document42GPT-5.6 Terra model ID openai/gpt-5.6-terra; reasoning_effort=xhigh; max_output_tokens=30000 43.41
Printed as 43.41%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5.6-terra; reasoning_effort=xhigh; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 42 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["openai/gpt-5.6-terra"].42 | GPT-5.6 Terra | 43.41%±2.17 | $2/$12 | 18.41s
Every result from this document43Grok 4.5 model ID grok/grok-4.5; reasoning_effort=high; temperature=1; top_p=0.95; max_o… 43.29
Printed as 43.29%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4.5; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 43 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["grok/grok-4.5"].43 | Grok 4.5 | 43.29%±2.31 | $2/$6 | 59.64s
Every result from this document44Hy4 Preview model ID tencent/hy4-preview; temperature=1; top_p=1; max_output_tokens=64000 43.25
Printed as 43.25%Official leaderboard, measured Sep 2026Configuration: model ID tencent/hy4-preview; temperature=1; top_p=1; max_output_tokens=64000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 44 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["tencent/hy4-preview"].44 | Hy4 Preview | 43.25%±2.13 | $0.834/$2.501 | 6m39s
Every result from this document45GPT 5 Mini model ID openai/gpt-5-mini-2025-08-07; reasoning_effort=high; max_output_tokens… 43.05
Printed as 43.05%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5-mini-2025-08-07; reasoning_effort=high; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 45 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["openai/gpt-5-mini-2025-08-07"].45 | GPT 5 Mini | 43.05%±2.04 | $0.25/$2 | 28.17s
Every result from this document46GLM 5.3 model ID zai/glm-5.3; reasoning_effort=max; temperature=1; top_p=0.95; max_outp… 42.86
Printed as 42.86%Official leaderboard, measured Sep 2026Configuration: model ID zai/glm-5.3; reasoning_effort=max; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 46 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["zai/glm-5.3"].46 | GLM 5.3 | 42.86%±2.11 | $1.4/$4.4 | 2m49s
Every result from this document47DeepSeek V4 Pro 0813 model ID deepseek/deepseek-v4-pro-0813; reasoning_effort=max; max_output_tokens… 42.47
Printed as 42.47%Official leaderboard, measured Sep 2026Configuration: model ID deepseek/deepseek-v4-pro-0813; reasoning_effort=max; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 47 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["deepseek/deepseek-v4-pro-0813"].47 | DeepSeek V4 Pro 0813 | 42.47%±2.16 | $1.32/$3.96 | 3m01s
Every result from this document48GPT-5.6 Luna model ID openai/gpt-5.6-luna; reasoning_effort=max; max_output_tokens=30000 42.39
Printed as 42.39%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5.6-luna; reasoning_effort=max; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 48 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["openai/gpt-5.6-luna"].48 | GPT-5.6 Luna | 42.39%±2.27 | $0.2/$1.2 | 81.28s
Every result from this document49GLM 5.1 model ID zai/glm-5.1; temperature=1; top_p=0.95; max_output_tokens=30000 41.60
Printed as 41.60%Official leaderboard, measured Sep 2026Configuration: model ID zai/glm-5.1; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 49 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["zai/glm-5.1"].49 | GLM 5.1 | 41.60%±2.12 | $1/$3.2 | 77.58s
Every result from this document50DeepSeek V4 Flash 0731 model ID deepseek/deepseek-v4-flash-0731; reasoning_effort=high; max_output_tok… 41.41
Printed as 41.41%Official leaderboard, measured Sep 2026Configuration: model ID deepseek/deepseek-v4-flash-0731; reasoning_effort=high; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 50 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["deepseek/deepseek-v4-flash-0731"].50 | DeepSeek V4 Flash 0731 | 41.41%±2.15 | $0.44/$1.32 | 2m07s
Every result from this document51Claude Opus 4.1 (Nonthinking) model ID anthropic/claude-opus-4-1-20250805; temperature=1; max_output_tokens=3… 41.37
Printed as 41.37%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-1-20250805; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 51 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-opus-4-1-20250805"].51 | Claude Opus 4.1 (Nonthinking) | 41.37%±1.96 | $15/$75 | 13.08s
Every result from this document52GPT 5.4 (xhigh) model ID openai/gpt-5.4-2026-03-05; reasoning_effort=xhigh; max_output_tokens=3… 41.29
Printed as 41.29%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5.4-2026-03-05; reasoning_effort=xhigh; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 52 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["openai/gpt-5.4-2026-03-05"].52 | GPT 5.4 (xhigh) | 41.29%±2.15 | $2.5/$15 | 3m07s
Every result from this document53Inkling model ID thinkingmachines/inkling; reasoning_effort=0.99; temperature=1; top_p=… 41.19
Printed as 41.19%Official leaderboard, measured Sep 2026Configuration: model ID thinkingmachines/inkling; reasoning_effort=0.99; temperature=1; top_p=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 53 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["thinkingmachines/inkling"].53 | Inkling | 41.19%±2.23 | $1/$4.05 | 2m47s
Every result from this document54DeepSeek V4.1 Flash model ID deepseek/deepseek-v4.1-flash; reasoning_effort=high; max_output_tokens… 41.17
Printed as 41.17%Official leaderboard, measured Sep 2026Configuration: model ID deepseek/deepseek-v4.1-flash; reasoning_effort=high; max_output_tokens=384000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 54 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["deepseek/deepseek-v4.1-flash"].54 | DeepSeek V4.1 Flash | 41.17%±2.04 | $0.3/$1.2 | 35.13s
Every result from this document55MiMo V2.6 Flash model ID xiaomi/mimo-v2.6-flash; temperature=1; top_p=0.95; max_output_tokens=1… 41.06
Printed as 41.06%Official leaderboard, measured Sep 2026Configuration: model ID xiaomi/mimo-v2.6-flash; temperature=1; top_p=0.95; max_output_tokens=128000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 55 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["xiaomi/mimo-v2.6-flash"].55 | MiMo V2.6 Flash | 41.06%±2.03 | $0.14/$0.28 | 93.37s
Every result from this document56GPT 5.4 Nano model ID openai/gpt-5.4-nano-2026-03-17; reasoning_effort=high; max_output_toke… 41.03
Printed as 41.03%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5.4-nano-2026-03-17; reasoning_effort=high; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 56 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["openai/gpt-5.4-nano-2026-03-17"].56 | GPT 5.4 Nano | 41.03%±2.26 | $0.2/$1.25 | 8.04s
Every result from this document57GLM 5.2 model ID zai/glm-5.2; temperature=1; max_output_tokens=30000 40.77
Printed as 40.77%Official leaderboard, measured Sep 2026Configuration: model ID zai/glm-5.2; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 57 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["zai/glm-5.2"].57 | GLM 5.2 | 40.77%±2.17 | $1.4/$4.4 | 95.27s
Every result from this document58Qwen 3.8 Max model ID alibaba/qwen3.8-max; temperature=1; max_output_tokens=30000 40.67
Printed as 40.67%Official leaderboard, measured Sep 2026Configuration: model ID alibaba/qwen3.8-max; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 58 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["alibaba/qwen3.8-max"].58 | Qwen 3.8 Max | 40.67%±2.03 | $2/$6 | 6m18s
Every result from this document59Claude Sonnet 4.5 (Nonthinking) model ID anthropic/claude-sonnet-4-5-20250929; temperature=1; max_output_tokens… 40.57
Printed as 40.57%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-sonnet-4-5-20250929; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 59 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-sonnet-4-5-20250929"].59 | Claude Sonnet 4.5 (Nonthinking) | 40.57%±2.00 | $3/$15 | 12.01s
Every result from this document60Gemini 2.5 Flash Preview (9/25) (Nonthinking) model ID google/gemini-2.5-flash-preview-09-2025; temperature=1; max_output_tok… 40.54
Printed as 40.54%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-2.5-flash-preview-09-2025; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 60 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["google/gemini-2.5-flash-preview-09-2025"].60 | Gemini 2.5 Flash Preview (9/25) (Nonthinking) | 40.54%±1.93 | $0.3/$2.5 | 12.70s
Every result from this document61DeepSeek V4 model ID deepseek/deepseek-v4-pro; reasoning_effort=max; max_output_tokens=1280… 40.45
Printed as 40.45%Official leaderboard, measured Sep 2026Configuration: model ID deepseek/deepseek-v4-pro; reasoning_effort=max; max_output_tokens=128000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 61 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["deepseek/deepseek-v4-pro"].61 | DeepSeek V4 | 40.45%±2.12 | $1.32/$3.96 | 6m23s
Every result from this document62Gemini 2.5 Flash (7/17) (Thinking) model ID google/gemini-2.5-flash-thinking; temperature=1; max_output_tokens=300… 40.36
Printed as 40.36%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-2.5-flash-thinking; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 62 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["google/gemini-2.5-flash-thinking"].62 | Gemini 2.5 Flash (7/17) (Thinking) | 40.36%±1.95 | $0.3/$2.5 | 22.83s
Every result from this document63Gemini 2.5 Flash Preview (9/25) (Thinking) model ID google/gemini-2.5-flash-preview-09-2025-thinking; temperature=1; max_o… 40.33
Printed as 40.33%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-2.5-flash-preview-09-2025-thinking; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 63 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["google/gemini-2.5-flash-preview-09-2025-thinking"].63 | Gemini 2.5 Flash Preview (9/25) (Thinking) | 40.33%±1.92 | $0.3/$2.5 | 16.11s
Every result from this document64Kimi K2.6 model ID kimi/kimi-k2.6; top_p=0.95; max_output_tokens=30000 40.14
Printed as 40.14%Official leaderboard, measured Sep 2026Configuration: model ID kimi/kimi-k2.6; top_p=0.95; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 64 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["kimi/kimi-k2.6"].64 | Kimi K2.6 | 40.14%±2.04 | $0.95/$4 | 5m05s
Every result from this document65Kimi K2.5 model ID kimi/kimi-k2.5-thinking; top_p=0.95; max_output_tokens=30000 39.32
Printed as 39.32%Official leaderboard, measured Sep 2026Configuration: model ID kimi/kimi-k2.5-thinking; top_p=0.95; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 65 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["kimi/kimi-k2.5-thinking"].65 | Kimi K2.5 | 39.32%±2.12 | $0.6/$3 | 75.45s
Every result from this document66Qwen 3.7 Max model ID alibaba/qwen3.7-max; temperature=1; max_output_tokens=30000 38.75
Printed as 38.75%Official leaderboard, measured Sep 2026Configuration: model ID alibaba/qwen3.7-max; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 66 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["alibaba/qwen3.7-max"].66 | Qwen 3.7 Max | 38.75%±2.20 | $2.5/$7.5 | 28.75s
Every result from this document67Nemotron 3 Ultra model ID nvidia/nemotron-3-ultra-550b-a55b; temperature=1; top_p=0.95; max_outp… 38.62
Printed as 38.62%Official leaderboard, measured Sep 2026Configuration: model ID nvidia/nemotron-3-ultra-550b-a55b; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 67 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["nvidia/nemotron-3-ultra-550b-a55b"].67 | Nemotron 3 Ultra | 38.62%±2.00 | N/A | 18.68s
Every result from this document68Gemini 2.5 Flash (7/17) (Nonthinking) model ID google/gemini-2.5-flash; temperature=1; max_output_tokens=30000 38.42
Printed as 38.42%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-2.5-flash; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 68 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["google/gemini-2.5-flash"].68 | Gemini 2.5 Flash (7/17) (Nonthinking) | 38.42%±1.92 | $0.3/$2.5 | 22.14s
Every result from this document69Grok 4 model ID grok/grok-4-0709; temperature=1; top_p=0.95; max_output_tokens=30000 38.08
Printed as 38.08%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4-0709; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 69 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["grok/grok-4-0709"].69 | Grok 4 | 38.08%±2.21 | $3/$15 | 89.00s
Every result from this document70Grok 4.3 model ID grok/grok-4.3; temperature=1; top_p=0.95; max_output_tokens=30000 38.07
Printed as 38.07%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4.3; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 70 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["grok/grok-4.3"].70 | Grok 4.3 | 38.07%±2.08 | $1.25/$2.5 | 43.92s
Every result from this document71Inkling Small model ID thinkingmachines/inkling-small; reasoning_effort=0.99; temperature=1;… 37.89
Printed as 37.89%Official leaderboard, measured Sep 2026Configuration: model ID thinkingmachines/inkling-small; reasoning_effort=0.99; temperature=1; top_p=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 71 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["thinkingmachines/inkling-small"].71 | Inkling Small | 37.89%±2.21 | $0.3/$1.2 | 3m11s
Every result from this document72Grok 4 Fast (Reasoning) model ID grok/grok-4-fast-reasoning; temperature=1; top_p=0.95; max_output_toke… 37.38
Printed as 37.38%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 72 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["grok/grok-4-fast-reasoning"].72 | Grok 4 Fast (Reasoning) | 37.38%±1.94 | $0.2/$0.5 | 17.52s
Every result from this document73Qwen 3.6 Plus model ID alibaba/qwen3.6-plus; temperature=1; max_output_tokens=30000 36.89
Printed as 36.89%Official leaderboard, measured Sep 2026Configuration: model ID alibaba/qwen3.6-plus; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 73 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["alibaba/qwen3.6-plus"].73 | Qwen 3.6 Plus | 36.89%±2.02 | $0.5/$3 | 56.22s
Every result from this document74Llama 4 Maverick model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_to… 36.51
Printed as 36.51%Official leaderboard, measured Sep 2026Configuration: model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 74 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["fireworks/llama4-maverick-instruct-basic"].74 | Llama 4 Maverick | 36.51%±1.99 | $0.22/$0.88 | 21.24s
Every result from this document75Claude Sonnet 4 (Thinking) model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=30000 34.96
Printed as 34.96%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 75 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-sonnet-4-20250514-thinking"].75 | Claude Sonnet 4 (Thinking) | 34.96%±1.94 | $3/$15 | 39.80s
Every result from this document76MiniMax-M2.7 model ID minimax/MiniMax-M2.7; temperature=1; top_p=0.95; max_output_tokens=300… 34.44
Printed as 34.44%Official leaderboard, measured Sep 2026Configuration: model ID minimax/MiniMax-M2.7; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 76 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["minimax/MiniMax-M2.7"].76 | MiniMax-M2.7 | 34.44%±1.99 | $0.3/$1.2 | 30.25s
Every result from this document77Gemini 2.5 Flash Lite (9/25) (Thinking) model ID google/gemini-2.5-flash-lite-preview-09-2025-thinking; temperature=1;… 34.19
Printed as 34.19%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-2.5-flash-lite-preview-09-2025-thinking; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 77 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["google/gemini-2.5-flash-lite-preview-09-2025-thinking"].77 | Gemini 2.5 Flash Lite (9/25) (Thinking) | 34.19%±1.74 | $0.1/$0.4 | 10.49s
Every result from this document78MiniMax-M2.1 model ID minimax/MiniMax-M2.1; temperature=1; top_p=0.95; max_output_tokens=300… 34.08
Printed as 34.08%Official leaderboard, measured Sep 2026Configuration: model ID minimax/MiniMax-M2.1; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 78 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["minimax/MiniMax-M2.1"].78 | MiniMax-M2.1 | 34.08%±1.94 | $0.3/$1.2 | 19.55s
Every result from this document79Claude Sonnet 4 (Nonthinking) model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=3… 33.94
Printed as 33.94%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 79 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-sonnet-4-20250514"].79 | Claude Sonnet 4 (Nonthinking) | 33.94%±1.91 | $3/$15 | 7.30s
Every result from this document80o4 Mini model ID openai/o4-mini-2025-04-16; reasoning_effort=high; max_output_tokens=30… 33.79
Printed as 33.79%Official leaderboard, measured Sep 2026Configuration: model ID openai/o4-mini-2025-04-16; reasoning_effort=high; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 80 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["openai/o4-mini-2025-04-16"].80 | o4 Mini | 33.79%±2.02 | $1.1/$4.4 | 21.06s
Every result from this document81Mistral Medium 3.5 model ID mistralai/mistral-medium-3.5; reasoning_effort=high; temperature=1; to… 33.75
Printed as 33.75%Official leaderboard, measured Sep 2026Configuration: model ID mistralai/mistral-medium-3.5; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 81 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["mistralai/mistral-medium-3.5"].81 | Mistral Medium 3.5 | 33.75%±1.15 | $1.5/$7.5 | 34.88s
Every result from this document82Qwen 3.5 Flash model ID alibaba/qwen3.5-flash; temperature=1; max_output_tokens=30000 33.00
Printed as 33.00%Official leaderboard, measured Sep 2026Configuration: model ID alibaba/qwen3.5-flash; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 82 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["alibaba/qwen3.5-flash"].82 | Qwen 3.5 Flash | 33.00%±1.79 | $0.1/$0.4 | 63.07s
Every result from this document83GLM 4.7 model ID zai/glm-4.7; temperature=1; top_p=1; max_output_tokens=30000 32.77
Printed as 32.77%Official leaderboard, measured Sep 2026Configuration: model ID zai/glm-4.7; temperature=1; top_p=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 83 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["zai/glm-4.7"].83 | GLM 4.7 | 32.77%±2.00 | $0.6/$2.2 | 2m03s
Every result from this document84Claude Haiku 4.5 (Thinking) model ID anthropic/claude-haiku-4-5-20251001-thinking; temperature=1; max_outpu… 32.68
Printed as 32.68%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-haiku-4-5-20251001-thinking; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 84 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-haiku-4-5-20251001-thinking"].84 | Claude Haiku 4.5 (Thinking) | 32.68%±2.00 | $1/$5 | 34.29s
Every result from this document85MiMo V2.5 Pro model ID xiaomi/mimo-v2.5-pro; temperature=1; top_p=0.95; max_output_tokens=300… 32.48
Printed as 32.48%Official leaderboard, measured Sep 2026Configuration: model ID xiaomi/mimo-v2.5-pro; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 85 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["xiaomi/mimo-v2.5-pro"].85 | MiMo V2.5 Pro | 32.48%±1.91 | $0.435/$0.87 | 38.45s
Every result from this document86Ling 3.0 Flash model ID ant/ling-3.0-flash-2607; temperature=1; top_p=0.95; max_output_tokens=… 32.27
Printed as 32.27%Official leaderboard, measured Sep 2026Configuration: model ID ant/ling-3.0-flash-2607; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 86 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["ant/ling-3.0-flash-2607"].86 | Ling 3.0 Flash | 32.27%±1.91 | $0.075/$0.22 | 9.65s
Every result from this document87Grok 4.20 (Reasoning) model ID grok/grok-4.20-0309-reasoning; temperature=1; top_p=0.95; max_output_t… 32.16
Printed as 32.16%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4.20-0309-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 87 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["grok/grok-4.20-0309-reasoning"].87 | Grok 4.20 (Reasoning) | 32.16%±2.12 | $2/$6 | 16.55s
Every result from this document88MiMo V2.5 model ID xiaomi/mimo-v2.5; temperature=1; top_p=0.95; max_output_tokens=30000 31.89
Printed as 31.89%Official leaderboard, measured Sep 2026Configuration: model ID xiaomi/mimo-v2.5; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 88 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["xiaomi/mimo-v2.5"].88 | MiMo V2.5 | 31.89%±2.02 | $0.14/$0.28 | 17.70s
Every result from this document89Qwen 3 VL Plus model ID alibaba/qwen3-vl-plus-2025-09-23; temperature=1; max_output_tokens=300… 31.65
Printed as 31.65%Official leaderboard, measured Sep 2026Configuration: model ID alibaba/qwen3-vl-plus-2025-09-23; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 89 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["alibaba/qwen3-vl-plus-2025-09-23"].89 | Qwen 3 VL Plus | 31.65%±1.84 | $0.2/$1.6 | 9.95s
Every result from this document90Qwen 3 Max Thinking model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=30000 31.37
Printed as 31.37%Official leaderboard, measured Sep 2026Configuration: model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 90 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["alibaba/qwen3-max-2026-01-23"].90 | Qwen 3 Max Thinking | 31.37%±1.89 | $1.2/$6 | 3m02s
Every result from this document91Mercury 2.5 model ID inception/mercury-2.5; reasoning_effort=high; temperature=1; max_outpu… 31.33
Printed as 31.33%Official leaderboard, measured Sep 2026Configuration: model ID inception/mercury-2.5; reasoning_effort=high; temperature=1; max_output_tokens=65536Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 91 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["inception/mercury-2.5"].91 | Mercury 2.5 | 31.33%±1.95 | $0.2/$0.75 | 7.93s
Every result from this document92GPT 5 Nano model ID openai/gpt-5-nano-2025-08-07; reasoning_effort=high; max_output_tokens… 30.44
Printed as 30.44%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5-nano-2025-08-07; reasoning_effort=high; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 92 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["openai/gpt-5-nano-2025-08-07"].92 | GPT 5 Nano | 30.44%±1.95 | $0.05/$0.4 | 29.74s
Every result from this document93Grok 4 Fast (Non-Reasoning) model ID grok/grok-4-fast-non-reasoning; temperature=1; top_p=0.95; max_output_… 30.04
Printed as 30.04%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 93 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["grok/grok-4-fast-non-reasoning"].93 | Grok 4 Fast (Non-Reasoning) | 30.04%±1.97 | $0.2/$0.5 | 18.21s
Every result from this document94Ling 3.0 Flash Fin model ID ant/ling-3.0-flash-af-rc3; temperature=1; top_p=0.95; max_output_token… 29.30
Printed as 29.30%Official leaderboard, measured Sep 2026Configuration: model ID ant/ling-3.0-flash-af-rc3; temperature=1; top_p=0.95; max_output_tokens=131072Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 94 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["ant/ling-3.0-flash-af-rc3"].94 | Ling 3.0 Flash Fin | 29.30%±1.94 | $0.06/$0.18 | 31.17s
Every result from this document95Qwen 3.8 27B model ID alibaba/qwen3.8-27b; reasoning_effort=xhigh; temperature=1; top_p=0.95… 28.70
Printed as 28.70%Official leaderboard, measured Sep 2026Configuration: model ID alibaba/qwen3.8-27b; reasoning_effort=xhigh; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 95 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["alibaba/qwen3.8-27b"].95 | Qwen 3.8 27B | 28.70%±1.97 | $0.5/$3 | 89.38s
Every result from this document96Grok 4.1 Fast Non-Reasoning model ID grok/grok-4-1-fast-non-reasoning; temperature=1; top_p=0.95; max_outpu… 28.35
Printed as 28.35%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4-1-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 96 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["grok/grok-4-1-fast-non-reasoning"].96 | Grok 4.1 Fast Non-Reasoning | 28.35%±1.92 | $0.2/$0.5 | 4.50s
Every result from this document97Grok 4.1 Fast (Reasoning) model ID grok/grok-4-1-fast-reasoning; temperature=1; top_p=0.95; max_output_to… 28.08
Printed as 28.08%Official leaderboard, measured Sep 2026Configuration: model ID grok/grok-4-1-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 97 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["grok/grok-4-1-fast-reasoning"].97 | Grok 4.1 Fast (Reasoning) | 28.08%±1.99 | $0.2/$0.5 | 46.50s
Every result from this document98Gemini 2.5 Flash Lite (Nonthinking) model ID google/gemini-2.5-flash-lite; temperature=1; max_output_tokens=30000 27.11
Printed as 27.11%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-2.5-flash-lite; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 98 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["google/gemini-2.5-flash-lite"].98 | Gemini 2.5 Flash Lite (Nonthinking) | 27.11%±1.84 | $0.1/$0.4 | 5.75s
Every result from this document99Gemini 2.5 Flash Lite (9/25) (Nonthinking) model ID google/gemini-2.5-flash-lite-preview-09-2025; temperature=1; max_outpu… 27.08
Printed as 27.08%Official leaderboard, measured Sep 2026Configuration: model ID google/gemini-2.5-flash-lite-preview-09-2025; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 99 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["google/gemini-2.5-flash-lite-preview-09-2025"].99 | Gemini 2.5 Flash Lite (9/25) (Nonthinking) | 27.08%±1.91 | $0.1/$0.4 | 6.24s
Every result from this document100Llama 4 Scout model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max… 23.31
Printed as 23.31%Official leaderboard, measured Sep 2026Configuration: model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 100 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["together/meta-llama/Llama-4-Scout-17B-16E-Instruct"].100 | Llama 4 Scout | 23.31%±1.75 | $0.18/$0.59 | 10.74s
Every result from this document101Laguna M.1 model ID poolside/laguna-m.1; temperature=1; max_output_tokens=30000 23.11
Printed as 23.11%Official leaderboard, measured Sep 2026Configuration: model ID poolside/laguna-m.1; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 101 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["poolside/laguna-m.1"].101 | Laguna M.1 | 23.11%±1.69 | N/A | 71.00s
Every result from this document102Laguna XS.2 model ID poolside/laguna-xs.2; temperature=1; max_output_tokens=30000 21.25
Printed as 21.25%Official leaderboard, measured Sep 2026Configuration: model ID poolside/laguna-xs.2; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 102 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["poolside/laguna-xs.2"].102 | Laguna XS.2 | 21.25%±1.70 | N/A | 33.09s
Every result from this document103Command A+ model ID cohere/command-a-plus-05-2026; temperature=1; top_p=0.95; max_output_t… 19.72
Printed as 19.72%Official leaderboard, measured Sep 2026Configuration: model ID cohere/command-a-plus-05-2026; temperature=1; top_p=0.95; max_output_tokens=64000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 103 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["cohere/command-a-plus-05-2026"].103 | Command A+ | 19.72%±1.83 | N/A | 103.75s
Every result from this document
Documents
1
- Vals AI MedCode leaderboardofficial leaderboard, Vals AI, 29 Sep 2026Results it supports