Clinical Benchmarks

Qwen 3 Max Thinking

Also written as alibaba/qwen3-max-2026-01-23, qwen3-max-2026-01-23Compare with other models

Clinical Benchmarks Index
42.9rank 85 of 148; 52.6 × 0.816 = 42.9, from 2 of 10 boards
Boards
2 of 13
Results
2
Latest measurement
Sep 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

Documentation and coding

  1. MedScribe (Vals AI)
    model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=30000
    72.71
    Rank 83 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  2. MedCode (Vals AI)
    model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=30000
    31.37
    Rank 90 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

Sources

Open a line for the quote and page.

  1. 83MedScribe (Vals AI) model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=30000 72.71
    Printed as 72.71%Official leaderboard, measured Sep 2026Configuration: model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 83 of 105 (Qwen 3 Max Thinking), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["alibaba/qwen3-max-2026-01-23"].
    83 | Qwen 3 Max Thinking | 72.71%±1.91 | $1.2/$6 | 6m02s
    Every result from this document
  2. 90MedCode (Vals AI) model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=30000 31.37
    Printed as 31.37%Official leaderboard, measured Sep 2026Configuration: model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=30000
    Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 90 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["alibaba/qwen3-max-2026-01-23"].
    90 | Qwen 3 Max Thinking | 31.37%±1.89 | $1.2/$6 | 3m02s
    Every result from this document

Other Alibaba models: Lingshu-32B, Lingshu-7B, Qwen3-14B, Qwen3-235B-A22B-Instruct-2507, Qwen3-32B, Qwen3-4B, Qwen 3.5, Qwen3.5-27B, Qwen3.5-35B-A3B, Qwen3.5 397B A17B, Qwen3.5-9B, Qwen 3.5 Flash, Qwen3.5-Plus, Qwen3.6-Max, Qwen3.6 Plus, Qwen 3.7 Max, Qwen3.7 Plus, Qwen 3.8 27B, Qwen3.8 Max, Qwen3-VL-235B-A22B, Qwen 3 VL Plus