Clinical Benchmarks

Alibaba

Alibaba models with results here, each placed on its own board.

Models
22
Results
38
Boards
10 of 13

qwen.ai

Models

ModelKindReleasedBoards
Qwen3.8 MaxModel3 Aug 20263
Qwen3.7 PlusModel31 May 20261
Qwen3.6 PlusModel1 Apr 20264
Qwen3.5 397B A17BModel16 Feb 20264
Qwen3-14BModel29 Apr 20251
Qwen3-32BModel29 Apr 20251
Qwen3-4BModel29 Apr 20251
Lingshu-32BModelNot stated1
Lingshu-7BModelNot stated1
Qwen 3 Max ThinkingModelNot stated2
Qwen 3 VL PlusModelNot stated2
Qwen 3.5ModelNot stated1
Qwen 3.5 FlashModelNot stated2
Qwen 3.7 MaxModelNot stated2
Qwen 3.8 27BModelNot stated2
Qwen3-235B-A22B-Instruct-2507ModelNot stated1
Qwen3-VL-235B-A22BModelNot stated1
Qwen3.5-27BModelNot stated1
Qwen3.5-35B-A3BModelNot stated1
Qwen3.5-9BModelNot stated1
Qwen3.5-PlusModelNot stated1
Qwen3.6-MaxModelNot stated1

Results

One line per result. Dark tick: the board leader.

Clinical reasoning and knowledge

  1. MedXpertQA (MM)
    Alibaba's own Qwen3.8 launch table; protocol not stated
    80.4
    Rank 3 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
    Vendor-reported
    Measured Aug 2026
  2. MedXpertQA (MM)
    Alibaba's own Qwen3.7 Plus launch table; protocol not stated
    71.0
    Rank 10 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
    Vendor-reported
    Measured May 2026
  3. MedXpertQA (MM)
    self-reported in the Qwen3.5-397B-A17B model card
    70.0
    Rank 11 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
    Vendor-reported
  4. MedXpertQA (MM)
    Qwen-run comparison in the Qwen3.7-Plus launch post
    68.7
    Rank 12 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
    Vendor-reported
  5. MedXpertQA (MM)
    Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors
    47.6
    Rank 20 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
    Vendor-reported
    Measured Feb 2026
  6. 57.9
    Rank 5 of 8 here, 11 models on the boardLeader GPT-5.6 Sol 60.2
    Official leaderboard
    Measured Aug 2026

Documentation and coding

  1. MedScribe (Vals AI)
    model ID alibaba/qwen3.8-max; temperature=1; max_output_tokens=30000
    84.95
    Rank 30 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  2. MedScribe (Vals AI)
    model ID alibaba/qwen3.8-27b; temperature=1; top_p=0.95; max_output_tokens=30000; reasoning_effort=xhigh
    83.85
    Rank 38 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  3. MedScribe (Vals AI)
    model ID alibaba/qwen3.7-max; temperature=1; max_output_tokens=30000
    79.40
    Rank 59 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  4. MedScribe (Vals AI)
    model ID alibaba/qwen3-vl-plus-2025-09-23; temperature=1; max_output_tokens=30000
    77.13
    Rank 67 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  5. MedScribe (Vals AI)
    model ID alibaba/qwen3.6-plus; temperature=1; max_output_tokens=30000
    76.96
    Rank 69 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  6. MedScribe (Vals AI)
    model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=30000
    72.71
    Rank 83 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  7. MedScribe (Vals AI)
    model ID alibaba/qwen3.5-flash; temperature=1; max_output_tokens=30000
    70.62
    Rank 90 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  8. MedCode (Vals AI)
    model ID alibaba/qwen3.8-max; temperature=1; max_output_tokens=30000
    40.67
    Rank 58 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  9. MedCode (Vals AI)
    model ID alibaba/qwen3.7-max; temperature=1; max_output_tokens=30000
    38.75
    Rank 66 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  10. MedCode (Vals AI)
    model ID alibaba/qwen3.6-plus; temperature=1; max_output_tokens=30000
    36.89
    Rank 73 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  11. MedCode (Vals AI)
    model ID alibaba/qwen3.5-flash; temperature=1; max_output_tokens=30000
    33.00
    Rank 82 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  12. MedCode (Vals AI)
    model ID alibaba/qwen3-vl-plus-2025-09-23; temperature=1; max_output_tokens=30000
    31.65
    Rank 89 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  13. MedCode (Vals AI)
    model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=30000
    31.37
    Rank 90 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  14. MedCode (Vals AI)
    model ID alibaba/qwen3.8-27b; reasoning_effort=xhigh; temperature=1; top_p=0.95; max_output_tokens=30000
    28.70
    Rank 95 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

EHR and workflow agents

  1. 13.7
    Rank 18 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
    Official leaderboard
    Measured May 2026
  2. EHR-Complex
    headline evaluation
    0.620
    Rank 3 of 18Leader GPT-5.4 (high reasoning) 0.650
    Official leaderboard
    Measured Jun 2026
  3. EHR-Complex
    Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    0.530
    Rank 9 of 18Leader GPT-5.4 (high reasoning) 0.650
    Official leaderboard
    Measured Jun 2026
  4. EHR-Complex
    Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    0.360
    Rank 13 of 18Leader GPT-5.4 (high reasoning) 0.650
    Official leaderboard
    Measured Jun 2026
  5. EHR-Complex
    Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    0.300
    Rank 17 of 18Leader GPT-5.4 (high reasoning) 0.650
    Official leaderboard
    Measured Jun 2026
  6. EHR-Complex
    Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    0.160
    Rank 18 of 18Leader GPT-5.4 (high reasoning) 0.650
    Official leaderboard
    Measured Jun 2026
  7. CHI-Bench
    All Domains pass@1; hermes harness; Qwen3.6-Max preview endpoint
    16.4
    Rank 19 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026
  8. CHI-Bench
    All Domains pass@1; openai-agents harness; Qwen3.6-Max preview endpoint
    15.6
    Rank 21 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026
  9. CHI-Bench
    All Domains pass@1; deepagents harness; Qwen3.6-Max preview endpoint
    9.3
    Rank 33 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026
  10. CHI-Bench
    All Domains pass@1; openclaw harness; Qwen3.6-Max preview endpoint
    4.9
    Rank 38 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026
  11. HealthAdminBench
    screenshot-only, detailed prompting
    13.3
    Rank 5 of 7 here, 5 models on the boardLeader Claude Opus 4.6 (computer-use agent) 36.3
    Official leaderboard
    Measured Apr 2026

Safety

  1. MedPIC
    zero-shot; independent questions; exact option-set match
    72.4
    Rank 3 of 28Leader Gemini-3.1-Pro 80.7
    Independent run
    Measured Aug 2026
  2. MedPIC
    zero-shot; independent questions; exact option-set match
    69.4
    Rank 6 of 28Leader Gemini-3.1-Pro 80.7
    Independent run
    Measured Aug 2026
  3. MedPIC
    zero-shot; independent questions; exact option-set match
    68.5
    Rank 7 of 28Leader Gemini-3.1-Pro 80.7
    Independent run
    Measured Aug 2026
  4. MedPIC
    zero-shot; independent questions; exact option-set match
    63.8
    Rank 10 of 28Leader Gemini-3.1-Pro 80.7
    Independent run
    Measured Aug 2026
  5. MedPIC
    zero-shot; independent questions; exact option-set match
    59.5
    Rank 13 of 28Leader Gemini-3.1-Pro 80.7
    Independent run
    Measured Aug 2026
  6. MedPIC
    zero-shot; independent questions; exact option-set match
    43.3
    Rank 21 of 28Leader Gemini-3.1-Pro 80.7
    Independent run
    Measured Aug 2026
  7. 61.1
    Rank 14 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2
    Official leaderboard
    Measured Aug 2026