Clinical Benchmarks

Qwen3.5 397B A17B

Released 16 Feb 2026256K context397B total / 17B activeapache-2.0Also written as Qwen3.5-397BCompare with other models

Clinical Benchmarks Index
65.2rank 31 of 148; 65.2 × 1 = 65.2, from 4 of 10 boards
Boards
4 of 13
Results
4
Latest measurement
Aug 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

Clinical reasoning and knowledge

  1. MedXpertQA (MM)
    self-reported in the Qwen3.5-397B-A17B model card
    70.0
    Rank 11 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
    Vendor-reported
  2. 57.9
    Rank 5 of 8 here, 11 models on the boardLeader GPT-5.6 Sol 60.2
    Official leaderboard
    Measured Aug 2026

EHR and workflow agents

  1. EHR-Complex
    headline evaluation
    0.620
    Rank 3 of 18Leader GPT-5.4 (high reasoning) 0.650
    Official leaderboard
    Measured Jun 2026

Safety

  1. 61.1
    Rank 14 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2
    Official leaderboard
    Measured Aug 2026

Sources

Open a line for the quote and page.

  1. 11MedXpertQA (MM) self-reported in the Qwen3.5-397B-A17B model card 70.0
    Printed as 70.0Vendor-reportedConfiguration: self-reported in the Qwen3.5-397B-A17B model card
    Qwen/Qwen3.5-397B-A17B model card model card, Alibaba / Qwen, 16 Feb 2026. Benchmark Results > Vision Language > Medical VQA, row MedXpertQA-MM; columns GPT5.2 / Claude 4.5 Opus / Gemini-3 Pro / Qwen3-VL-235B-A22B / K2.5-1T-A32B / Qwen3.5-397B-A17B
    MedXpertQA-MM 73.3 63.6 76.0 47.6 65.3 70.0
    Every result from this document
  2. 5MAST (Medical AI Superintelligence Test) 57.9
    Printed as 57.9%Official leaderboard, measured Aug 2026
    MAST: Medical AI Superintelligence Test leaderboard (General board) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
    5 Qwen3.5 397B A17B Alibaba 57.9%
    Every result from this document
  3. 14First, Do NOHARM (v2) 61.1
    Printed as 61.1%Official leaderboard, measured Aug 2026
    MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
    13Qwen3.5 397B A17BOSSAlibaba 61.1%
    Every result from this document
  4. 3EHR-Complex headline evaluation 0.620
    Printed as 0.62Official leaderboard, measured Jun 2026Configuration: headline evaluation
    EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. p. 6, Table 3 (Evaluation Results on the EHR-Complex Test Set), Avg. column
    Qwen3.5-397B 0.88 0.39 0.72 0.3 0.8 0.38 0.91 0.63 0.91 0.45 0.78 0.32 0.62
    Every result from this document

Other Alibaba models: Lingshu-32B, Lingshu-7B, Qwen3-14B, Qwen3-235B-A22B-Instruct-2507, Qwen3-32B, Qwen3-4B, Qwen 3.5, Qwen3.5-27B, Qwen3.5-35B-A3B, Qwen3.5-9B, Qwen 3.5 Flash, Qwen3.5-Plus, Qwen3.6-Max, Qwen3.6 Plus, Qwen 3.7 Max, Qwen3.7 Plus, Qwen 3.8 27B, Qwen3.8 Max, Qwen 3 Max Thinking, Qwen3-VL-235B-A22B, Qwen 3 VL Plus