Qwen3.5 397B A17B
Released 16 Feb 2026256K context397B total / 17B activeapache-2.0Also written as Qwen3.5-397BCompare with other models
- Clinical Benchmarks Index
- 65.2rank 31 of 148; 65.2 × 1 = 65.2, from 4 of 10 boards
- Boards
- 4 of 13
- Results
- 4
- Latest measurement
- Aug 2026
Results
Each line is placed on its own board. Dark tick: the board leader.
Clinical reasoning and knowledge
- MedXpertQA (MM)self-reported in the Qwen3.5-397B-A17B model card70.0Rank 11 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
- 57.9Rank 5 of 8 here, 11 models on the boardLeader GPT-5.6 Sol 60.2
EHR and workflow agents
- EHR-Complexheadline evaluation0.620Rank 3 of 18Leader GPT-5.4 (high reasoning) 0.650
Safety
- 61.1Rank 14 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2
Sources
Open a line for the quote and page.
11MedXpertQA (MM) self-reported in the Qwen3.5-397B-A17B model card 70.0
Printed as 70.0Vendor-reportedConfiguration: self-reported in the Qwen3.5-397B-A17B model cardQwen/Qwen3.5-397B-A17B model card model card, Alibaba / Qwen, 16 Feb 2026. Benchmark Results > Vision Language > Medical VQA, row MedXpertQA-MM; columns GPT5.2 / Claude 4.5 Opus / Gemini-3 Pro / Qwen3-VL-235B-A22B / K2.5-1T-A32B / Qwen3.5-397B-A17BMedXpertQA-MM 73.3 63.6 76.0 47.6 65.3 70.0
Every result from this document5MAST (Medical AI Superintelligence Test) 57.9
Printed as 57.9%Official leaderboard, measured Aug 2026MAST: Medical AI Superintelligence Test leaderboard (General board) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')5 Qwen3.5 397B A17B Alibaba 57.9%
Every result from this document14First, Do NOHARM (v2) 61.1
Printed as 61.1%Official leaderboard, measured Aug 2026MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)13Qwen3.5 397B A17BOSSAlibaba 61.1%
Every result from this document3EHR-Complex headline evaluation 0.620
Printed as 0.62Official leaderboard, measured Jun 2026Configuration: headline evaluationEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. p. 6, Table 3 (Evaluation Results on the EHR-Complex Test Set), Avg. columnQwen3.5-397B 0.88 0.39 0.72 0.3 0.8 0.38 0.91 0.63 0.91 0.45 0.78 0.32 0.62
Every result from this document- Also printed in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301): 0.62, arXiv abs page (landing page for the PDF)
Other Alibaba models: Lingshu-32B, Lingshu-7B, Qwen3-14B, Qwen3-235B-A22B-Instruct-2507, Qwen3-32B, Qwen3-4B, Qwen 3.5, Qwen3.5-27B, Qwen3.5-35B-A3B, Qwen3.5-9B, Qwen 3.5 Flash, Qwen3.5-Plus, Qwen3.6-Max, Qwen3.6 Plus, Qwen 3.7 Max, Qwen3.7 Plus, Qwen 3.8 27B, Qwen3.8 Max, Qwen 3 Max Thinking, Qwen3-VL-235B-A22B, Qwen 3 VL Plus