Qwen3-VL-235B-A22B
- Clinical Benchmarks Index
- 24.0rank 124 of 148; 41.6 × 0.577 = 24.0, from 1 of 10 boards
- Boards
- 1 of 13
- Results
- 1
- Latest measurement
- Feb 2026
Results
Each line is placed on its own board. Dark tick: the board leader.
Clinical reasoning and knowledge
- MedXpertQA (MM)Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors47.6Rank 20 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
Sources
Open a line for the quote and page.
20MedXpertQA (MM) Qwen3.5 model card comparison; source-specific evaluation, not harmonized acros… 47.6
Printed as 47.6Vendor-reported, measured Feb 2026Configuration: Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendorsQwen/Qwen3.5-397B-A17B model card model card, Alibaba / Qwen, 16 Feb 2026. Benchmark Results > Vision Language > Medical VQA; MedXpertQA-MM row, Qwen3-VL-235B-A22B column.| GPT5.2 | Claude 4.5 Opus | Gemini-3 Pro | Qwen3-VL-235B-A22B | K2.5-1T-A32B | Qwen3.5-397B-A17B MedXpertQA-MM | 73.3 | 63.6 | 76.0 | 47.6 | 65.3 | 70.0
Every result from this document
Other Alibaba models: Lingshu-32B, Lingshu-7B, Qwen3-14B, Qwen3-235B-A22B-Instruct-2507, Qwen3-32B, Qwen3-4B, Qwen 3.5, Qwen3.5-27B, Qwen3.5-35B-A3B, Qwen3.5 397B A17B, Qwen3.5-9B, Qwen 3.5 Flash, Qwen3.5-Plus, Qwen3.6-Max, Qwen3.6 Plus, Qwen 3.7 Max, Qwen3.7 Plus, Qwen 3.8 27B, Qwen3.8 Max, Qwen 3 Max Thinking, Qwen 3 VL Plus