Alibaba
Alibaba models with results here, each placed on its own board.
- Models
- 22
- Results
- 38
- Boards
- 10 of 13
Models
| Model | Kind | Released | Boards |
|---|---|---|---|
| Qwen3.8 Max | Model | 3 Aug 2026 | 3 |
| Qwen3.7 Plus | Model | 31 May 2026 | 1 |
| Qwen3.6 Plus | Model | 1 Apr 2026 | 4 |
| Qwen3.5 397B A17B | Model | 16 Feb 2026 | 4 |
| Qwen3-14B | Model | 29 Apr 2025 | 1 |
| Qwen3-32B | Model | 29 Apr 2025 | 1 |
| Qwen3-4B | Model | 29 Apr 2025 | 1 |
| Lingshu-32B | Model | Not stated | 1 |
| Lingshu-7B | Model | Not stated | 1 |
| Qwen 3 Max Thinking | Model | Not stated | 2 |
| Qwen 3 VL Plus | Model | Not stated | 2 |
| Qwen 3.5 | Model | Not stated | 1 |
| Qwen 3.5 Flash | Model | Not stated | 2 |
| Qwen 3.7 Max | Model | Not stated | 2 |
| Qwen 3.8 27B | Model | Not stated | 2 |
| Qwen3-235B-A22B-Instruct-2507 | Model | Not stated | 1 |
| Qwen3-VL-235B-A22B | Model | Not stated | 1 |
| Qwen3.5-27B | Model | Not stated | 1 |
| Qwen3.5-35B-A3B | Model | Not stated | 1 |
| Qwen3.5-9B | Model | Not stated | 1 |
| Qwen3.5-Plus | Model | Not stated | 1 |
| Qwen3.6-Max | Model | Not stated | 1 |
Results
One line per result. Dark tick: the board leader.
Clinical reasoning and knowledge
- 80.4Rank 3 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
- 71.0Rank 10 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
- 70.0Rank 11 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
- 68.7Rank 12 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
- MedXpertQA (MM)Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors47.6Rank 20 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
- 57.9Rank 5 of 8 here, 11 models on the boardLeader GPT-5.6 Sol 60.2
Documentation and coding
- 84.95Rank 30 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID alibaba/qwen3.8-27b; temperature=1; top_p=0.95; max_output_tokens=30000; reasoning_effort=xhigh83.85Rank 38 of 105Leader Claude Opus 5.5 91.43
- 79.40Rank 59 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID alibaba/qwen3-vl-plus-2025-09-23; temperature=1; max_output_tokens=3000077.13Rank 67 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID alibaba/qwen3.6-plus; temperature=1; max_output_tokens=3000076.96Rank 69 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=3000072.71Rank 83 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID alibaba/qwen3.5-flash; temperature=1; max_output_tokens=3000070.62Rank 90 of 105Leader Claude Opus 5.5 91.43
- 40.67Rank 58 of 103Leader Claude Opus 5 63.57
- 38.75Rank 66 of 103Leader Claude Opus 5 63.57
- 36.89Rank 73 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID alibaba/qwen3.5-flash; temperature=1; max_output_tokens=3000033.00Rank 82 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID alibaba/qwen3-vl-plus-2025-09-23; temperature=1; max_output_tokens=3000031.65Rank 89 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=3000031.37Rank 90 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID alibaba/qwen3.8-27b; reasoning_effort=xhigh; temperature=1; top_p=0.95; max_output_tokens=3000028.70Rank 95 of 103Leader Claude Opus 5 63.57
EHR and workflow agents
- 13.7Rank 18 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
- 0.620Rank 3 of 18Leader GPT-5.4 (high reasoning) 0.650
- EHR-ComplexTable 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns0.530Rank 9 of 18Leader GPT-5.4 (high reasoning) 0.650
- EHR-ComplexTable 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns0.360Rank 13 of 18Leader GPT-5.4 (high reasoning) 0.650
- EHR-ComplexTable 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns0.300Rank 17 of 18Leader GPT-5.4 (high reasoning) 0.650
- EHR-ComplexTable 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns0.160Rank 18 of 18Leader GPT-5.4 (high reasoning) 0.650
- 16.4Rank 19 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
- CHI-BenchAll Domains pass@1; openai-agents harness; Qwen3.6-Max preview endpoint15.6Rank 21 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
- CHI-BenchAll Domains pass@1; deepagents harness; Qwen3.6-Max preview endpoint9.3Rank 33 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
- 4.9Rank 38 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
- 13.3Rank 5 of 7 here, 5 models on the boardLeader Claude Opus 4.6 (computer-use agent) 36.3
Safety
- 72.4Rank 3 of 28Leader Gemini-3.1-Pro 80.7
- 69.4Rank 6 of 28Leader Gemini-3.1-Pro 80.7
- 68.5Rank 7 of 28Leader Gemini-3.1-Pro 80.7
- 63.8Rank 10 of 28Leader Gemini-3.1-Pro 80.7
- 59.5Rank 13 of 28Leader Gemini-3.1-Pro 80.7
- 43.3Rank 21 of 28Leader Gemini-3.1-Pro 80.7
- 61.1Rank 14 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2