Clinical Benchmarks

Qwen3-4B

Released 29 Apr 202532K context4.0BopenCompare with other models

Clinical Benchmarks Index
0.0rank 148 of 148; 0.0 × 0.577 = 0.0, from 1 of 10 boards
Boards
1 of 13
Results
1
Latest measurement
Jun 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

EHR and workflow agents

  1. EHR-Complex
    Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    0.160
    Rank 18 of 18Leader GPT-5.4 (high reasoning) 0.650
    Official leaderboard
    Measured Jun 2026

Sources

Open a line for the quote and page.

  1. 18EHR-Complex Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.160
    Printed as 0.16Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, Qwen3-4B row, final Avg. column (PDF p. 6).
    Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. Qwen3-4B | 0.24 | 0.07 | 0.28 | 0.11 | 0.18 | 0.01 | 0.48 | 0.12 | 0.17 | 0.06 | 0.16 | 0.03 | 0.16
    Every result from this document

Other Alibaba models: Lingshu-32B, Lingshu-7B, Qwen3-14B, Qwen3-235B-A22B-Instruct-2507, Qwen3-32B, Qwen 3.5, Qwen3.5-27B, Qwen3.5-35B-A3B, Qwen3.5 397B A17B, Qwen3.5-9B, Qwen 3.5 Flash, Qwen3.5-Plus, Qwen3.6-Max, Qwen3.6 Plus, Qwen 3.7 Max, Qwen3.7 Plus, Qwen 3.8 27B, Qwen3.8 Max, Qwen 3 Max Thinking, Qwen3-VL-235B-A22B, Qwen 3 VL Plus