Clinical Benchmarks

Qwen3-32B-SFT

Compare with other models

Clinical Benchmarks Index
45.9rank 74 of 148; 79.6 × 0.577 = 45.9, from 1 of 10 boards
Boards
1 of 13
Results
1
Latest measurement
Jun 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

EHR and workflow agents

  1. EHR-Complex
    Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns; fine-tuned on trajectories from EHR-Complex training set
    0.550
    Rank 8 of 18Leader GPT-5.4 (high reasoning) 0.650
    Official leaderboard
    Measured Jun 2026

Sources

Open a line for the quote and page.

  1. 8EHR-Complex Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.550
    Printed as 0.55Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns; fine-tuned on trajectories from EHR-Complex training set
    EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, Qwen3-32B-SFT row, final Avg. column (PDF p. 6).
    Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. Qwen3-32B-SFT | 0.83 | 0.36 | 0.55 | 0.29 | 0.72 | 0.3 | 0.86 | 0.55 | 0.8 | 0.34 | 0.69 | 0.26 | 0.55
    Every result from this document

Other Ant Group models: Ling 3.0 Flash, Ling 3.0 Flash Fin, Qwen3-14B-SFT