Qwen3-14B-SFT
- Clinical Benchmarks Index
- 34.2rank 108 of 148; 59.2 × 0.577 = 34.2, from 1 of 10 boards
- Boards
- 1 of 13
- Results
- 1
- Latest measurement
- Jun 2026
Results
Each line is placed on its own board. Dark tick: the board leader.
EHR and workflow agents
- EHR-ComplexTable 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns; fine-tuned on trajectories from EHR-Complex training set0.450Rank 12 of 18Leader GPT-5.4 (high reasoning) 0.650
Sources
Open a line for the quote and page.
12EHR-Complex Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.450
Printed as 0.45Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns; fine-tuned on trajectories from EHR-Complex training setEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, Qwen3-14B-SFT row, final Avg. column (PDF p. 6).Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. Qwen3-14B-SFT | 0.69 | 0.22 | 0.43 | 0.15 | 0.6 | 0.19 | 0.84 | 0.51 | 0.74 | 0.22 | 0.63 | 0.18 | 0.45
Every result from this document
Other Ant Group models: Ling 3.0 Flash, Ling 3.0 Flash Fin, Qwen3-32B-SFT