EHR-Complex
Agentic clinical reasoning over MIMIC-IV records through SQL and Python, at patient and population level.
More about this boardLess
Agentic clinical reasoning over MIMIC-IV EHR databases via SQL and Python across six clinical intents, at patient and population level with temporal evidence paths.
Scores come from the paper, which reports both a headline 12-model evaluation and human-validated configurations; rows here mix the two, labeled in the config column. Consistency drops below 50 percent at Pass^4 for nearly every model.
Published by Academic team (Qiao et al., Ant Group-affiliated; arXiv 2606.23301), released Jun 2026. ~52,000 tasks (3,915-task test set) over 365K patients, 31 tables, 500M+ records. Exact-match accuracy, 0-1, higher better.
- Headline metric
- exact-match success, Avg. over 12 intent x scope columns
- Tables
- Table 3 (p. 6) headline; Table 10 (p. 15) strong commercial models
- Paper date
- 22 Jun 2026
- Paper version
- v1
- Rows
- 18
- Models
- 17 of 18on the board
- Labs
- 7
- Last measured
- Jun 2026
- Leader
- 0.650GPT-5.4 (high reasoning)
Ranking
Exact-match accuracy, 0-1, higher better
- Anthropic
- OpenAI
- Alibaba
- Moonshot AI
- Other labs
18 of 18 rows
Rows and sources
Open a row for the quote, the page and the document.
1GPT-5.4 (high reasoning) average over 12 intent columns; run as human-validation configuration, not in t… 0.650
Printed as 0.65Official leaderboard, measured Jun 2026Configuration: average over 12 intent columns; run as human-validation configuration, not in the headline 12-model tableEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. p. 15, Table 10 (Strong commercial model results), Avg. columnGPT-5.4 (high) 0.85 0.36 0.78 0.34 0.8 0.37 0.93 0.76 0.91 0.49 0.84 0.36 0.65
Every result from this document- Also printed in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301): 0.65, arXiv abs page (landing page for the PDF)
2Gemini 3.1 Pro validation configuration 0.630
Printed as 0.63Official leaderboard, measured Jun 2026Configuration: validation configurationEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. p. 15, Table 10 (Strong commercial model results), Avg. columnGemini 3.1 Pro 0.87 0.32 0.67 0.38 0.84 0.37 0.92 0.63 0.92 0.44 0.83 0.34 0.63
Every result from this document- Also printed in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301): 0.63, arXiv abs page (landing page for the PDF)
3Kimi-K2.5 headline 12-model evaluation, top open-weight 0.620
Printed as 0.62Official leaderboard, measured Jun 2026Configuration: headline 12-model evaluation, top open-weightEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. p. 6, Table 3 (Evaluation Results on the EHR-Complex Test Set), Avg. columnKimi-K2.5 0.89 0.34 0.8 0.27 0.73 0.35 0.92 0.72 0.89 0.42 0.84 0.29 0.62
Every result from this document- Also printed in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301): 0.62, arXiv abs page (landing page for the PDF)
3Qwen3.5-397B headline evaluation 0.620
Printed as 0.62Official leaderboard, measured Jun 2026Configuration: headline evaluationEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. p. 6, Table 3 (Evaluation Results on the EHR-Complex Test Set), Avg. columnQwen3.5-397B 0.88 0.39 0.72 0.3 0.8 0.38 0.91 0.63 0.91 0.45 0.78 0.32 0.62
Every result from this document- Also printed in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301): 0.62, arXiv abs page (landing page for the PDF)
5DeepSeek-V3.2-Exp Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.590
Printed as 0.59Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turnsEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, DeepSeek-V3.2-Exp row, final Avg. column (PDF p. 6).Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. DeepSeek-V3.2-Exp | 0.83 | 0.36 | 0.75 | 0.27 | 0.7 | 0.3 | 0.88 | 0.65 | 0.89 | 0.44 | 0.77 | 0.26 | 0.59
Every result from this document6GPT-5.4 (low reasoning) validation configuration 0.580
Printed as 0.58Official leaderboard, measured Jun 2026Configuration: validation configurationEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. p. 15, Table 10 (Strong commercial model results), Avg. columnGPT-5.4 (low) 0.81 0.32 0.65 0.26 0.73 0.31 0.88 0.66 0.86 0.43 0.8 0.29 0.58
Every result from this document- Also printed in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301): 0.58, arXiv abs page (landing page for the PDF)
7DeepSeek-V3.1 Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.560
Printed as 0.56Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turnsEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, DeepSeek-V3.1 row, final Avg. column (PDF p. 6).Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. DeepSeek-V3.1 | 0.81 | 0.33 | 0.7 | 0.27 | 0.65 | 0.24 | 0.89 | 0.68 | 0.82 | 0.39 | 0.72 | 0.21 | 0.56
Every result from this document8Qwen3-32B-SFT Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.550
Printed as 0.55Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns; fine-tuned on trajectories from EHR-Complex training setEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, Qwen3-32B-SFT row, final Avg. column (PDF p. 6).Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. Qwen3-32B-SFT | 0.83 | 0.36 | 0.55 | 0.29 | 0.72 | 0.3 | 0.86 | 0.55 | 0.8 | 0.34 | 0.69 | 0.26 | 0.55
Every result from this document9Qwen3-235B Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.530
Printed as 0.53Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turnsEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, Qwen3-235B row, final Avg. column (PDF p. 6).Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. Qwen3-235B | 0.75 | 0.27 | 0.65 | 0.26 | 0.72 | 0.25 | 0.91 | 0.55 | 0.8 | 0.36 | 0.7 | 0.21 | 0.53
Every result from this document10GPT-4.1 mini Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.490
Printed as 0.49Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turnsEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, GPT-4.1 mini row, final Avg. column (PDF p. 6).Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. GPT-4.1 mini | 0.75 | 0.27 | 0.6 | 0.22 | 0.67 | 0.21 | 0.88 | 0.6 | 0.65 | 0.3 | 0.59 | 0.17 | 0.49
Every result from this document11GPT-4.1 Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.470
Printed as 0.47Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turnsEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, GPT-4.1 row, final Avg. column (PDF p. 6).Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. GPT-4.1 | 0.76 | 0.21 | 0.55 | 0.16 | 0.59 | 0.17 | 0.89 | 0.5 | 0.72 | 0.26 | 0.63 | 0.18 | 0.47
Every result from this document12Qwen3-14B-SFT Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.450
Printed as 0.45Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns; fine-tuned on trajectories from EHR-Complex training setEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, Qwen3-14B-SFT row, final Avg. column (PDF p. 6).Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. Qwen3-14B-SFT | 0.69 | 0.22 | 0.43 | 0.15 | 0.6 | 0.19 | 0.84 | 0.51 | 0.74 | 0.22 | 0.63 | 0.18 | 0.45
Every result from this document13Claude Sonnet 4.6 validation configuration 0.360
Printed as 0.36Official leaderboard, measured Jun 2026Configuration: validation configurationEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. p. 15, Table 10 (Strong commercial model results), Avg. columnClaude Sonnet 4.6 0.59 0.21 0.45 0.12 0.45 0.12 0.83 0.2 0.49 0.33 0.4 0.11 0.36
Every result from this document- Also printed in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301): 0.36, arXiv abs page (landing page for the PDF)
13Qwen3-32B Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.360
Printed as 0.36Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turnsEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, Qwen3-32B row, final Avg. column (PDF p. 6).Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. Qwen3-32B | 0.51 | 0.2 | 0.3 | 0.24 | 0.47 | 0.2 | 0.73 | 0.35 | 0.5 | 0.21 | 0.42 | 0.15 | 0.36
Every result from this document15Gemini 2.5 Pro Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.310
Printed as 0.31Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turnsEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, Gemini 2.5 Pro row, final Avg. column (PDF p. 6).Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. Gemini 2.5 Pro | 0.44 | 0.19 | 0.18 | 0.2 | 0.43 | 0.15 | 0.59 | 0.38 | 0.38 | 0.27 | 0.34 | 0.15 | 0.31
Every result from this document15GPT-4o Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.310
Printed as 0.31Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turnsEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, GPT-4o row, final Avg. column (PDF p. 6).Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. GPT-4o | 0.52 | 0.1 | 0.45 | 0.26 | 0.33 | 0.07 | 0.53 | 0.33 | 0.46 | 0.19 | 0.38 | 0.08 | 0.31
Every result from this document17Qwen3-14B Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.300
Printed as 0.3Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turnsEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, Qwen3-14B row, final Avg. column (PDF p. 6).Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. Qwen3-14B | 0.44 | 0.15 | 0.21 | 0.18 | 0.43 | 0.15 | 0.63 | 0.33 | 0.4 | 0.19 | 0.39 | 0.1 | 0.3
Every result from this document18Qwen3-4B Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.160
Printed as 0.16Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turnsEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, Qwen3-4B row, final Avg. column (PDF p. 6).Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. Qwen3-4B | 0.24 | 0.07 | 0.28 | 0.11 | 0.18 | 0.01 | 0.48 | 0.12 | 0.17 | 0.06 | 0.16 | 0.03 | 0.16
Every result from this document
Documents
2
- EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301)paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026Results it supports
- EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF)paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026Results it supports