Clinical Benchmarks

EHR-Complex

Agentic clinical reasoning over MIMIC-IV records through SQL and Python, at patient and population level.

academic teamOfficial pagePaper
More about this boardLess

Agentic clinical reasoning over MIMIC-IV EHR databases via SQL and Python across six clinical intents, at patient and population level with temporal evidence paths.

Scores come from the paper, which reports both a headline 12-model evaluation and human-validated configurations; rows here mix the two, labeled in the config column. Consistency drops below 50 percent at Pass^4 for nearly every model.

Published by Academic team (Qiao et al., Ant Group-affiliated; arXiv 2606.23301), released Jun 2026. ~52,000 tasks (3,915-task test set) over 365K patients, 31 tables, 500M+ records. Exact-match accuracy, 0-1, higher better.

Headline metric
exact-match success, Avg. over 12 intent x scope columns
Tables
Table 3 (p. 6) headline; Table 10 (p. 15) strong commercial models
Paper date
22 Jun 2026
Paper version
v1
Rows
18
Models
17 of 18on the board
Labs
7
Last measured
Jun 2026
Leader
0.650GPT-5.4 (high reasoning)

Ranking

Exact-match accuracy, 0-1, higher better

  • Anthropic
  • Google
  • OpenAI
  • Alibaba
  • Moonshot AI
  • Other labs
All official leaderboard
  1. 1GPT-5.4 (high reasoning)average over 12 intent columns0.650
  2. 2Gemini 3.1 Provalidation configuration0.630
  3. 3Kimi-K2.5headline 12-model evaluation0.620
  4. 3Qwen3.5-397Bheadline evaluation0.620
  5. 5DeepSeek-V3.2-ExpTable 30.590
  6. 6GPT-5.4 (low reasoning)validation configuration0.580
  7. 7DeepSeek-V3.1Table 30.560
  8. 8Qwen3-32B-SFTTable 30.550
  9. 9Qwen3-235BTable 30.530
  10. 10GPT-4.1 miniTable 30.490
  11. 11GPT-4.1Table 30.470
  12. 12Qwen3-14B-SFTTable 30.450
  13. 13Claude Sonnet 4.6validation configuration0.360
  14. 13Qwen3-32BTable 30.360
  15. 15Gemini 2.5 ProTable 30.310
  16. 15GPT-4oTable 30.310
  17. 17Qwen3-14BTable 30.300
  18. 18Qwen3-4BTable 30.160

18 of 18 rows

Rows and sources

Open a row for the quote, the page and the document.

  1. 1GPT-5.4 (high reasoning) average over 12 intent columns; run as human-validation configuration, not in t… 0.650
    Printed as 0.65Official leaderboard, measured Jun 2026Configuration: average over 12 intent columns; run as human-validation configuration, not in the headline 12-model table
    EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. p. 15, Table 10 (Strong commercial model results), Avg. column
    GPT-5.4 (high) 0.85 0.36 0.78 0.34 0.8 0.37 0.93 0.76 0.91 0.49 0.84 0.36 0.65
    Every result from this document
  2. 2Gemini 3.1 Pro validation configuration 0.630
    Printed as 0.63Official leaderboard, measured Jun 2026Configuration: validation configuration
    EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. p. 15, Table 10 (Strong commercial model results), Avg. column
    Gemini 3.1 Pro 0.87 0.32 0.67 0.38 0.84 0.37 0.92 0.63 0.92 0.44 0.83 0.34 0.63
    Every result from this document
  3. 3Kimi-K2.5 headline 12-model evaluation, top open-weight 0.620
    Printed as 0.62Official leaderboard, measured Jun 2026Configuration: headline 12-model evaluation, top open-weight
    EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. p. 6, Table 3 (Evaluation Results on the EHR-Complex Test Set), Avg. column
    Kimi-K2.5 0.89 0.34 0.8 0.27 0.73 0.35 0.92 0.72 0.89 0.42 0.84 0.29 0.62
    Every result from this document
  4. 3Qwen3.5-397B headline evaluation 0.620
    Printed as 0.62Official leaderboard, measured Jun 2026Configuration: headline evaluation
    EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. p. 6, Table 3 (Evaluation Results on the EHR-Complex Test Set), Avg. column
    Qwen3.5-397B 0.88 0.39 0.72 0.3 0.8 0.38 0.91 0.63 0.91 0.45 0.78 0.32 0.62
    Every result from this document
  5. 5DeepSeek-V3.2-Exp Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.590
    Printed as 0.59Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, DeepSeek-V3.2-Exp row, final Avg. column (PDF p. 6).
    Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. DeepSeek-V3.2-Exp | 0.83 | 0.36 | 0.75 | 0.27 | 0.7 | 0.3 | 0.88 | 0.65 | 0.89 | 0.44 | 0.77 | 0.26 | 0.59
    Every result from this document
  6. 6GPT-5.4 (low reasoning) validation configuration 0.580
    Printed as 0.58Official leaderboard, measured Jun 2026Configuration: validation configuration
    EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. p. 15, Table 10 (Strong commercial model results), Avg. column
    GPT-5.4 (low) 0.81 0.32 0.65 0.26 0.73 0.31 0.88 0.66 0.86 0.43 0.8 0.29 0.58
    Every result from this document
  7. 7DeepSeek-V3.1 Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.560
    Printed as 0.56Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, DeepSeek-V3.1 row, final Avg. column (PDF p. 6).
    Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. DeepSeek-V3.1 | 0.81 | 0.33 | 0.7 | 0.27 | 0.65 | 0.24 | 0.89 | 0.68 | 0.82 | 0.39 | 0.72 | 0.21 | 0.56
    Every result from this document
  8. 8Qwen3-32B-SFT Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.550
    Printed as 0.55Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns; fine-tuned on trajectories from EHR-Complex training set
    EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, Qwen3-32B-SFT row, final Avg. column (PDF p. 6).
    Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. Qwen3-32B-SFT | 0.83 | 0.36 | 0.55 | 0.29 | 0.72 | 0.3 | 0.86 | 0.55 | 0.8 | 0.34 | 0.69 | 0.26 | 0.55
    Every result from this document
  9. 9Qwen3-235B Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.530
    Printed as 0.53Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, Qwen3-235B row, final Avg. column (PDF p. 6).
    Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. Qwen3-235B | 0.75 | 0.27 | 0.65 | 0.26 | 0.72 | 0.25 | 0.91 | 0.55 | 0.8 | 0.36 | 0.7 | 0.21 | 0.53
    Every result from this document
  10. 10GPT-4.1 mini Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.490
    Printed as 0.49Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, GPT-4.1 mini row, final Avg. column (PDF p. 6).
    Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. GPT-4.1 mini | 0.75 | 0.27 | 0.6 | 0.22 | 0.67 | 0.21 | 0.88 | 0.6 | 0.65 | 0.3 | 0.59 | 0.17 | 0.49
    Every result from this document
  11. 11GPT-4.1 Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.470
    Printed as 0.47Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, GPT-4.1 row, final Avg. column (PDF p. 6).
    Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. GPT-4.1 | 0.76 | 0.21 | 0.55 | 0.16 | 0.59 | 0.17 | 0.89 | 0.5 | 0.72 | 0.26 | 0.63 | 0.18 | 0.47
    Every result from this document
  12. 12Qwen3-14B-SFT Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.450
    Printed as 0.45Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns; fine-tuned on trajectories from EHR-Complex training set
    EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, Qwen3-14B-SFT row, final Avg. column (PDF p. 6).
    Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. Qwen3-14B-SFT | 0.69 | 0.22 | 0.43 | 0.15 | 0.6 | 0.19 | 0.84 | 0.51 | 0.74 | 0.22 | 0.63 | 0.18 | 0.45
    Every result from this document
  13. 13Claude Sonnet 4.6 validation configuration 0.360
    Printed as 0.36Official leaderboard, measured Jun 2026Configuration: validation configuration
    EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. p. 15, Table 10 (Strong commercial model results), Avg. column
    Claude Sonnet 4.6 0.59 0.21 0.45 0.12 0.45 0.12 0.83 0.2 0.49 0.33 0.4 0.11 0.36
    Every result from this document
  14. 13Qwen3-32B Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.360
    Printed as 0.36Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, Qwen3-32B row, final Avg. column (PDF p. 6).
    Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. Qwen3-32B | 0.51 | 0.2 | 0.3 | 0.24 | 0.47 | 0.2 | 0.73 | 0.35 | 0.5 | 0.21 | 0.42 | 0.15 | 0.36
    Every result from this document
  15. 15Gemini 2.5 Pro Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.310
    Printed as 0.31Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, Gemini 2.5 Pro row, final Avg. column (PDF p. 6).
    Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. Gemini 2.5 Pro | 0.44 | 0.19 | 0.18 | 0.2 | 0.43 | 0.15 | 0.59 | 0.38 | 0.38 | 0.27 | 0.34 | 0.15 | 0.31
    Every result from this document
  16. 15GPT-4o Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.310
    Printed as 0.31Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, GPT-4o row, final Avg. column (PDF p. 6).
    Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. GPT-4o | 0.52 | 0.1 | 0.45 | 0.26 | 0.33 | 0.07 | 0.53 | 0.33 | 0.46 | 0.19 | 0.38 | 0.08 | 0.31
    Every result from this document
  17. 17Qwen3-14B Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.300
    Printed as 0.3Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, Qwen3-14B row, final Avg. column (PDF p. 6).
    Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. Qwen3-14B | 0.44 | 0.15 | 0.21 | 0.18 | 0.43 | 0.15 | 0.63 | 0.33 | 0.4 | 0.19 | 0.39 | 0.1 | 0.3
    Every result from this document
  18. 18Qwen3-4B Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.160
    Printed as 0.16Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, Qwen3-4B row, final Avg. column (PDF p. 6).
    Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. Qwen3-4B | 0.24 | 0.07 | 0.28 | 0.11 | 0.18 | 0.01 | 0.48 | 0.12 | 0.17 | 0.06 | 0.16 | 0.03 | 0.16
    Every result from this document

Documents

2