Clinical Benchmarks

EHR-Complex: current results

Academic team (Qiao et al., Ant Group-affiliated; arXiv 2606.23301) · ~52,000 tasks (3,915-task test set) over 365K patients, 31 tables, 500M+ records · index updated August 16, 2026

GPT-5.4 (high reasoning) holds the top current result on EHR-Complex, 0.65 as of 2026-06, per EHR-Complex paper. Agentic clinical reasoning over MIMIC-IV EHR databases via SQL and Python across six clinical intents, at patient and population level with temporal evidence paths.

Current results

Result detail

#modelscoreas of
1OpenAI logoGPT-5.4 (high reasoning) OpenAI
average over 12 intent columns; run as human-validation configuration, not in the headline 12-model table
0.652026-06
2Google logoGemini 3.1 Pro Google
validation configuration
0.632026-06
3Moonshot AI logoKimi-K2.5 Moonshot AI
headline 12-model evaluation, top open-weight
0.622026-06
4AQwen3.5-397B Alibaba
headline evaluation
0.622026-06
5OpenAI logoGPT-5.4 (low reasoning) OpenAI
validation configuration
0.582026-06
6Anthropic logoClaude Sonnet 4.6 Anthropic
validation configuration
0.362026-06

Scores appear exactly as EHR-Complex paper publishes them (independently run). Scores come from the paper, which reports both a headline 12-model evaluation and human-validated configurations; rows here mix the two, labeled in the config column. Consistency drops below 50 percent at Pass^4 for nearly every model.

About the benchmark

publisherAcademic team (Qiao et al., Ant Group-affiliated; arXiv 2606.23301)
categoryagentic and workflow benchmarks
released2026-06
size~52,000 tasks (3,915-task test set) over 365K patients, 31 tables, 500M+ records
scaleexact-match accuracy, 0-1, higher better
result basisindependently run
sourceEHR-Complex paper
last frontier result2026-06

What is EHR-Complex?

EHR-Complex is a agentic and workflow benchmark from academic team, released 2026-06: ~52,000 tasks (3,915-task test set) over 365K patients, 31 tables, 500M+ records, scored on a exact-match accuracy scale. Agentic clinical reasoning over MIMIC-IV EHR databases via SQL and Python across six clinical intents, at patient and population level with temporal evidence paths.

Which model leads EHR-Complex?

GPT-5.4 (high reasoning) (OpenAI) holds the top current result on EHR-Complex at 0.65, per EHR-Complex paper, as of 2026-06.

Where do the EHR-Complex numbers come from?

From EHR-Complex paper (independently run). Scores come from the paper, which reports both a headline 12-model evaluation and human-validated configurations; rows here mix the two, labeled in the config column. Consistency drops below 50 percent at Pass^4 for nearly every model.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.