EHR-Complex: current results
Academic team (Qiao et al., Ant Group-affiliated; arXiv 2606.23301) · ~52,000 tasks (3,915-task test set) over 365K patients, 31 tables, 500M+ records · index updated August 16, 2026
GPT-5.4 (high reasoning) holds the top current result on EHR-Complex, 0.65 as of 2026-06, per EHR-Complex paper. Agentic clinical reasoning over MIMIC-IV EHR databases via SQL and Python across six clinical intents, at patient and population level with temporal evidence paths.
Current results
- 1
GPT-5.4 (high reasoning)0.65
- 2
Gemini 3.1 Pro0.63
- 3
Kimi-K2.50.62
- 4AQwen3.5-397B0.62
- 5
GPT-5.4 (low reasoning)0.58
- 6
Claude Sonnet 4.60.36
Result detail
| # | model | score | as of | |
|---|---|---|---|---|
| 1 | GPT-5.4 (high reasoning) OpenAI average over 12 intent columns; run as human-validation configuration, not in the headline 12-model table | 0.65 | 2026-06 | |
| 2 | Gemini 3.1 Pro Google validation configuration | 0.63 | 2026-06 | |
| 3 | Kimi-K2.5 Moonshot AI headline 12-model evaluation, top open-weight | 0.62 | 2026-06 | |
| 4 | A | Qwen3.5-397B Alibaba headline evaluation | 0.62 | 2026-06 |
| 5 | GPT-5.4 (low reasoning) OpenAI validation configuration | 0.58 | 2026-06 | |
| 6 | Claude Sonnet 4.6 Anthropic validation configuration | 0.36 | 2026-06 | |
Scores appear exactly as EHR-Complex paper publishes them (independently run). Scores come from the paper, which reports both a headline 12-model evaluation and human-validated configurations; rows here mix the two, labeled in the config column. Consistency drops below 50 percent at Pass^4 for nearly every model.
About the benchmark
| publisher | Academic team (Qiao et al., Ant Group-affiliated; arXiv 2606.23301) |
|---|---|
| category | agentic and workflow benchmarks |
| released | 2026-06 |
| size | ~52,000 tasks (3,915-task test set) over 365K patients, 31 tables, 500M+ records |
| scale | exact-match accuracy, 0-1, higher better |
| result basis | independently run |
| source | EHR-Complex paper |
| last frontier result | 2026-06 |
What is EHR-Complex?
EHR-Complex is a agentic and workflow benchmark from academic team, released 2026-06: ~52,000 tasks (3,915-task test set) over 365K patients, 31 tables, 500M+ records, scored on a exact-match accuracy scale. Agentic clinical reasoning over MIMIC-IV EHR databases via SQL and Python across six clinical intents, at patient and population level with temporal evidence paths.
Which model leads EHR-Complex?
GPT-5.4 (high reasoning) (OpenAI) holds the top current result on EHR-Complex at 0.65, per EHR-Complex paper, as of 2026-06.
Where do the EHR-Complex numbers come from?
From EHR-Complex paper (independently run). Scores come from the paper, which reports both a headline 12-model evaluation and human-validated configurations; rows here mix the two, labeled in the config column. Consistency drops below 50 percent at Pass^4 for nearly every model.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.