GPT-4.1 mini
Released 14 Apr 2025$0.4 input, $1.6 output per million tokens1M contextproprietaryCompare with other models
- Clinical Benchmarks Index
- 38.8rank 102 of 148; 67.3 × 0.577 = 38.8, from 1 of 10 boards
- Boards
- 1 of 13
- Results
- 1
- Latest measurement
- Jun 2026
Results
Each line is placed on its own board. Dark tick: the board leader.
EHR and workflow agents
- EHR-ComplexTable 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns0.490Rank 10 of 18Leader GPT-5.4 (high reasoning) 0.650
Sources
Open a line for the quote and page.
10EHR-Complex Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50… 0.490
Printed as 0.49Official leaderboard, measured Jun 2026Configuration: Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turnsEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) paper, Zhejiang University / Ant Group (Qiao, Liu, Chu et al.), 22 Jun 2026. arXiv HTML 2606.23301v1, Table 3, GPT-4.1 mini row, final Avg. column (PDF p. 6).Model | Demographics | Vitals | Medications | Cost | Labs | Diagnoses | Avg. GPT-4.1 mini | 0.75 | 0.27 | 0.6 | 0.22 | 0.67 | 0.21 | 0.88 | 0.6 | 0.65 | 0.3 | 0.59 | 0.17 | 0.49
Every result from this document
Other OpenAI models: GPT-4.1, GPT-4o, GPT-5, GPT-5.1, GPT-5.2, GPT-5.3-Codex, GPT-5.4, GPT-5.4 mini, GPT-5.4 nano, GPT-5.5, GPT-5.5 Instant, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, GPT-5 mini, GPT-5 nano, GPT-6.1 Sol, GPT-6 Astra, GPT-6 Luna, GPT-6 Sol, GPT OSS 120B, GPT OSS 20B, o3, o4-mini