PhysicianBench
Agents carrying out long-horizon physician workflows inside real EHR systems, verified by execution against those systems.
More about this boardLess
Agents carrying out long-horizon physician workflows inside real EHR systems, verified by execution against those systems.
LLM agents on long-horizon composite physician workflows inside real EHR environments, with execution-grounded verification against actual EHR systems via standard commercial APIs.
Scores come from the paper; there is no standalone public leaderboard, and the results also feed the ARISE MAST composite. The lower half of the table falls steeply, with several agents near 1 percent on Pass^3 consistency.
Published by Academic team (Ruoqi Liu, Imran Q. Mohiuddin et al., arXiv 2605.02240); also a MAST component, released May 2026. 100 real-world clinical tasks, 21 specialties, 670 structured checkpoints (~27 tool calls per task). Pass@1 success rate %, higher better (3 independent runs; Pass^3 also reported).
- Table
- Table 2, p. 8
- Headline metric
- Pass@1 (%), mean ± sd over 3 runs
- Paper date
- 4 May 2026
- Paper version
- v1
- Models in the paper's table
- 12
- Rows
- 21
- Models
- 17 of 12on the board
- Labs
- 9
- Last measured
- Sep 2026
- Leader
- 68.4Claude Opus 5.5 (max)
Ranking
Pass@1 success rate %, higher better (3 independent runs; Pass^3 also reported)
- Anthropic
- OpenAI
- Alibaba
- Moonshot AI
- Other labs
20 of 21 rows
Rows and sources
Open a row for the quote, the page and the document.
1Claude Opus 5.5 (max) Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety… 68.4
Printed as 68.4%Vendor-reported, measured Sep 2026Configuration: Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabledClaude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. pp. 138–139, Section 8.15.3; PhysicianBench pass@1, Claude Opus 5.5 (max).PhysicianBench is a public benchmark of 100 physician tasks carried out in an EHR. At max effort, Claude Sonnet 5.5 passes 63.2% of attempts, well above Claude Sonnet 5 (37.4%), level with Claude Fable 5.1 (61.0%), about 6 points above Claude Opus 5 (57.6%) and about 5 points below Claude Opus 5.5 (68.4%).
Every result from this document2Claude Sonnet 5.5 (max) Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety… 63.2
Printed as 63.2%Vendor-reported, measured Sep 2026Configuration: Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabledClaude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. pp. 138–139, Section 8.15.3; PhysicianBench pass@1, Claude Sonnet 5.5 (max).PhysicianBench is a public benchmark of 100 physician tasks carried out in an EHR. At max effort, Claude Sonnet 5.5 passes 63.2% of attempts, well above Claude Sonnet 5 (37.4%), level with Claude Fable 5.1 (61.0%), about 6 points above Claude Opus 5 (57.6%) and about 5 points below Claude Opus 5.5 (68.4%).
Every result from this document3Claude Fable 5.1 (max) Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety… 61.0
Printed as 61.0%Vendor-reported, measured Sep 2026Configuration: Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabledClaude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. pp. 138–139, Section 8.15.3; PhysicianBench pass@1, Claude Fable 5.1 (max).PhysicianBench is a public benchmark of 100 physician tasks carried out in an EHR. At max effort, Claude Sonnet 5.5 passes 63.2% of attempts, well above Claude Sonnet 5 (37.4%), level with Claude Fable 5.1 (61.0%), about 6 points above Claude Opus 5 (57.6%) and about 5 points below Claude Opus 5.5 (68.4%).
Every result from this document4Claude Opus 5 (max) Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety… 57.6
Printed as 57.6%Vendor-reported, measured Sep 2026Configuration: Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabledClaude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. pp. 138–139, Section 8.15.3; PhysicianBench pass@1, Claude Opus 5 (max).PhysicianBench is a public benchmark of 100 physician tasks carried out in an EHR. At max effort, Claude Sonnet 5.5 passes 63.2% of attempts, well above Claude Sonnet 5 (37.4%), level with Claude Fable 5.1 (61.0%), about 6 points above Claude Opus 5 (57.6%) and about 5 points below Claude Opus 5.5 (68.4%).
Every result from this document5Claude Sonnet 5.5 (xhigh) Anthropic harness; 100 tasks; pass@1; xhigh effort; Opus 5 rubric grader; safet… 56.4
Printed as 56.4%Vendor-reported, measured Sep 2026Configuration: Anthropic harness; 100 tasks; pass@1; xhigh effort; Opus 5 rubric grader; safety classifiers enabledClaude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. p. 138, Figure 8.15.B, right panel PhysicianBench (pass@1); orange Sonnet 5.5 xhigh label (visually read printed labels).Sonnet 5.5 | PhysicianBench (pass@1) | xhigh 56.4%
Every result from this document6Claude Sonnet 5.5 (high) Anthropic harness; 100 tasks; pass@1; high effort; Opus 5 rubric grader; safety… 47.6
Printed as 47.6%Vendor-reported, measured Sep 2026Configuration: Anthropic harness; 100 tasks; pass@1; high effort; Opus 5 rubric grader; safety classifiers enabledClaude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. p. 138, Figure 8.15.B, right panel PhysicianBench (pass@1); orange Sonnet 5.5 high label (visually read printed labels).Sonnet 5.5 | PhysicianBench (pass@1) | high 47.6%
Every result from this document7GPT-5.5 pass@1; Pass^3 28.0 46.3
Printed as 46.3 ± 1.2Official leaderboard, measured May 2026Configuration: pass@1; Pass^3 28.0PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026. p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)GPT-5.5 46.3 ± 1.2 57.4 28.0 41.9
Every result from this document- Also printed in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240): 46.3 ± 1.2, arXiv abs page (landing page for the PDF)
8Claude Sonnet 5 (max) Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety… 37.4
Printed as 37.4%Vendor-reported, measured Sep 2026Configuration: Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabledClaude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. pp. 138–139, Section 8.15.3; PhysicianBench pass@1, Claude Sonnet 5 (max).PhysicianBench is a public benchmark of 100 physician tasks carried out in an EHR. At max effort, Claude Sonnet 5.5 passes 63.2% of attempts, well above Claude Sonnet 5 (37.4%), level with Claude Fable 5.1 (61.0%), about 6 points above Claude Opus 5 (57.6%) and about 5 points below Claude Opus 5.5 (68.4%).
Every result from this document9Claude Opus 4.6 Pass^3 18.0 31.7
Printed as 31.7 ± 2.3Official leaderboard, measured May 2026Configuration: Pass^3 18.0PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026. p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)Claude Opus 4.6 31.7 ± 2.3 41.5 18.0 25.2
Every result from this document- Also printed in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240): 31.7 ± 2.3, arXiv abs page (landing page for the PDF)
10Claude Sonnet 5.5 (medium) Anthropic harness; 100 tasks; pass@1; medium effort; Opus 5 rubric grader; safe… 30.0
Printed as 30.0%Vendor-reported, measured Sep 2026Configuration: Anthropic harness; 100 tasks; pass@1; medium effort; Opus 5 rubric grader; safety classifiers enabledClaude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. p. 138, Figure 8.15.B, right panel PhysicianBench (pass@1); orange Sonnet 5.5 medium label (visually read printed labels).Sonnet 5.5 | PhysicianBench (pass@1) | medium 30.0%
Every result from this document11Claude Opus 4.7 Pass^3 18.0 29.3
Printed as 29.3 ± 2.5Official leaderboard, measured May 2026Configuration: Pass^3 18.0PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026. p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)Claude Opus 4.7 29.3 ± 2.5 37.9 18.0 16.2
Every result from this document- Also printed in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240): 29.3 ± 2.5, arXiv abs page (landing page for the PDF)
12GPT-5.4 27.7
Printed as 27.7 ± 1.5Official leaderboard, measured May 2026PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026. p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)GPT-5.4 27.7 ± 1.5 37.7 13.0 39.8
Every result from this document- Also printed in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240): 27.7 ± 1.5, arXiv abs page (landing page for the PDF)
13Claude Sonnet 5.5 (low) Anthropic harness; 100 tasks; pass@1; low effort; Opus 5 rubric grader; safety… 27.2
Printed as 27.2%Vendor-reported, measured Sep 2026Configuration: Anthropic harness; 100 tasks; pass@1; low effort; Opus 5 rubric grader; safety classifiers enabledClaude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. p. 138, Figure 8.15.B, right panel PhysicianBench (pass@1); orange Sonnet 5.5 low label (visually read printed labels).Sonnet 5.5 | PhysicianBench (pass@1) | low 27.2%
Every result from this document14Claude Sonnet 4.6 23.0
Printed as 23.0 ± 2.6Official leaderboard, measured May 2026PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026. p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)Claude Sonnet 4.6 23.0 ± 2.6 33.2 9.0 22.3
Every result from this document- Also printed in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240): 23.0 ± 2.6, arXiv abs page (landing page for the PDF)
15DeepSeek V4-Pro Pass@1 over 3 runs; Pass^3 6.0; shared FHIR tool harness; up to 100 turns; high… 18.7
Printed as 18.7 ± 2.9Official leaderboard, measured May 2026Configuration: Pass@1 over 3 runs; Pass^3 6.0; shared FHIR tool harness; up to 100 turns; high reasoning when supportedPhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026. arXiv HTML 2605.02240v1, Section 5.2, Table 2; DeepSeek V4-Pro row, Pass@1 column (PDF p. 8).Model | Pass@1 | Pass@3 | Pass^3 | #Turns DeepSeek V4-Pro | 18.7 ± 2.9 | 27.9 | 6.0 | 35.3
Every result from this document16Kimi-K2.6 open source 17.0
Printed as 17.0 ± 2.6Official leaderboard, measured May 2026Configuration: open sourcePhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026. p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)Kimi-K2.6 17.0 ± 2.6 26.3 5.0 42.4
Every result from this document- Also printed in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240): 17.0 ± 2.6, arXiv abs page (landing page for the PDF)
17MiMo-v2.5-Pro Pass@1 over 3 runs; Pass^3 6.0; shared FHIR tool harness; up to 100 turns; high… 16.7
Printed as 16.7 ± 4.0Official leaderboard, measured May 2026Configuration: Pass@1 over 3 runs; Pass^3 6.0; shared FHIR tool harness; up to 100 turns; high reasoning when supportedPhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026. arXiv HTML 2605.02240v1, Section 5.2, Table 2; MiMo-v2.5-Pro row, Pass@1 column (PDF p. 8).Model | Pass@1 | Pass@3 | Pass^3 | #Turns MiMo-v2.5-Pro | 16.7 ± 4.0 | 23.6 | 6.0 | 29.5
Every result from this document18Qwen3.6-Plus 13.7
Printed as 13.7 ± 4.0Official leaderboard, measured May 2026PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026. p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)Qwen3.6-Plus 13.7 ± 4.0 22.6 2.0 28.0
Every result from this document- Also printed in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240): 13.7 ± 4.0, arXiv abs page (landing page for the PDF)
19MiniMax M2.7 Pass@1 over 3 runs; Pass^3 1.0; shared FHIR tool harness; up to 100 turns; high… 8.7
Printed as 8.7 ± 1.2Official leaderboard, measured May 2026Configuration: Pass@1 over 3 runs; Pass^3 1.0; shared FHIR tool harness; up to 100 turns; high reasoning when supportedPhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026. arXiv HTML 2605.02240v1, Section 5.2, Table 2; MiniMax M2.7 row, Pass@1 column (PDF p. 8).Model | Pass@1 | Pass@3 | Pass^3 | #Turns MiniMax M2.7 | 8.7 ± 1.2 | 15.9 | 1.0 | 29.7
Every result from this document20Gemini Pro 3.1 6.0
Printed as 6.0 ± 1.0Official leaderboard, measured May 2026PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026. p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)Gemini Pro 3.1 6.0 ± 1.0 9.3 3.0 30.4
Every result from this document- Also printed in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240): 6.0 ± 1.0, arXiv abs page (landing page for the PDF)
21Grok-4.20 5.3
Printed as 5.3 ± 3.2Official leaderboard, measured May 2026PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026. p. 8, Table 2 (Proprietary Models block)Grok-4.20 5.3 ± 3.2 9.7 1.0 16.7
Every result from this document
Documents
3
- Claude Sonnet 5.5 System Cardsystem card, Anthropic, 28 Sep 2026Results it supports
- PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240)paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026Results it supports
- PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF)paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026Results it supports