Clinical Benchmarks

PhysicianBench

Agents carrying out long-horizon physician workflows inside real EHR systems, verified by execution against those systems.

academic teamOfficial pagePaper
More about this boardLess

Agents carrying out long-horizon physician workflows inside real EHR systems, verified by execution against those systems.

LLM agents on long-horizon composite physician workflows inside real EHR environments, with execution-grounded verification against actual EHR systems via standard commercial APIs.

Scores come from the paper; there is no standalone public leaderboard, and the results also feed the ARISE MAST composite. The lower half of the table falls steeply, with several agents near 1 percent on Pass^3 consistency.

Published by Academic team (Ruoqi Liu, Imran Q. Mohiuddin et al., arXiv 2605.02240); also a MAST component, released May 2026. 100 real-world clinical tasks, 21 specialties, 670 structured checkpoints (~27 tool calls per task). Pass@1 success rate %, higher better (3 independent runs; Pass^3 also reported).

Table
Table 2, p. 8
Headline metric
Pass@1 (%), mean ± sd over 3 runs
Paper date
4 May 2026
Paper version
v1
Models in the paper's table
12
Rows
21
Models
17 of 12on the board
Labs
9
Last measured
Sep 2026
Leader
68.4Claude Opus 5.5 (max)

Ranking

Pass@1 success rate %, higher better (3 independent runs; Pass^3 also reported)

  • Anthropic
  • Google
  • OpenAI
  • Alibaba
  • Moonshot AI
  • Other labs
  1. 1Claude Opus 5.5 (max)Anthropic harness68.4
  2. 2Claude Sonnet 5.5 (max)Anthropic harness63.2
  3. 3Claude Fable 5.1 (max)Anthropic harness61.0
  4. 4Claude Opus 5 (max)Anthropic harness57.6
  5. 5Claude Sonnet 5.5 (xhigh)Anthropic harness56.4
  6. 6Claude Sonnet 5.5 (high)Anthropic harness47.6
  7. 7GPT-5.5pass@146.3
  8. 8Claude Sonnet 5 (max)Anthropic harness37.4
  9. 9Claude Opus 4.6Pass^3 1831.7
  10. 10Claude Sonnet 5.5 (medium)Anthropic harness30.0
  11. 11Claude Opus 4.7Pass^3 1829.3
  12. 12GPT-5.427.7
  13. 13Claude Sonnet 5.5 (low)Anthropic harness27.2
  14. 14Claude Sonnet 4.623.0
  15. 15DeepSeek V4-ProPass@1 over 3 runs18.7
  16. 16Kimi-K2.6open source17.0
  17. 17MiMo-v2.5-ProPass@1 over 3 runs16.7
  18. 18Qwen3.6-Plus13.7
  19. 19MiniMax M2.7Pass@1 over 3 runs8.7
  20. 20Gemini Pro 3.16.0

20 of 21 rows

Rows and sources

Open a row for the quote, the page and the document.

  1. 1Claude Opus 5.5 (max) Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety… 68.4
    Printed as 68.4%Vendor-reported, measured Sep 2026Configuration: Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabled
    Claude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. pp. 138–139, Section 8.15.3; PhysicianBench pass@1, Claude Opus 5.5 (max).
    PhysicianBench is a public benchmark of 100 physician tasks carried out in an EHR. At max effort, Claude Sonnet 5.5 passes 63.2% of attempts, well above Claude Sonnet 5 (37.4%), level with Claude Fable 5.1 (61.0%), about 6 points above Claude Opus 5 (57.6%) and about 5 points below Claude Opus 5.5 (68.4%).
    Every result from this document
  2. 2Claude Sonnet 5.5 (max) Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety… 63.2
    Printed as 63.2%Vendor-reported, measured Sep 2026Configuration: Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabled
    Claude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. pp. 138–139, Section 8.15.3; PhysicianBench pass@1, Claude Sonnet 5.5 (max).
    PhysicianBench is a public benchmark of 100 physician tasks carried out in an EHR. At max effort, Claude Sonnet 5.5 passes 63.2% of attempts, well above Claude Sonnet 5 (37.4%), level with Claude Fable 5.1 (61.0%), about 6 points above Claude Opus 5 (57.6%) and about 5 points below Claude Opus 5.5 (68.4%).
    Every result from this document
  3. 3Claude Fable 5.1 (max) Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety… 61.0
    Printed as 61.0%Vendor-reported, measured Sep 2026Configuration: Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabled
    Claude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. pp. 138–139, Section 8.15.3; PhysicianBench pass@1, Claude Fable 5.1 (max).
    PhysicianBench is a public benchmark of 100 physician tasks carried out in an EHR. At max effort, Claude Sonnet 5.5 passes 63.2% of attempts, well above Claude Sonnet 5 (37.4%), level with Claude Fable 5.1 (61.0%), about 6 points above Claude Opus 5 (57.6%) and about 5 points below Claude Opus 5.5 (68.4%).
    Every result from this document
  4. 4Claude Opus 5 (max) Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety… 57.6
    Printed as 57.6%Vendor-reported, measured Sep 2026Configuration: Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabled
    Claude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. pp. 138–139, Section 8.15.3; PhysicianBench pass@1, Claude Opus 5 (max).
    PhysicianBench is a public benchmark of 100 physician tasks carried out in an EHR. At max effort, Claude Sonnet 5.5 passes 63.2% of attempts, well above Claude Sonnet 5 (37.4%), level with Claude Fable 5.1 (61.0%), about 6 points above Claude Opus 5 (57.6%) and about 5 points below Claude Opus 5.5 (68.4%).
    Every result from this document
  5. 5Claude Sonnet 5.5 (xhigh) Anthropic harness; 100 tasks; pass@1; xhigh effort; Opus 5 rubric grader; safet… 56.4
    Printed as 56.4%Vendor-reported, measured Sep 2026Configuration: Anthropic harness; 100 tasks; pass@1; xhigh effort; Opus 5 rubric grader; safety classifiers enabled
    Claude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. p. 138, Figure 8.15.B, right panel PhysicianBench (pass@1); orange Sonnet 5.5 xhigh label (visually read printed labels).
    Sonnet 5.5 | PhysicianBench (pass@1) | xhigh 56.4%
    Every result from this document
  6. 6Claude Sonnet 5.5 (high) Anthropic harness; 100 tasks; pass@1; high effort; Opus 5 rubric grader; safety… 47.6
    Printed as 47.6%Vendor-reported, measured Sep 2026Configuration: Anthropic harness; 100 tasks; pass@1; high effort; Opus 5 rubric grader; safety classifiers enabled
    Claude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. p. 138, Figure 8.15.B, right panel PhysicianBench (pass@1); orange Sonnet 5.5 high label (visually read printed labels).
    Sonnet 5.5 | PhysicianBench (pass@1) | high 47.6%
    Every result from this document
  7. 7GPT-5.5 pass@1; Pass^3 28.0 46.3
    Printed as 46.3 ± 1.2Official leaderboard, measured May 2026Configuration: pass@1; Pass^3 28.0
    PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026. p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
    GPT-5.5 46.3 ± 1.2 57.4 28.0 41.9
    Every result from this document
  8. 8Claude Sonnet 5 (max) Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety… 37.4
    Printed as 37.4%Vendor-reported, measured Sep 2026Configuration: Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabled
    Claude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. pp. 138–139, Section 8.15.3; PhysicianBench pass@1, Claude Sonnet 5 (max).
    PhysicianBench is a public benchmark of 100 physician tasks carried out in an EHR. At max effort, Claude Sonnet 5.5 passes 63.2% of attempts, well above Claude Sonnet 5 (37.4%), level with Claude Fable 5.1 (61.0%), about 6 points above Claude Opus 5 (57.6%) and about 5 points below Claude Opus 5.5 (68.4%).
    Every result from this document
  9. 9Claude Opus 4.6 Pass^3 18.0 31.7
    Printed as 31.7 ± 2.3Official leaderboard, measured May 2026Configuration: Pass^3 18.0
    PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026. p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
    Claude Opus 4.6 31.7 ± 2.3 41.5 18.0 25.2
    Every result from this document
  10. 10Claude Sonnet 5.5 (medium) Anthropic harness; 100 tasks; pass@1; medium effort; Opus 5 rubric grader; safe… 30.0
    Printed as 30.0%Vendor-reported, measured Sep 2026Configuration: Anthropic harness; 100 tasks; pass@1; medium effort; Opus 5 rubric grader; safety classifiers enabled
    Claude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. p. 138, Figure 8.15.B, right panel PhysicianBench (pass@1); orange Sonnet 5.5 medium label (visually read printed labels).
    Sonnet 5.5 | PhysicianBench (pass@1) | medium 30.0%
    Every result from this document
  11. 11Claude Opus 4.7 Pass^3 18.0 29.3
    Printed as 29.3 ± 2.5Official leaderboard, measured May 2026Configuration: Pass^3 18.0
    PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026. p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
    Claude Opus 4.7 29.3 ± 2.5 37.9 18.0 16.2
    Every result from this document
  12. 12GPT-5.4 27.7
    Printed as 27.7 ± 1.5Official leaderboard, measured May 2026
    PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026. p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
    GPT-5.4 27.7 ± 1.5 37.7 13.0 39.8
    Every result from this document
  13. 13Claude Sonnet 5.5 (low) Anthropic harness; 100 tasks; pass@1; low effort; Opus 5 rubric grader; safety… 27.2
    Printed as 27.2%Vendor-reported, measured Sep 2026Configuration: Anthropic harness; 100 tasks; pass@1; low effort; Opus 5 rubric grader; safety classifiers enabled
    Claude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. p. 138, Figure 8.15.B, right panel PhysicianBench (pass@1); orange Sonnet 5.5 low label (visually read printed labels).
    Sonnet 5.5 | PhysicianBench (pass@1) | low 27.2%
    Every result from this document
  14. 14Claude Sonnet 4.6 23.0
    Printed as 23.0 ± 2.6Official leaderboard, measured May 2026
    PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026. p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
    Claude Sonnet 4.6 23.0 ± 2.6 33.2 9.0 22.3
    Every result from this document
  15. 15DeepSeek V4-Pro Pass@1 over 3 runs; Pass^3 6.0; shared FHIR tool harness; up to 100 turns; high… 18.7
    Printed as 18.7 ± 2.9Official leaderboard, measured May 2026Configuration: Pass@1 over 3 runs; Pass^3 6.0; shared FHIR tool harness; up to 100 turns; high reasoning when supported
    PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026. arXiv HTML 2605.02240v1, Section 5.2, Table 2; DeepSeek V4-Pro row, Pass@1 column (PDF p. 8).
    Model | Pass@1 | Pass@3 | Pass^3 | #Turns DeepSeek V4-Pro | 18.7 ± 2.9 | 27.9 | 6.0 | 35.3
    Every result from this document
  16. 16Kimi-K2.6 open source 17.0
    Printed as 17.0 ± 2.6Official leaderboard, measured May 2026Configuration: open source
    PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026. p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
    Kimi-K2.6 17.0 ± 2.6 26.3 5.0 42.4
    Every result from this document
  17. 17MiMo-v2.5-Pro Pass@1 over 3 runs; Pass^3 6.0; shared FHIR tool harness; up to 100 turns; high… 16.7
    Printed as 16.7 ± 4.0Official leaderboard, measured May 2026Configuration: Pass@1 over 3 runs; Pass^3 6.0; shared FHIR tool harness; up to 100 turns; high reasoning when supported
    PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026. arXiv HTML 2605.02240v1, Section 5.2, Table 2; MiMo-v2.5-Pro row, Pass@1 column (PDF p. 8).
    Model | Pass@1 | Pass@3 | Pass^3 | #Turns MiMo-v2.5-Pro | 16.7 ± 4.0 | 23.6 | 6.0 | 29.5
    Every result from this document
  18. 18Qwen3.6-Plus 13.7
    Printed as 13.7 ± 4.0Official leaderboard, measured May 2026
    PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026. p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
    Qwen3.6-Plus 13.7 ± 4.0 22.6 2.0 28.0
    Every result from this document
  19. 19MiniMax M2.7 Pass@1 over 3 runs; Pass^3 1.0; shared FHIR tool harness; up to 100 turns; high… 8.7
    Printed as 8.7 ± 1.2Official leaderboard, measured May 2026Configuration: Pass@1 over 3 runs; Pass^3 1.0; shared FHIR tool harness; up to 100 turns; high reasoning when supported
    PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026. arXiv HTML 2605.02240v1, Section 5.2, Table 2; MiniMax M2.7 row, Pass@1 column (PDF p. 8).
    Model | Pass@1 | Pass@3 | Pass^3 | #Turns MiniMax M2.7 | 8.7 ± 1.2 | 15.9 | 1.0 | 29.7
    Every result from this document
  20. 20Gemini Pro 3.1 6.0
    Printed as 6.0 ± 1.0Official leaderboard, measured May 2026
    PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026. p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
    Gemini Pro 3.1 6.0 ± 1.0 9.3 3.0 30.4
    Every result from this document
  21. 21Grok-4.20 5.3
    Printed as 5.3 ± 3.2Official leaderboard, measured May 2026
    PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026. p. 8, Table 2 (Proprietary Models block)
    Grok-4.20 5.3 ± 3.2 9.7 1.0 16.7
    Every result from this document

Documents

3