Anthropic
Anthropic models with results here, each placed on its own board.
- Models
- 16
- Results
- 92
- Boards
- 13 of 13
Models
| Model | Kind | Released | Boards |
|---|---|---|---|
| Claude Sonnet 5.5 | Model | 28 Sep 2026 | 4 |
| Claude Opus 5.5 | Model | 22 Sep 2026 | 4 |
| Claude Fable 5.1 | Model | 1 Sep 2026 | 4 |
| Claude Opus 5 | Model | 24 Jul 2026 | 8 |
| Claude Sonnet 5 | Model | 30 Jun 2026 | 6 |
| Claude Fable 5 | Model | 9 Jun 2026 | 6 |
| Claude Opus 4.8 | Model | 28 May 2026 | 6 |
| Claude Opus 4.7 | Model | 16 Apr 2026 | 6 |
| Claude Sonnet 4.6 | Model | 17 Feb 2026 | 6 |
| Claude Opus 4.6 | Model | 5 Feb 2026 | 8 |
| Claude Opus 4.5 | Model | 24 Nov 2025 | 3 |
| Claude Haiku 4.5 | Model | 15 Oct 2025 | 3 |
| Claude Sonnet 4.5 | Model | 29 Sep 2025 | 2 |
| Claude Opus 4.1 | Model | 5 Aug 2025 | 2 |
| Claude Sonnet 4 | Model | 22 May 2025 | 2 |
| Claude 3.7 Sonnet | Model | 24 Feb 2025 | 1 |
Results
One line per result. Dark tick: the board leader.
Clinical reasoning and knowledge
- HealthBench ProfessionalAnthropic; max effort; Opus 4.8 grader; safety classifiers enabled; length-adjusted; raw 77.1%0.692Rank 2 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- HealthBench Professionallength-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 70.3%). Measured as Claude Mythos 5; Anthropic itself prints 66.0 in the Fable 5 column of the Opus 5 card with footnote Mythos 5.0.660Rank 3 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- HealthBench ProfessionalAnthropic; adaptive max effort; Opus 4.8 grader; five trials; no tools or custom system prompt; safety classifiers and refusal fallback to Opus 5; length-adjusted; raw 77.1%0.656Rank 4 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- HealthBench ProfessionalMeasured as Claude Fable 5; Anthropic September card; length-adjusted; adaptive max effort; Opus 4.8 grader; five trials; no tools or custom system prompt; raw 68.9%0.633Rank 6 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- HealthBench Professionallength-adjusted (method published in the HealthBench Professional paper); Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over five trials, no tools or customized system prompt; Fable 5.1 run with safety classifiers active and a refusal-fallback to Claude Opus 5 (raw 74.2%).0.621Rank 7 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- HealthBench Professionallength-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 73.4%).0.598Rank 11 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- HealthBench Professionallength-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 62.4%).0.578Rank 13 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- HealthBench ProfessionalAnthropic June evaluation; length-adjusted; adaptive max effort; Opus 4.8 grader; five trials; no tools or custom system prompt0.574Rank 15 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- HealthBench Professionallength-adjusted, adaptive thinking at max effort, Claude Sonnet 4.6 grader (the grader used by the Opus 4.8 card itself); the other Claude rows on this board use the Claude Opus 4.8 grader, under which Anthropic later prints 57.4 for this model.0.558Rank 17 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- HealthBench Professionallength-adjusted; adaptive thinking at max effort; Claude Sonnet 4.6 grader; 5 trials; comparison model in the Opus 4.8 card0.519Rank 21 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- HealthBench Professionallength-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt; comparison column in the Sonnet 5 card. Other Anthropic prints: 44.4% (Fable card Figure 8.18.2.A), 41.7% (Opus 4.8 card, Sonnet 4.6 grader)0.442Rank 27 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- 80.0Rank 4 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
- 71.7Rank 9 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
- MedXpertQA (MM)Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)64.8Rank 15 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
- MedXpertQA (MM)Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors63.6Rank 16 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
- 57.1Rank 6 of 8 here, 11 models on the boardLeader GPT-5.6 Sol 60.2
- 56.6Rank 7 of 8 here, 11 models on the boardLeader GPT-5.6 Sol 60.2
- 0.456Rank 8 of 10 here, 11 models on the boardLeader Gemini 3.1 Pro (Preview) 0.652
- 0.450Rank 9 of 10 here, 11 models on the boardLeader Gemini 3.1 Pro (Preview) 0.652
Documentation and coding
- MedScribe (Vals AI)model ID anthropic/claude-opus-5-5; temperature=1; max_output_tokens=128000; compute_effort=max91.43Rank 1 of 105Leads this board
- 91.29Rank 2 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID anthropic/claude-sonnet-5-5; temperature=1; max_output_tokens=128000; compute_effort=max91.10Rank 3 of 105Leader Claude Opus 5.5 91.43
- 90.98Rank 4 of 105Leader Claude Opus 5.5 91.43
- 88.52Rank 10 of 105Leader Claude Opus 5.5 91.43
- 86.74Rank 18 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID anthropic/claude-opus-4-6-thinking; temperature=1; max_output_tokens=30000; compute_effort=max86.13Rank 21 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID anthropic/claude-opus-4-8; temperature=1; max_output_tokens=30000; compute_effort=max85.75Rank 23 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID anthropic/claude-opus-4-5-20251101-thinking; temperature=1; max_output_tokens=30000; compute_effort=high85.32Rank 26 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID anthropic/claude-haiku-4-5-20251001-thinking; temperature=1; max_output_tokens=3000085.23Rank 28 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID anthropic/claude-sonnet-4-5-20250929; temperature=1; max_output_tokens=3000084.52Rank 31 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID anthropic/claude-sonnet-4-5-20250929-thinking; temperature=1; max_output_tokens=3000084.10Rank 36 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID anthropic/claude-opus-4-5-20251101; temperature=1; max_output_tokens=30000; compute_effort=high83.25Rank 44 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID anthropic/claude-opus-4-7; temperature=1; max_output_tokens=30000; compute_effort=max82.95Rank 46 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID anthropic/claude-sonnet-5; temperature=1; max_output_tokens=30000; compute_effort=max76.05Rank 74 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID anthropic/claude-opus-4-1-20250805-thinking; temperature=1; max_output_tokens=3000073.90Rank 79 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=3000072.41Rank 84 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID anthropic/claude-opus-4-1-20250805; temperature=1; max_output_tokens=3000071.75Rank 88 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=3000069.35Rank 92 of 105Leader Claude Opus 5.5 91.43
- 63.57Rank 1 of 103Leads this board
- 56.07Rank 3 of 103Leader Claude Opus 5 63.57
- 54.86Rank 6 of 103Leader Claude Opus 5 63.57
- 53.51Rank 7 of 103Leader Claude Opus 5 63.57
- 53.22Rank 9 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID anthropic/claude-sonnet-5-5; compute_effort=max; temperature=1; max_output_tokens=12800052.92Rank 11 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID anthropic/claude-opus-5-5; compute_effort=max; temperature=1; max_output_tokens=12800049.80Rank 16 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID anthropic/claude-opus-4-5-20251101-thinking; compute_effort=high; temperature=1; max_output_tokens=3000049.16Rank 21 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID anthropic/claude-opus-4-6-thinking; compute_effort=max; temperature=1; max_output_tokens=3000049.13Rank 22 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID anthropic/claude-opus-4-6; compute_effort=max; temperature=1; max_output_tokens=3000048.24Rank 27 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID anthropic/claude-sonnet-5; compute_effort=max; temperature=1; max_output_tokens=3000047.54Rank 30 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID anthropic/claude-opus-4-1-20250805-thinking; temperature=1; max_output_tokens=3000047.23Rank 32 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID anthropic/claude-opus-4-5-20251101; compute_effort=high; temperature=1; max_output_tokens=3000045.17Rank 35 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID anthropic/claude-sonnet-4-5-20250929-thinking; temperature=1; max_output_tokens=3000044.13Rank 39 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID anthropic/claude-opus-4-1-20250805; temperature=1; max_output_tokens=3000041.37Rank 51 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID anthropic/claude-sonnet-4-5-20250929; temperature=1; max_output_tokens=3000040.57Rank 59 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=3000034.96Rank 75 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=3000033.94Rank 79 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID anthropic/claude-haiku-4-5-20251001-thinking; temperature=1; max_output_tokens=3000032.68Rank 84 of 103Leader Claude Opus 5 63.57
EHR and workflow agents
- PhysicianBenchAnthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabled68.4Rank 1 of 21 here, 12 models on the boardLeads this board
- PhysicianBenchAnthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabled63.2Rank 2 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
- PhysicianBenchAnthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabled61.0Rank 3 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
- PhysicianBenchAnthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabled57.6Rank 4 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
- PhysicianBenchAnthropic harness; 100 tasks; pass@1; xhigh effort; Opus 5 rubric grader; safety classifiers enabled56.4Rank 5 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
- PhysicianBenchAnthropic harness; 100 tasks; pass@1; high effort; Opus 5 rubric grader; safety classifiers enabled47.6Rank 6 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
- PhysicianBenchAnthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabled37.4Rank 8 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
- 31.7Rank 9 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
- PhysicianBenchAnthropic harness; 100 tasks; pass@1; medium effort; Opus 5 rubric grader; safety classifiers enabled30.0Rank 10 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
- 29.3Rank 11 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
- PhysicianBenchAnthropic harness; 100 tasks; pass@1; low effort; Opus 5 rubric grader; safety classifiers enabled27.2Rank 13 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
- 23.0Rank 14 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
- 0.360Rank 13 of 18Leader GPT-5.4 (high reasoning) 0.650
- 55Rank 1 of 12Leads this board
- 36Rank 4 of 12Leader Claude Code (Opus 5) 55
- 32Rank 6 of 12Leader Claude Code (Opus 5) 55
- 27Rank 8 of 12Leader Claude Code (Opus 5) 55
- 19Rank 10 of 12Leader Claude Code (Opus 5) 55
- 17Rank 11 of 12Leader Claude Code (Opus 5) 55
- CHI-Benchcommunity-submitted harness config validated by automated workspace judge54.7Rank 1 of 43 here, 45 models on the boardLeads this board
- 37.3Rank 2 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
- 37.3Rank 2 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
- 33.3Rank 4 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
- 28.0Rank 5 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
- 26.2Rank 6 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
- 24.4Rank 9 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
- 24.0Rank 10 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
- 20.0Rank 13 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
- 17.3Rank 17 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
- 6.2Rank 36 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
- HealthAdminBenchscreenshot-only, detailed prompting; native CUA harness; subtask rate 78.4%36.3Rank 1 of 7 here, 5 models on the boardLeads this board
- HealthAdminBenchscreenshot-only, detailed prompting; authors' standardized harness, no native CUA14.8Rank 4 of 7 here, 5 models on the boardLeader Claude Opus 4.6 (computer-use agent) 36.3
Safety
- 64.2Rank 9 of 28Leader Gemini-3.1-Pro 80.7
- 74.6Rank 6 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2
- 65.0Rank 11 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2