Clinical Benchmarks

CHI-Bench: current results

actAVA.ai · 75 workflows (25 per domain), 21 healthcare applications, 200+ MCP tools · index updated August 16, 2026

erius + claude-opus-5 holds the top current result on CHI-Bench, 54.7% as of 2026-08, per CHI-Bench leaderboard (actAVA). Long-horizon US healthcare operations workflows for agents (prior authorization, utilization management, care management), 60-80 step tasks across 4-6 stages, judged by deterministic unit tests plus an LLM judge for evidence grounding, consent, and cross-stage consistency.

Current results

Result detail

#modelscoreas of
1Humana (harness) / Anthropic (model) logoerius + claude-opus-5 Humana (harness) / Anthropic (model)
community-submitted harness config validated by automated workspace judge
54.7%2026-08
2Humana / Anthropic logoerius + claude-opus-4-8 Humana / Anthropic37.3%2026-08
3Anthropic logoclaude-code + claude-opus-5 Anthropic37.3%2026-08
4Anthropic logoclaude-code + claude-opus-4-8 Anthropic33.3%2026-08
5Anthropic logoclaude-code + claude-opus-4-6 Anthropic28.0%2026-08
6Anthropic logoclaude-code + claude-sonnet-4-6 Anthropic26.2%2026-08
7OpenAI logocodex + gpt-5.6-sol OpenAI25.3%2026-08
8Moonshot AI logoopenai-agents + kimi-k3 Moonshot AI25.3%2026-08
9Anthropic logoclaude-code + claude-opus-4-7 Anthropic24.4%2026-08
10Anthropic logoclaude-code + claude-fable-5 Anthropic24.0%2026-08

Scores appear exactly as CHI-Bench leaderboard (actAVA) publishes them (mixed sources). Released May 20, 2026; updated August 12, 2026 across 45 harness configurations. The launch report led with reliability, not capability: no agent stayed above 20 percent across three identical runs. Harness choice matters as much as model choice, and the board accepts community submissions, so rows mix author-run and submitted results.

About the benchmark

publisheractAVA.ai
categoryagentic and workflow benchmarks
released2026-05
size75 workflows (25 per domain), 21 healthcare applications, 200+ MCP tools
scalepass@1 with binary 0/1 reward, higher better
result basismixed sources
sourceCHI-Bench leaderboard (actAVA)
last frontier result2026-08

What is CHI-Bench?

CHI-Bench is a agentic and workflow benchmark from actAVA, released 2026-05: 75 workflows (25 per domain), 21 healthcare applications, 200+ MCP tools, scored on a pass@1 with binary 0/1 reward scale. Long-horizon US healthcare operations workflows for agents (prior authorization, utilization management, care management), 60-80 step tasks across 4-6 stages, judged by deterministic unit tests plus an LLM judge for evidence grounding, consent, and cross-stage consistency.

Which model leads CHI-Bench?

erius + claude-opus-5 (Humana (harness) / Anthropic (model)) holds the top current result on CHI-Bench at 54.7%, per CHI-Bench leaderboard (actAVA), as of 2026-08.

Where do the CHI-Bench numbers come from?

From CHI-Bench leaderboard (actAVA) (mixed sources). Released May 20, 2026; updated August 12, 2026 across 45 harness configurations. The launch report led with reliability, not capability: no agent stayed above 20 percent across three identical runs. Harness choice matters as much as model choice, and the board accepts community submissions, so rows mix author-run and submitted results.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.