HealthAdminBench
Computer-use agents completing healthcare administration workflows: prior authorizations, denial appeals, and DME ordering.
More about this boardLess
Computer-use agents completing healthcare administration workflows: prior authorizations, denial appeals, and DME ordering.
End-to-end task success of computer-use LLM agents on healthcare administration workflows: prior authorizations, denial appeals, and DME ordering; success requires completing every subtask in a task.
Run by Kinetic Systems' research team: independent of the model vendors, though published by a company selling healthcare-admin automation. No GPT-5.6 or Claude 5 rows yet, and no stated refresh cadence.
Published by Kinetic Systems (with Stanford Hospital domain experts), released Apr 2026. 135 tasks / 1,698 rubric-scored subtasks. Percentage end-to-end task success 0-100, higher better.
- Headline metric
- end-to-end task success rate, n=135 tasks; screenshot-only observations, Task Description + Portal Guidance prompting
- Board last updated
- 10 Apr 2026
- Rows
- 7
- Models
- 5
- Labs
- 5
- Last measured
- Apr 2026
- Leader
- 36.3Claude Opus 4.6 (computer-use agent)
Ranking
Percentage end-to-end task success 0-100, higher better
- Anthropic
- OpenAI
- Alibaba
- Moonshot AI
7 of 7 rows
Rows and sources
Open a row for the quote, the page and the document.
1Claude Opus 4.6 (computer-use agent) screenshot-only, detailed prompting; native CUA harness; subtask rate 78.4% 36.3
Printed as 36.3%Official leaderboard, measured Apr 2026Configuration: screenshot-only, detailed prompting; native CUA harness; subtask rate 78.4%HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) paper, Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care), 10 Apr 2026. p. 8, Figure 3(a), Task Success Rate, bar "Claude Opus 4.6 CUA"; section 4, p. 7 gives screenshot-only, Task Description + Portal Guidance; arXiv 2604.09937v1 dated 2026-04-10; checked 2026-09-30.Claude Opus 4.6 CUA | 36.3%
Every result from this document- Also printed in Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog): 36.3%, Section 'LLMs struggle with long-horizon tasks'
2GPT-5.4 (computer-use agent) screenshot-only, detailed prompting; subtask rate 82.8% 26.7
Printed as 26.7%Official leaderboard, measured Apr 2026Configuration: screenshot-only, detailed prompting; subtask rate 82.8%HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) paper, Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care), 10 Apr 2026. p. 8, Figure 3(a), Task Success Rate, bar "GPT-5.4 CUA"; section 4, p. 7 gives screenshot-only, Task Description + Portal Guidance; arXiv 2604.09937v1 dated 2026-04-10; checked 2026-09-30.GPT-5.4 CUA | 26.7%
Every result from this document- Also printed in Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog): 26.7%, Section 'LLMs struggle with long-horizon tasks'
3Kimi K2.5 screenshot-only, detailed prompting 15.6
Printed as 15.6%Official leaderboard, measured Apr 2026Configuration: screenshot-only, detailed promptingHealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) paper, Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care), 10 Apr 2026. p. 8, Figure 3(a), Task Success Rate, bar "Kimi K2.5"; section 4, p. 7 gives screenshot-only, Task Description + Portal Guidance; arXiv 2604.09937v1 dated 2026-04-10; checked 2026-09-30.Kimi K2.5 | 15.6%
Every result from this document- Also printed in Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog): 15.6%, Section 'LLMs struggle with long-horizon tasks'
4Claude Opus 4.6 (standardized harness) screenshot-only, detailed prompting; authors' standardized harness, no native C… 14.8
Printed as 14.8%Official leaderboard, measured Apr 2026Configuration: screenshot-only, detailed prompting; authors' standardized harness, no native CUAHealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) paper, Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care), 10 Apr 2026. p. 8, Figure 3(a), Task Success Rate, bar "Claude Opus 4.6"; section 4, p. 7 gives screenshot-only, Task Description + Portal Guidance; arXiv 2604.09937v1 dated 2026-04-10; checked 2026-09-30.Claude Opus 4.6 | 14.8%
Every result from this document5Qwen 3.5 screenshot-only, detailed prompting 13.3
Printed as 13.3%Official leaderboard, measured Apr 2026Configuration: screenshot-only, detailed promptingHealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) paper, Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care), 10 Apr 2026. p. 8, Figure 3(a), Task Success Rate, bar "Qwen 3.5"; section 4, p. 7 gives screenshot-only, Task Description + Portal Guidance; arXiv 2604.09937v1 dated 2026-04-10; checked 2026-09-30.Qwen 3.5 | 13.3%
Every result from this document- Also printed in Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog): 13.3%, Section 'LLMs struggle with long-horizon tasks'
6Gemini 3.1 Pro screenshot-only, detailed prompting 11.9
Printed as 11.9%Official leaderboard, measured Apr 2026Configuration: screenshot-only, detailed promptingHealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) paper, Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care), 10 Apr 2026. p. 8, Figure 3(a), Task Success Rate, bar "Gemini 3.1 Pro"; section 4, p. 7 gives screenshot-only, Task Description + Portal Guidance; arXiv 2604.09937v1 dated 2026-04-10; checked 2026-09-30.Gemini 3.1 Pro | 11.9%
Every result from this document- Also printed in Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog): 11.9%, Section 'LLMs struggle with long-horizon tasks'
7GPT-5.4 (standardized harness) screenshot-only, detailed prompting; authors' standardized harness, no native C… 5.9
Printed as 5.9%Official leaderboard, measured Apr 2026Configuration: screenshot-only, detailed prompting; authors' standardized harness, no native CUAHealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) paper, Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care), 10 Apr 2026. p. 8, Figure 3(a), Task Success Rate, bar "GPT-5.4"; section 4, p. 7 gives screenshot-only, Task Description + Portal Guidance; arXiv 2604.09937v1 dated 2026-04-10; checked 2026-09-30.GPT-5.4 | 5.9%
Every result from this document
Documents
2
- HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1)paper, Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care), 10 Apr 2026Results it supports
- Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog)blog post, Kinetic Systems, 14 Apr 2026Results it supports