Clinical Benchmarks

HealthAdminBench

Computer-use agents completing healthcare administration workflows: prior authorizations, denial appeals, and DME ordering.

Kinetic SystemsOfficial pagePaper
More about this boardLess

Computer-use agents completing healthcare administration workflows: prior authorizations, denial appeals, and DME ordering.

End-to-end task success of computer-use LLM agents on healthcare administration workflows: prior authorizations, denial appeals, and DME ordering; success requires completing every subtask in a task.

Run by Kinetic Systems' research team: independent of the model vendors, though published by a company selling healthcare-admin automation. No GPT-5.6 or Claude 5 rows yet, and no stated refresh cadence.

Published by Kinetic Systems (with Stanford Hospital domain experts), released Apr 2026. 135 tasks / 1,698 rubric-scored subtasks. Percentage end-to-end task success 0-100, higher better.

Headline metric
end-to-end task success rate, n=135 tasks; screenshot-only observations, Task Description + Portal Guidance prompting
Board last updated
10 Apr 2026
Rows
7
Models
5
Labs
5
Last measured
Apr 2026
Leader
36.3Claude Opus 4.6 (computer-use agent)

Ranking

Percentage end-to-end task success 0-100, higher better

  • Anthropic
  • Google
  • OpenAI
  • Alibaba
  • Moonshot AI
All official leaderboard
  1. 1Claude Opus 4.6 (computer-use agent)screenshot-only36.3
  2. 2GPT-5.4 (computer-use agent)screenshot-only26.7
  3. 3Kimi K2.5screenshot-only15.6
  4. 4Claude Opus 4.6 (standardized harness)screenshot-only14.8
  5. 5Qwen 3.5screenshot-only13.3
  6. 6Gemini 3.1 Proscreenshot-only11.9
  7. 7GPT-5.4 (standardized harness)screenshot-only5.9

7 of 7 rows

Rows and sources

Open a row for the quote, the page and the document.

  1. 1Claude Opus 4.6 (computer-use agent) screenshot-only, detailed prompting; native CUA harness; subtask rate 78.4% 36.3
    Printed as 36.3%Official leaderboard, measured Apr 2026Configuration: screenshot-only, detailed prompting; native CUA harness; subtask rate 78.4%
    HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) paper, Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care), 10 Apr 2026. p. 8, Figure 3(a), Task Success Rate, bar "Claude Opus 4.6 CUA"; section 4, p. 7 gives screenshot-only, Task Description + Portal Guidance; arXiv 2604.09937v1 dated 2026-04-10; checked 2026-09-30.
    Claude Opus 4.6 CUA | 36.3%
    Every result from this document
  2. 2GPT-5.4 (computer-use agent) screenshot-only, detailed prompting; subtask rate 82.8% 26.7
    Printed as 26.7%Official leaderboard, measured Apr 2026Configuration: screenshot-only, detailed prompting; subtask rate 82.8%
    HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) paper, Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care), 10 Apr 2026. p. 8, Figure 3(a), Task Success Rate, bar "GPT-5.4 CUA"; section 4, p. 7 gives screenshot-only, Task Description + Portal Guidance; arXiv 2604.09937v1 dated 2026-04-10; checked 2026-09-30.
    GPT-5.4 CUA | 26.7%
    Every result from this document
  3. 3Kimi K2.5 screenshot-only, detailed prompting 15.6
    Printed as 15.6%Official leaderboard, measured Apr 2026Configuration: screenshot-only, detailed prompting
    HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) paper, Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care), 10 Apr 2026. p. 8, Figure 3(a), Task Success Rate, bar "Kimi K2.5"; section 4, p. 7 gives screenshot-only, Task Description + Portal Guidance; arXiv 2604.09937v1 dated 2026-04-10; checked 2026-09-30.
    Kimi K2.5 | 15.6%
    Every result from this document
  4. 4Claude Opus 4.6 (standardized harness) screenshot-only, detailed prompting; authors' standardized harness, no native C… 14.8
    Printed as 14.8%Official leaderboard, measured Apr 2026Configuration: screenshot-only, detailed prompting; authors' standardized harness, no native CUA
    HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) paper, Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care), 10 Apr 2026. p. 8, Figure 3(a), Task Success Rate, bar "Claude Opus 4.6"; section 4, p. 7 gives screenshot-only, Task Description + Portal Guidance; arXiv 2604.09937v1 dated 2026-04-10; checked 2026-09-30.
    Claude Opus 4.6 | 14.8%
    Every result from this document
  5. 5Qwen 3.5 screenshot-only, detailed prompting 13.3
    Printed as 13.3%Official leaderboard, measured Apr 2026Configuration: screenshot-only, detailed prompting
    HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) paper, Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care), 10 Apr 2026. p. 8, Figure 3(a), Task Success Rate, bar "Qwen 3.5"; section 4, p. 7 gives screenshot-only, Task Description + Portal Guidance; arXiv 2604.09937v1 dated 2026-04-10; checked 2026-09-30.
    Qwen 3.5 | 13.3%
    Every result from this document
  6. 6Gemini 3.1 Pro screenshot-only, detailed prompting 11.9
    Printed as 11.9%Official leaderboard, measured Apr 2026Configuration: screenshot-only, detailed prompting
    HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) paper, Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care), 10 Apr 2026. p. 8, Figure 3(a), Task Success Rate, bar "Gemini 3.1 Pro"; section 4, p. 7 gives screenshot-only, Task Description + Portal Guidance; arXiv 2604.09937v1 dated 2026-04-10; checked 2026-09-30.
    Gemini 3.1 Pro | 11.9%
    Every result from this document
  7. 7GPT-5.4 (standardized harness) screenshot-only, detailed prompting; authors' standardized harness, no native C… 5.9
    Printed as 5.9%Official leaderboard, measured Apr 2026Configuration: screenshot-only, detailed prompting; authors' standardized harness, no native CUA
    HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) paper, Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care), 10 Apr 2026. p. 8, Figure 3(a), Task Success Rate, bar "GPT-5.4"; section 4, p. 7 gives screenshot-only, Task Description + Portal Guidance; arXiv 2604.09937v1 dated 2026-04-10; checked 2026-09-30.
    GPT-5.4 | 5.9%
    Every result from this document

Documents

2