Clinical Benchmarks

MAST (Medical AI Superintelligence Test): current results

ARISE AI Research Network (multi-institutional) · composite of 6 component benchmarks; 11 models · index updated August 16, 2026

GPT-5.6 Sol holds the top current result on MAST (Medical AI Superintelligence Test), 60.2% as of 2026-08, per ARISE MAST leaderboard. Composite score across curated clinical benchmarks spanning diagnostic reasoning, management reasoning, safety, multimodal images, multimodal radiology, and agentic capability. Components: First Do NOHARM v2, SCT-Bench, MedAgentBench v2, PhysicianBench, ReXrank Mini, CPC-Bench.

Current results

Result detail

#modelscoreas of
1OpenAI logoGPT-5.6 Sol OpenAI
MAST in preview; 'exact scores may change'
60.2%2026-08
2Moonshot AI logoKimi K3 Moonshot AI60.1%2026-08
3Google logoGemini 3.6 Flash Google59.3%2026-08
4Google logoGemini 3.1 Pro Google58.9%2026-08
5AQwen3.5 397B A17B Alibaba57.9%2026-08
6Anthropic logoClaude Opus 5 Anthropic57.1%2026-08
7Anthropic logoClaude Sonnet 5 Anthropic56.6%2026-08
8xAI logoGrok 4.3 xAI53.7%2026-08

Scores appear exactly as ARISE MAST leaderboard publishes them (independently run). The board is marked as a preview and was last updated August 15, 2026; component-level breakdowns are published only for First, Do NOHARM v2. Scores may move before the full release.

About the benchmark

publisherARISE AI Research Network (multi-institutional)
categorycomposite indices
released2026-08
sizecomposite of 6 component benchmarks; 11 models
scalepercentage composite, higher better
result basisindependently run
sourceARISE MAST leaderboard
last frontier result2026-08

What is MAST (Medical AI Superintelligence Test)?

MAST (Medical AI Superintelligence Test) is a composite benchmark from ARISE AI Research Network, released 2026-08: composite of 6 component benchmarks; 11 models, scored on a percentage composite scale. Composite score across curated clinical benchmarks spanning diagnostic reasoning, management reasoning, safety, multimodal images, multimodal radiology, and agentic capability. Components: First Do NOHARM v2, SCT-Bench, MedAgentBench v2, PhysicianBench, ReXrank Mini, CPC-Bench.

Which model leads MAST (Medical AI Superintelligence Test)?

GPT-5.6 Sol (OpenAI) holds the top current result on MAST (Medical AI Superintelligence Test) at 60.2%, per ARISE MAST leaderboard, as of 2026-08.

Where do the MAST (Medical AI Superintelligence Test) numbers come from?

From ARISE MAST leaderboard (independently run). The board is marked as a preview and was last updated August 15, 2026; component-level breakdowns are published only for First, Do NOHARM v2. Scores may move before the full release.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.