MAST (Medical AI Superintelligence Test): current results
ARISE AI Research Network (multi-institutional) · composite of 6 component benchmarks; 11 models · index updated August 16, 2026
GPT-5.6 Sol holds the top current result on MAST (Medical AI Superintelligence Test), 60.2% as of 2026-08, per ARISE MAST leaderboard. Composite score across curated clinical benchmarks spanning diagnostic reasoning, management reasoning, safety, multimodal images, multimodal radiology, and agentic capability. Components: First Do NOHARM v2, SCT-Bench, MedAgentBench v2, PhysicianBench, ReXrank Mini, CPC-Bench.
Current results
- 1
GPT-5.6 Sol60.2%
- 2
Kimi K360.1%
- 3
Gemini 3.6 Flash59.3%
- 4
Gemini 3.1 Pro58.9%
- 5AQwen3.5 397B A17B57.9%
- 6
Claude Opus 557.1%
- 7
Claude Sonnet 556.6%
- 8
Grok 4.353.7%
Result detail
| # | model | score | as of | |
|---|---|---|---|---|
| 1 | GPT-5.6 Sol OpenAI MAST in preview; 'exact scores may change' | 60.2% | 2026-08 | |
| 2 | Kimi K3 Moonshot AI | 60.1% | 2026-08 | |
| 3 | Gemini 3.6 Flash Google | 59.3% | 2026-08 | |
| 4 | Gemini 3.1 Pro Google | 58.9% | 2026-08 | |
| 5 | A | Qwen3.5 397B A17B Alibaba | 57.9% | 2026-08 |
| 6 | Claude Opus 5 Anthropic | 57.1% | 2026-08 | |
| 7 | Claude Sonnet 5 Anthropic | 56.6% | 2026-08 | |
| 8 | Grok 4.3 xAI | 53.7% | 2026-08 | |
Scores appear exactly as ARISE MAST leaderboard publishes them (independently run). The board is marked as a preview and was last updated August 15, 2026; component-level breakdowns are published only for First, Do NOHARM v2. Scores may move before the full release.
About the benchmark
| publisher | ARISE AI Research Network (multi-institutional) |
|---|---|
| category | composite indices |
| released | 2026-08 |
| size | composite of 6 component benchmarks; 11 models |
| scale | percentage composite, higher better |
| result basis | independently run |
| source | ARISE MAST leaderboard |
| last frontier result | 2026-08 |
What is MAST (Medical AI Superintelligence Test)?
MAST (Medical AI Superintelligence Test) is a composite benchmark from ARISE AI Research Network, released 2026-08: composite of 6 component benchmarks; 11 models, scored on a percentage composite scale. Composite score across curated clinical benchmarks spanning diagnostic reasoning, management reasoning, safety, multimodal images, multimodal radiology, and agentic capability. Components: First Do NOHARM v2, SCT-Bench, MedAgentBench v2, PhysicianBench, ReXrank Mini, CPC-Bench.
Which model leads MAST (Medical AI Superintelligence Test)?
GPT-5.6 Sol (OpenAI) holds the top current result on MAST (Medical AI Superintelligence Test) at 60.2%, per ARISE MAST leaderboard, as of 2026-08.
Where do the MAST (Medical AI Superintelligence Test) numbers come from?
From ARISE MAST leaderboard (independently run). The board is marked as a preview and was last updated August 15, 2026; component-level breakdowns are published only for First, Do NOHARM v2. Scores may move before the full release.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.