Moonshot AI: healthcare benchmark results
4 models · 6 results · updated August 16, 2026
Everything the index holds for Moonshot AI, gathered in one place: Kimi K3, Kimi-K2.5, openai-agents + kimi-k3, Kimi-K2.6. Scores sit on each benchmark's own scale and never compare across rows from different boards.
Every result
| model | benchmark | score | position |
|---|---|---|---|
| Kimi K3 | MAST (Medical AI Superintelligence Test) | 60.1% | 2 of 11 |
| Kimi K3 | First, Do NOHARM (v2) | 74.0% | 2 of 19 |
| Kimi-K2.5 | EHR-Complex | 0.62 | 3 of 12 |
| Kimi-K2.5 | HealthAdminBench | 15.6% | 3 of 5 |
| openai-agents + kimi-k3 | CHI-Bench | 25.3% | 8 of 45 |
| Kimi-K2.6 | PhysicianBench | 17.0 ± 2.6 | 6 of 13 |
Which healthcare benchmarks does Moonshot AI appear on?
As of August 16, 2026, Moonshot AI models hold 6 current results across 6 tracked benchmarks, through Kimi K3, Kimi-K2.5, openai-agents + kimi-k3, Kimi-K2.6.
Where does Moonshot AI lead?
Moonshot AI holds no first places on the tracked boards right now.
Other labs with pages: Anthropic, OpenAI, Google, Alibaba, Meta, xAI. The full field is on the index.