Clinical Benchmarks

Kimi K3

Released 16 Jul 2026$3 input, $15 output per million tokens1M context2.8T total / 104B activeKimi K3 License (open weights)Compare with other models

Clinical Benchmarks Index
84.3rank 7 of 148; 84.3 × 1 = 84.3, from 4 of 10 boards
Boards
5 of 13
Results
5
Latest measurement
Sep 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

Clinical reasoning and knowledge

  1. 60.1
    Rank 2 of 8 here, 11 models on the boardLeader GPT-5.6 Sol 60.2
    Official leaderboard
    Measured Aug 2026

Documentation and coding

  1. 87.96
    Rank 13 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  2. MedCode (Vals AI)
    model ID kimi/kimi-k3; temperature=1; max_output_tokens=30000
    48.88
    Rank 24 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

EHR and workflow agents

  1. CHI-Bench
    Listed as openai-agents + kimi-k3
    25.3
    Rank 7 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured Jul 2026

Safety

  1. 74.0
    Rank 7 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2
    Official leaderboard
    Measured Aug 2026

Sources

Open a line for the quote and page.

  1. 13MedScribe (Vals AI) 87.96
    Printed as 87.96%Official leaderboard, measured Sep 2026
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 13 of 105 (Kimi K3), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["kimi/kimi-k3"].
    13 | Kimi K3 | 87.96%±1.89 | $3/$15 | 2m17s
    Every result from this document
  2. 24MedCode (Vals AI) model ID kimi/kimi-k3; temperature=1; max_output_tokens=30000 48.88
    Printed as 48.88%Official leaderboard, measured Sep 2026Configuration: model ID kimi/kimi-k3; temperature=1; max_output_tokens=30000
    Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 24 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["kimi/kimi-k3"].
    24 | Kimi K3 | 48.88%±2.19 | $3/$15 | 116.18s
    Every result from this document
  3. 2MAST (Medical AI Superintelligence Test) 60.1
    Printed as 60.1%Official leaderboard, measured Aug 2026
    MAST: Medical AI Superintelligence Test leaderboard (General board) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
    2 Kimi K3 Moonshot AI 60.1%
    Every result from this document
  4. 7First, Do NOHARM (v2) 74.0
    Printed as 74.0%Official leaderboard, measured Aug 2026
    MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column
    7Kimi K3OSSMoonshot AI 74.0%
    Every result from this document
  5. 7CHI-Bench 25.3
    Printed as 25.3%Official leaderboard, measured Jul 2026
    CHI-Bench leaderboard (actAVA) official leaderboard, actAVA, 12 Aug 2026. CHI-Bench v1.0.0, All Domains, rank 08, Agent/Model and Accuracy columns; PA/UM/CM follow; board last updated 2026-08-12; board Date column 2026-07-24; accessed 2026-09-30. Linked submission provenance: run 2026-07-22T08:03:23.788878Z to 2026-07-22T14:22:38.269923Z; submitted_at 2026-08-12T20:20:00Z.
    08 | openai-agents | kimi-k3 | Open-source | 25.3% | 28.0% | 32.0% | 16.0% | 2026-07-24
    Every result from this document

Other Moonshot AI models: Kimi K2.5, Kimi K2.6