Clinical Benchmarks

Moonshot AI

Moonshot AI models with results here, each placed on its own board.

Models
3
Results
18
Boards
9 of 13

www.moonshot.ai

Models

ModelKindReleasedBoards
Kimi K3Model16 Jul 20265
Kimi K2.6Model20 Apr 20265
Kimi K2.5Model27 Jan 20265

Results

One line per result. Dark tick: the board leader.

Clinical reasoning and knowledge

  1. MedXpertQA (MM)
    Qwen-run comparison in the Qwen3.5-397B-A17B model card (K2.5-1T-A32B column)
    65.3
    Rank 14 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
    Independent run
  2. 60.1
    Rank 2 of 8 here, 11 models on the boardLeader GPT-5.6 Sol 60.2
    Official leaderboard
    Measured Aug 2026

Documentation and coding

  1. 87.96
    Rank 13 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  2. MedScribe (Vals AI)
    model ID kimi/kimi-k2.6; top_p=0.95; max_output_tokens=30000
    78.15
    Rank 62 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  3. MedScribe (Vals AI)
    model ID kimi/kimi-k2.5-thinking; top_p=0.95; max_output_tokens=30000
    76.44
    Rank 72 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  4. MedCode (Vals AI)
    model ID kimi/kimi-k3; temperature=1; max_output_tokens=30000
    48.88
    Rank 24 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  5. MedCode (Vals AI)
    model ID kimi/kimi-k2.6; top_p=0.95; max_output_tokens=30000
    40.14
    Rank 64 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  6. MedCode (Vals AI)
    model ID kimi/kimi-k2.5-thinking; top_p=0.95; max_output_tokens=30000
    39.32
    Rank 65 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

EHR and workflow agents

  1. 17.0
    Rank 16 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
    Official leaderboard
    Measured May 2026
  2. EHR-Complex
    headline 12-model evaluation, top open-weight
    0.620
    Rank 3 of 18Leader GPT-5.4 (high reasoning) 0.650
    Official leaderboard
    Measured Jun 2026
  3. 25.3
    Rank 7 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured Jul 2026
  4. CHI-Bench
    All Domains pass@1; hermes harness
    15.6
    Rank 21 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026
  5. CHI-Bench
    All Domains pass@1; openai-agents harness
    15.1
    Rank 23 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026
  6. CHI-Bench
    All Domains pass@1; openclaw harness
    10.2
    Rank 32 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026
  7. CHI-Bench
    All Domains pass@1; deepagents harness
    3.1
    Rank 40 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026
  8. HealthAdminBench
    screenshot-only, detailed prompting
    15.6
    Rank 3 of 7 here, 5 models on the boardLeader Claude Opus 4.6 (computer-use agent) 36.3
    Official leaderboard
    Measured Apr 2026

Safety

  1. 74.0
    Rank 7 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2
    Official leaderboard
    Measured Aug 2026
  2. 59.1
    Rank 15 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2
    Official leaderboard
    Measured Aug 2026