Kimi K3
Released 16 Jul 2026$3 input, $15 output per million tokens1M context2.8T total / 104B activeKimi K3 License (open weights)Compare with other models
- Clinical Benchmarks Index
- 84.3rank 7 of 148; 84.3 × 1 = 84.3, from 4 of 10 boards
- Boards
- 5 of 13
- Results
- 5
- Latest measurement
- Sep 2026
Results
Each line is placed on its own board. Dark tick: the board leader.
Clinical reasoning and knowledge
- 60.1Rank 2 of 8 here, 11 models on the boardLeader GPT-5.6 Sol 60.2
Documentation and coding
- 87.96Rank 13 of 105Leader Claude Opus 5.5 91.43
- MedCode (Vals AI)model ID kimi/kimi-k3; temperature=1; max_output_tokens=3000048.88Rank 24 of 103Leader Claude Opus 5 63.57
EHR and workflow agents
- CHI-BenchListed as openai-agents + kimi-k325.3Rank 7 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
Safety
- 74.0Rank 7 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2
Sources
Open a line for the quote and page.
13MedScribe (Vals AI) 87.96
Printed as 87.96%Official leaderboard, measured Sep 2026Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 13 of 105 (Kimi K3), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["kimi/kimi-k3"].13 | Kimi K3 | 87.96%±1.89 | $3/$15 | 2m17s
Every result from this document24MedCode (Vals AI) model ID kimi/kimi-k3; temperature=1; max_output_tokens=30000 48.88
Printed as 48.88%Official leaderboard, measured Sep 2026Configuration: model ID kimi/kimi-k3; temperature=1; max_output_tokens=30000Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 24 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["kimi/kimi-k3"].24 | Kimi K3 | 48.88%±2.19 | $3/$15 | 116.18s
Every result from this document2MAST (Medical AI Superintelligence Test) 60.1
Printed as 60.1%Official leaderboard, measured Aug 2026MAST: Medical AI Superintelligence Test leaderboard (General board) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')2 Kimi K3 Moonshot AI 60.1%
Every result from this document7First, Do NOHARM (v2) 74.0
Printed as 74.0%Official leaderboard, measured Aug 2026MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column7Kimi K3OSSMoonshot AI 74.0%
Every result from this document7CHI-Bench 25.3
Printed as 25.3%Official leaderboard, measured Jul 2026CHI-Bench leaderboard (actAVA) official leaderboard, actAVA, 12 Aug 2026. CHI-Bench v1.0.0, All Domains, rank 08, Agent/Model and Accuracy columns; PA/UM/CM follow; board last updated 2026-08-12; board Date column 2026-07-24; accessed 2026-09-30. Linked submission provenance: run 2026-07-22T08:03:23.788878Z to 2026-07-22T14:22:38.269923Z; submitted_at 2026-08-12T20:20:00Z.08 | openai-agents | kimi-k3 | Open-source | 25.3% | 28.0% | 32.0% | 16.0% | 2026-07-24
Every result from this document- Also printed in CHI-Bench submission manifest: openai-agents + moonshotai/kimi-k3, submission.model; provenance.started_at / finished_at; results.overall.pass_at_1; retrieved 2026-09-30