Moonshot AI
Moonshot AI models with results here, each placed on its own board.
- Models
- 3
- Results
- 18
- Boards
- 9 of 13
Models
Results
One line per result. Dark tick: the board leader.
Clinical reasoning and knowledge
- MedXpertQA (MM)Qwen-run comparison in the Qwen3.5-397B-A17B model card (K2.5-1T-A32B column)65.3Rank 14 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
- 60.1Rank 2 of 8 here, 11 models on the boardLeader GPT-5.6 Sol 60.2
Documentation and coding
- 87.96Rank 13 of 105Leader Claude Opus 5.5 91.43
- 78.15Rank 62 of 105Leader Claude Opus 5.5 91.43
- 76.44Rank 72 of 105Leader Claude Opus 5.5 91.43
- 48.88Rank 24 of 103Leader Claude Opus 5 63.57
- 40.14Rank 64 of 103Leader Claude Opus 5 63.57
- 39.32Rank 65 of 103Leader Claude Opus 5 63.57
EHR and workflow agents
- 17.0Rank 16 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
- 0.620Rank 3 of 18Leader GPT-5.4 (high reasoning) 0.650
- 25.3Rank 7 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
- 15.6Rank 21 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
- 15.1Rank 23 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
- 10.2Rank 32 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
- 3.1Rank 40 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
- 15.6Rank 3 of 7 here, 5 models on the boardLeader Claude Opus 4.6 (computer-use agent) 36.3
Safety
- 74.0Rank 7 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2
- 59.1Rank 15 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2