Clinical Benchmarks

Zhipu

Zhipu models with results here, each placed on its own board.

Models
5
Results
15
Boards
4 of 13

www.zhipuai.cn

Models

ModelKindReleasedBoards
GLM 5.3 FlashModel26 Aug 20261
GLM 5.3Model18 Aug 20262
GLM 5.2Model16 Jun 20263
GLM 5.1Model7 Apr 20264
GLM 4.7Model22 Dec 20252

Results

One line per result. Dark tick: the board leader.

Documentation and coding

  1. MedScribe (Vals AI)
    model ID zai/glm-5.3-flash; temperature=1; top_p=0.95; max_output_tokens=30000; reasoning_effort=max
    88.94
    Rank 7 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  2. MedScribe (Vals AI)
    model ID zai/glm-5.3; temperature=1; top_p=0.95; max_output_tokens=30000; reasoning_effort=max
    88.81
    Rank 9 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  3. MedScribe (Vals AI)
    model ID zai/glm-5.2; temperature=1; max_output_tokens=30000
    83.53
    Rank 43 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  4. MedScribe (Vals AI)
    model ID zai/glm-5.1; temperature=1; top_p=0.95; max_output_tokens=30000
    72.27
    Rank 85 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  5. MedScribe (Vals AI)
    model ID zai/glm-4.7; temperature=1; top_p=1; max_output_tokens=30000
    68.63
    Rank 94 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  6. MedCode (Vals AI)
    model ID zai/glm-5.3; reasoning_effort=max; temperature=1; top_p=0.95; max_output_tokens=30000
    42.86
    Rank 46 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  7. MedCode (Vals AI)
    model ID zai/glm-5.1; temperature=1; top_p=0.95; max_output_tokens=30000
    41.60
    Rank 49 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  8. MedCode (Vals AI)
    model ID zai/glm-5.2; temperature=1; max_output_tokens=30000
    40.77
    Rank 57 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  9. MedCode (Vals AI)
    model ID zai/glm-4.7; temperature=1; top_p=1; max_output_tokens=30000
    32.77
    Rank 83 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

EHR and workflow agents

  1. CHI-Bench
    All Domains pass@1; openai-agents harness
    18.7
    Rank 14 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured Jul 2026
  2. CHI-Bench
    All Domains pass@1; hermes harness
    18.7
    Rank 14 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026
  3. CHI-Bench
    All Domains pass@1; openai-agents harness
    18.7
    Rank 14 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026
  4. CHI-Bench
    All Domains pass@1; openclaw harness
    16.9
    Rank 18 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026
  5. CHI-Bench
    All Domains pass@1; deepagents harness
    11.1
    Rank 29 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026

Safety

  1. First, Do NOHARM (v2)
    First Do NOHARM v2 overall; preview; open-weight model; run date unpublished
    57.9
    Rank 16 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2
    Official leaderboard
    Measured Sep 2026