Clinical Benchmarks

GLM 5.1

Released 7 Apr 2026$1.4 input, $4.4 output per million tokens200K contextopenAlso written as zai/glm-5.1Compare with other models

Clinical Benchmarks Index
45.6rank 75 of 148; 45.6 × 1 = 45.6, from 3 of 10 boards
Boards
4 of 13
Results
7
Latest measurement
Sep 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

Documentation and coding

  1. MedScribe (Vals AI)
    model ID zai/glm-5.1; temperature=1; top_p=0.95; max_output_tokens=30000
    72.27
    Rank 85 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  2. MedCode (Vals AI)
    model ID zai/glm-5.1; temperature=1; top_p=0.95; max_output_tokens=30000
    41.60
    Rank 49 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

EHR and workflow agents

  1. CHI-Bench
    Listed as hermes + glm-5.1
    All Domains pass@1; hermes harness
    18.7
    Rank 14 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026
  2. CHI-Bench
    Listed as openai-agents + glm-5.1
    All Domains pass@1; openai-agents harness
    18.7
    Rank 14 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026
  3. CHI-Bench
    Listed as openclaw + glm-5.1
    All Domains pass@1; openclaw harness
    16.9
    Rank 18 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026
  4. CHI-Bench
    Listed as deepagents + glm-5.1
    All Domains pass@1; deepagents harness
    11.1
    Rank 29 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026

Safety

  1. First, Do NOHARM (v2)
    First Do NOHARM v2 overall; preview; open-weight model; run date unpublished
    57.9
    Rank 16 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2
    Official leaderboard
    Measured Sep 2026

Sources

Open a line for the quote and page.

  1. 85MedScribe (Vals AI) model ID zai/glm-5.1; temperature=1; top_p=0.95; max_output_tokens=30000 72.27
    Printed as 72.27%Official leaderboard, measured Sep 2026Configuration: model ID zai/glm-5.1; temperature=1; top_p=0.95; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 85 of 105 (GLM 5.1), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["zai/glm-5.1"].
    85 | GLM 5.1 | 72.27%±2.06 | $1/$3.2 | 95.70s
    Every result from this document
  2. 49MedCode (Vals AI) model ID zai/glm-5.1; temperature=1; top_p=0.95; max_output_tokens=30000 41.60
    Printed as 41.60%Official leaderboard, measured Sep 2026Configuration: model ID zai/glm-5.1; temperature=1; top_p=0.95; max_output_tokens=30000
    Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 49 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["zai/glm-5.1"].
    49 | GLM 5.1 | 41.60%±2.12 | $1/$3.2 | 77.58s
    Every result from this document
  3. 16First, Do NOHARM (v2) First Do NOHARM v2 overall; preview; open-weight model; run date unpublished 57.9
    Printed as 57.9%Official leaderboard, measured Sep 2026Configuration: First Do NOHARM v2 overall; preview; open-weight model; run date unpublished
    ARISE MAST technical leaderboard official leaderboard, ARISE. First Do NOHARM v2, overall leaderboard; GLM 5.1 row; displayed 2026-09-30.
    First Do NOHARM v2 overall metric across 19 models 15GLM 5.1OSSZ.ai 57.9%
    Every result from this document
  4. 14CHI-Bench All Domains pass@1; hermes harness 18.7
    Printed as 18.7%Official leaderboard, measured May 2026Configuration: All Domains pass@1; hermes harness
    CHI-Bench leaderboard (actAVA) official leaderboard, actAVA, 12 Aug 2026. CHI-Bench v1.0.0, All Domains, rank 15, Agent/Model and Accuracy columns; PA/UM/CM follow; board last updated 2026-08-12; board Date column 2026-05-01; accessed 2026-09-30. Linked submission provenance: run 2026-05-03T07:01:48Z to 2026-05-04T05:47:30Z; submitted_at 2026-05-04T05:47:30Z.
    15 | hermes | glm-5.1 | Open-source | 18.7% | 10.7% | 34.7% | 10.7% | 2026-05-01
    Every result from this document
  5. 14CHI-Bench All Domains pass@1; openai-agents harness 18.7
    Printed as 18.7%Official leaderboard, measured May 2026Configuration: All Domains pass@1; openai-agents harness
    CHI-Bench leaderboard (actAVA) official leaderboard, actAVA, 12 Aug 2026. CHI-Bench v1.0.0, All Domains, rank 14, Agent/Model and Accuracy columns; PA/UM/CM follow; board last updated 2026-08-12; board Date column 2026-05-01; accessed 2026-09-30. Linked submission provenance: run 2026-05-03T08:41:54Z to 2026-05-04T07:03:18Z; submitted_at 2026-05-04T07:03:18Z.
    14 | openai-agents | glm-5.1 | Open-source | 18.7% | 18.7% | 33.3% | 4.0% | 2026-05-01
    Every result from this document
  6. 18CHI-Bench All Domains pass@1; openclaw harness 16.9
    Printed as 16.9%Official leaderboard, measured May 2026Configuration: All Domains pass@1; openclaw harness
    CHI-Bench leaderboard (actAVA) official leaderboard, actAVA, 12 Aug 2026. CHI-Bench v1.0.0, All Domains, rank 18, Agent/Model and Accuracy columns; PA/UM/CM follow; board last updated 2026-08-12; board Date column 2026-05-01; accessed 2026-09-30. Linked submission provenance: run 2026-05-02T09:10:57Z to 2026-05-04T02:49:29Z; submitted_at 2026-05-04T02:49:29Z.
    18 | openclaw | glm-5.1 | Open-source | 16.9% | 13.3% | 26.7% | 10.7% | 2026-05-01
    Every result from this document
  7. 29CHI-Bench All Domains pass@1; deepagents harness 11.1
    Printed as 11.1%Official leaderboard, measured May 2026Configuration: All Domains pass@1; deepagents harness
    CHI-Bench leaderboard (actAVA) official leaderboard, actAVA, 12 Aug 2026. CHI-Bench v1.0.0, All Domains, rank 30, Agent/Model and Accuracy columns; PA/UM/CM follow; board last updated 2026-08-12; board Date column 2026-05-01; accessed 2026-09-30. Linked submission provenance: run 2026-05-03T21:51:53Z to 2026-05-04T13:10:34Z; submitted_at 2026-05-04T13:10:34Z.
    30 | deepagents | glm-5.1 | Open-source | 11.1% | 17.3% | 10.7% | 5.3% | 2026-05-01
    Every result from this document

Other Zhipu models: GLM 4.7, GLM 5.2, GLM 5.3, GLM 5.3 Flash