Clinical Benchmarks

MiMo V2.5 Pro

1M context1.02T A42BopenAlso written as xiaomi/mimo-v2.5-proCompare with other models

Clinical Benchmarks Index
46.1rank 72 of 148; 46.1 × 1 = 46.1, from 3 of 10 boards
Boards
3 of 13
Results
3
Latest measurement
Sep 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

Documentation and coding

  1. MedScribe (Vals AI)
    model ID xiaomi/mimo-v2.5-pro; temperature=1; top_p=0.95; max_output_tokens=30000
    83.73
    Rank 39 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  2. MedCode (Vals AI)
    model ID xiaomi/mimo-v2.5-pro; temperature=1; top_p=0.95; max_output_tokens=30000
    32.48
    Rank 85 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

EHR and workflow agents

  1. PhysicianBench
    Pass@1 over 3 runs; Pass^3 6.0; shared FHIR tool harness; up to 100 turns; high reasoning when supported
    16.7
    Rank 17 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
    Official leaderboard
    Measured May 2026

Sources

Open a line for the quote and page.

  1. 39MedScribe (Vals AI) model ID xiaomi/mimo-v2.5-pro; temperature=1; top_p=0.95; max_output_tokens=300… 83.73
    Printed as 83.73%Official leaderboard, measured Sep 2026Configuration: model ID xiaomi/mimo-v2.5-pro; temperature=1; top_p=0.95; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 39 of 105 (MiMo V2.5 Pro), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["xiaomi/mimo-v2.5-pro"].
    39 | MiMo V2.5 Pro | 83.73%±2.06 | $0.435/$0.87 | 90.71s
    Every result from this document
  2. 85MedCode (Vals AI) model ID xiaomi/mimo-v2.5-pro; temperature=1; top_p=0.95; max_output_tokens=300… 32.48
    Printed as 32.48%Official leaderboard, measured Sep 2026Configuration: model ID xiaomi/mimo-v2.5-pro; temperature=1; top_p=0.95; max_output_tokens=30000
    Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 85 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["xiaomi/mimo-v2.5-pro"].
    85 | MiMo V2.5 Pro | 32.48%±1.91 | $0.435/$0.87 | 38.45s
    Every result from this document
  3. 17PhysicianBench Pass@1 over 3 runs; Pass^3 6.0; shared FHIR tool harness; up to 100 turns; high… 16.7
    Printed as 16.7 ± 4.0Official leaderboard, measured May 2026Configuration: Pass@1 over 3 runs; Pass^3 6.0; shared FHIR tool harness; up to 100 turns; high reasoning when supported
    PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) paper, Stanford University (HealthRex; Liu, Chen et al.), 4 May 2026. arXiv HTML 2605.02240v1, Section 5.2, Table 2; MiMo-v2.5-Pro row, Pass@1 column (PDF p. 8).
    Model | Pass@1 | Pass@3 | Pass^3 | #Turns MiMo-v2.5-Pro | 16.7 ± 4.0 | 23.6 | 6.0 | 29.5
    Every result from this document

Other Xiaomi models: MiMo V2.5, MiMo V2.6 Flash, MiMo V2.6 Pro