Clinical Benchmarks

Ant Group

Ant Group models with results here, each placed on its own board.

Models
4
Results
6
Boards
3 of 13

www.antgroup.com

Models

ModelKindReleasedBoards
Ling 3.0 FlashModelNot stated2
Ling 3.0 Flash FinModelNot stated2
Qwen3-14B-SFTModelNot stated1
Qwen3-32B-SFTModelNot stated1

Results

One line per result. Dark tick: the board leader.

Documentation and coding

  1. MedScribe (Vals AI)
    model ID ant/ling-3.0-flash-2607; temperature=1; top_p=0.95; max_output_tokens=30000
    80.90
    Rank 51 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  2. MedScribe (Vals AI)
    model ID ant/ling-3.0-flash-af-rc3; temperature=1; top_p=0.95; max_output_tokens=131072
    75.59
    Rank 76 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  3. MedCode (Vals AI)
    model ID ant/ling-3.0-flash-2607; temperature=1; top_p=0.95; max_output_tokens=30000
    32.27
    Rank 86 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  4. MedCode (Vals AI)
    model ID ant/ling-3.0-flash-af-rc3; temperature=1; top_p=0.95; max_output_tokens=131072
    29.30
    Rank 94 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

EHR and workflow agents

  1. EHR-Complex
    Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns; fine-tuned on trajectories from EHR-Complex training set
    0.550
    Rank 8 of 18Leader GPT-5.4 (high reasoning) 0.650
    Official leaderboard
    Measured Jun 2026
  2. EHR-Complex
    Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns; fine-tuned on trajectories from EHR-Complex training set
    0.450
    Rank 12 of 18Leader GPT-5.4 (high reasoning) 0.650
    Official leaderboard
    Measured Jun 2026