Clinical Benchmarks

GPT-5.6 Terra

Released 9 Jul 2026$2 input, $12 output per million tokens1.05M contextproprietaryCompare with other models

Clinical Benchmarks Index
69.5rank 22 of 148; 69.5 × 1 = 69.5, from 3 of 10 boards
Boards
4 of 13
Results
4
Latest measurement
Sep 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

Clinical reasoning and knowledge

  1. HealthBench Professional
    length-adjusted, max reasoning effort (62.4 unadjusted, 3,618 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-TERRA.
    0.577
    Rank 14 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Jun 2026

Documentation and coding

  1. MedScribe (Vals AI)
    model ID openai/gpt-5.6-terra; max_output_tokens=30000; reasoning_effort=xhigh
    82.87
    Rank 47 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  2. MedCode (Vals AI)
    model ID openai/gpt-5.6-terra; reasoning_effort=xhigh; max_output_tokens=30000
    43.41
    Rank 42 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

EHR and workflow agents

  1. CHI-Bench
    Listed as codex + gpt-5.6-terra
    All Domains pass@1; codex harness
    13.3
    Rank 26 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured Jul 2026

Sources

Open a line for the quote and page.

  1. 14HealthBench Professional length-adjusted, max reasoning effort (62.4 unadjusted, 3,618 mean response cha… 0.577
    Printed as 57.7Vendor-reported, measured Jun 2026Configuration: length-adjusted, max reasoning effort (62.4 unadjusted, 3,618 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-TERRA.
    GPT-5.6 System Card system card, OpenAI, 9 Jul 2026. Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-TERRA
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    Every result from this document
    • Also printed in GPT-5.6 Preview System Card: 57.7, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-TERRA
  2. 47MedScribe (Vals AI) model ID openai/gpt-5.6-terra; max_output_tokens=30000; reasoning_effort=xhigh 82.87
    Printed as 82.87%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5.6-terra; max_output_tokens=30000; reasoning_effort=xhigh
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 48 of 105 (GPT-5.6 Terra), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-5.6-terra"].
    48 | GPT-5.6 Terra | 82.87%±1.95 | $2/$12 | 35.53s
    Every result from this document
  3. 42MedCode (Vals AI) model ID openai/gpt-5.6-terra; reasoning_effort=xhigh; max_output_tokens=30000 43.41
    Printed as 43.41%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-5.6-terra; reasoning_effort=xhigh; max_output_tokens=30000
    Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 42 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["openai/gpt-5.6-terra"].
    42 | GPT-5.6 Terra | 43.41%±2.17 | $2/$12 | 18.41s
    Every result from this document
  4. 26CHI-Bench All Domains pass@1; codex harness 13.3
    Printed as 13.3%Official leaderboard, measured Jul 2026Configuration: All Domains pass@1; codex harness
    CHI-Bench leaderboard (actAVA) official leaderboard, actAVA, 12 Aug 2026. CHI-Bench v1.0.0, All Domains, rank 26, Agent/Model and Accuracy columns; PA/UM/CM follow; board last updated 2026-08-12; board Date column 2026-07-24; accessed 2026-09-30. Linked submission provenance: run 2026-07-22T04:52:42.763467Z to 2026-07-22T06:23:17.742720Z; submitted_at 2026-07-24T23:01:25Z.
    26 | codex | gpt-5.6-terra | Proprietary | 13.3% | 12.0% | 20.0% | 8.0% | 2026-07-24
    Every result from this document

Other OpenAI models: GPT-4.1, GPT-4.1 mini, GPT-4o, GPT-5, GPT-5.1, GPT-5.2, GPT-5.3-Codex, GPT-5.4, GPT-5.4 mini, GPT-5.4 nano, GPT-5.5, GPT-5.5 Instant, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5 mini, GPT-5 nano, GPT-6.1 Sol, GPT-6 Astra, GPT-6 Luna, GPT-6 Sol, GPT OSS 120B, GPT OSS 20B, o3, o4-mini