Clinical Benchmarks

MedXpertQA (MM)

Expert-level multiple-choice questions over clinical images across 17 specialties.

Tsinghua UniversityOfficial pagePaper
More about this boardLess

Expert-level multimodal medical multiple-choice QA covering clinical images (X-ray, histology, dermatology, charts) across 17 specialties; MM subset of the 4,460-question MedXpertQA benchmark.

Assembled table, not a single board. Rows come from several vendors' own launch documents: Meta's Muse Spark evaluation (whose values live only in a rendered table image and cover Meta's own model plus four competitors it re-ran or quoted), Alibaba's Qwen3.5, 3.7 and 3.8 posts (Alibaba's own runs of its models and of competitors), and Google's Gemma 4 model card. Protocols are not known to match across rows; each row's source line says which document it came from and who ran it. benchlm.ai presents a subset as one mirrored view while citing only Meta's methodology.

Published by TsinghuaC3I (Tsinghua University), released Jan 2025. 2,000 multimodal questions (MM subset). Percentage accuracy 0-100, higher better.

  • Accuracy. Same 2025 paper; multimodal test under its original prompt and answer-extraction protocol. Unit: %.
Board last updated
8 Apr 2026
Paper version
ICML 2025 / arXiv v3
Rows
22
Models
22 of 5on the board
Labs
7
Last measured
Aug 2026
Leader
81.5GPT-5.6 Sol

Ranking

Percentage accuracy 0-100, higher better

  • Anthropic
  • Google
  • OpenAI
  • Alibaba
  • Meta
  • Moonshot AI
  • Other labs
  1. 1GPT-5.6 SolQwen-run comparison in the Qwen381.5
  2. 2Gemini 3.1 ProMeta's Muse Spark launch table81.3
  3. 3Qwen3.8 MaxAlibaba's own Qwen380.4
  4. 4Claude Fable 5Qwen-run comparison in the Qwen380.0
  5. 5Muse SparkMeta's Muse Spark launch table78.4
  6. 6GPT-5.4Meta's Muse Spark launch table77.1
  7. 7Gemini-3 ProQwen376.0
  8. 8GPT-5.2Qwen-run comparison in the Qwen373.3
  9. 9Claude Opus 4.8Qwen-run comparison in the Qwen371.7
  10. 10Qwen3.7 PlusAlibaba's own Qwen371.0
  11. 11Qwen3.5 397B A17Bself-reported in the Qwen370.0
  12. 12Qwen3.6 PlusQwen-run comparison in the Qwen368.7
  13. 13Grok 4.20Meta's Muse Spark launch table65.8
  14. 14Kimi K2.5Qwen-run comparison in the Qwen365.3
  15. 15Claude Opus 4.6Meta's Muse Spark launch table64.8
  16. 16Claude 4.5 OpusQwen363.6
  17. 17Gemma 4 31BGoogle model card61.3
  18. 18Gemma 4 26B A4BGoogle model card58.1
  19. 19Gemma 4 12BGoogle's Gemma 4 model card48.7
  20. 20Qwen3-VL-235B-A22BQwen347.6

20 of 22 rows

Rows and sources

Open a row for the quote, the page and the document.

  1. 1GPT-5.6 Sol Qwen-run comparison in the Qwen3.8-Max launch post 81.5
    Printed as 81.5Independent runConfiguration: Qwen-run comparison in the Qwen3.8-Max launch post
    Qwen3.8-Max: A New Bar for Coding and Cowork launch post, Alibaba, 2 Aug 2026. Full Benchmark Table, second (multimodal) table, row MedXpertQA-MM; columns Opus4.8 / Fable5 / Gemini3.1-Pro / GPT5.6-Sol / Qwen3.7-Plus / Qwen3.8-Max
    | MedXpertQA-MM | 71.7 | 80.0 | 80.7 | 81.5 | 71.0 | 80.4 |
    Every result from this document
  2. 2Gemini 3.1 Pro Meta's Muse Spark launch table (Meta reports the better of vendor self-reports… 81.3
    Printed as 81.3%Independent run, measured Apr 2026Configuration: Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
    Introducing Muse Spark: Scaling Towards Personal Superintelligence launch post, Meta, 8 Apr 2026. Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column Gemini 3.1 Pro High; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'
    MedXpertQA (MM) | Medical Multiple Choice | 78.4 | 64.8 | 81.3 | 77.1 | 65.8 (image table row; columns Muse Spark Thinking, Opus 4.6 Max, Gemini 3.1 Pro High, GPT 5.4 Xhigh, Grok 4.2 Reasoning)
    Every result from this document
  3. 3Qwen3.8 Max Alibaba's own Qwen3.8 launch table; protocol not stated 80.4
    Printed as 80.4%Vendor-reported, measured Aug 2026Configuration: Alibaba's own Qwen3.8 launch table; protocol not stated
    Qwen3.8-Max: A New Bar for Coding and Cowork launch post, Alibaba, 2 Aug 2026. Multimodal Benchmarks table, Multimodal Reasoning section, row MedXpertQA-MM, column Qwen3.8-Max; header row: | | Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus | Qwen3.8-Max |
    | MedXpertQA-MM | 71.7 | 80.0 | 80.7 | 81.5 | 71.0 | 80.4 |
    Every result from this document
  4. 4Claude Fable 5 Qwen-run comparison in the Qwen3.8-Max launch post 80.0
    Printed as 80.0Independent runConfiguration: Qwen-run comparison in the Qwen3.8-Max launch post
    Qwen3.8-Max: A New Bar for Coding and Cowork launch post, Alibaba, 2 Aug 2026. Full Benchmark Table, second (multimodal) table, row MedXpertQA-MM; columns Opus4.8 / Fable5 / Gemini3.1-Pro / GPT5.6-Sol / Qwen3.7-Plus / Qwen3.8-Max
    | MedXpertQA-MM | 71.7 | 80.0 | 80.7 | 81.5 | 71.0 | 80.4 |
    Every result from this document
  5. 5Muse Spark Meta's Muse Spark launch table (Meta reports the better of vendor self-reports… 78.4
    Printed as 78.4%Vendor-reported, measured Apr 2026Configuration: Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
    Introducing Muse Spark: Scaling Towards Personal Superintelligence launch post, Meta, 8 Apr 2026. Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column Muse Spark Thinking; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'
    MedXpertQA (MM) | Medical Multiple Choice | 78.4 | 64.8 | 81.3 | 77.1 | 65.8 (image table row; columns Muse Spark Thinking, Opus 4.6 Max, Gemini 3.1 Pro High, GPT 5.4 Xhigh, Grok 4.2 Reasoning)
    Every result from this document
  6. 6GPT-5.4 Meta's Muse Spark launch table (Meta reports the better of vendor self-reports… 77.1
    Printed as 77.1%Independent run, measured Apr 2026Configuration: Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
    Introducing Muse Spark: Scaling Towards Personal Superintelligence launch post, Meta, 8 Apr 2026. Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column GPT 5.4 Xhigh; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'
    MedXpertQA (MM) | Medical Multiple Choice | 78.4 | 64.8 | 81.3 | 77.1 | 65.8 (image table row; columns Muse Spark Thinking, Opus 4.6 Max, Gemini 3.1 Pro High, GPT 5.4 Xhigh, Grok 4.2 Reasoning)
    Every result from this document
  7. 7Gemini-3 Pro Qwen3.5 model card comparison; source-specific evaluation, not harmonized acros… 76.0
    Printed as 76.0Independent run, measured Feb 2026Configuration: Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors
    Qwen/Qwen3.5-397B-A17B model card model card, Alibaba / Qwen, 16 Feb 2026. Benchmark Results > Vision Language > Medical VQA; MedXpertQA-MM row, Gemini-3 Pro column.
    | GPT5.2 | Claude 4.5 Opus | Gemini-3 Pro | Qwen3-VL-235B-A22B | K2.5-1T-A32B | Qwen3.5-397B-A17B MedXpertQA-MM | 73.3 | 63.6 | 76.0 | 47.6 | 65.3 | 70.0
    Every result from this document
  8. 8GPT-5.2 Qwen-run comparison in the Qwen3.5-397B-A17B model card 73.3
    Printed as 73.3Independent runConfiguration: Qwen-run comparison in the Qwen3.5-397B-A17B model card
    Qwen/Qwen3.5-397B-A17B model card model card, Alibaba / Qwen, 16 Feb 2026. Benchmark Results > Vision Language > Medical VQA, row MedXpertQA-MM; columns GPT5.2 / Claude 4.5 Opus / Gemini-3 Pro / Qwen3-VL-235B-A22B / K2.5-1T-A32B / Qwen3.5-397B-A17B
    MedXpertQA-MM 73.3 63.6 76.0 47.6 65.3 70.0
    Every result from this document
  9. 9Claude Opus 4.8 Qwen-run comparison in the Qwen3.8-Max launch post 71.7
    Printed as 71.7Independent runConfiguration: Qwen-run comparison in the Qwen3.8-Max launch post
    Qwen3.8-Max: A New Bar for Coding and Cowork launch post, Alibaba, 2 Aug 2026. Full Benchmark Table, second (multimodal) table, row MedXpertQA-MM; columns Opus4.8 / Fable5 / Gemini3.1-Pro / GPT5.6-Sol / Qwen3.7-Plus / Qwen3.8-Max
    | MedXpertQA-MM | 71.7 | 80.0 | 80.7 | 81.5 | 71.0 | 80.4 |
    Every result from this document
  10. 10Qwen3.7 Plus Alibaba's own Qwen3.7 Plus launch table; protocol not stated 71.0
    Printed as 71.0%Vendor-reported, measured May 2026Configuration: Alibaba's own Qwen3.7 Plus launch table; protocol not stated
    Qwen3.7-Plus: Multimodal Agent Intelligence launch post, Alibaba, 31 May 2026. Multimodal Benchmarks table, Multimodal Reasoning section, row MedXpertQA-MM, column Qwen3.7-Plus; header row: | | GPT-5.4 (xhigh) | Opus-4.6 Max | Gemini-3.1 Pro | Qwen3.6-Plus | Qwen3.7-Plus |
    | MedXpertQA-MM | 77.3 | 64.4 | 80.7 | 68.7 | 71.0 |
    Every result from this document
  11. 11Qwen3.5 397B A17B self-reported in the Qwen3.5-397B-A17B model card 70.0
    Printed as 70.0Vendor-reportedConfiguration: self-reported in the Qwen3.5-397B-A17B model card
    Qwen/Qwen3.5-397B-A17B model card model card, Alibaba / Qwen, 16 Feb 2026. Benchmark Results > Vision Language > Medical VQA, row MedXpertQA-MM; columns GPT5.2 / Claude 4.5 Opus / Gemini-3 Pro / Qwen3-VL-235B-A22B / K2.5-1T-A32B / Qwen3.5-397B-A17B
    MedXpertQA-MM 73.3 63.6 76.0 47.6 65.3 70.0
    Every result from this document
  12. 12Qwen3.6 Plus Qwen-run comparison in the Qwen3.7-Plus launch post 68.7
    Printed as 68.7Vendor-reportedConfiguration: Qwen-run comparison in the Qwen3.7-Plus launch post
    Qwen3.7-Plus: Multimodal Agent Intelligence launch post, Alibaba, 31 May 2026. Multimodal Benchmarks table, row MedXpertQA-MM; columns GPT-5.4 (xhigh) / Opus-4.6 Max / Gemini-3.1 Pro / Qwen3.6-Plus / Qwen3.7-Plus
    | MedXpertQA-MM | 77.3 | 64.4 | 80.7 | 68.7 | 71.0 |
    Every result from this document
  13. 13Grok 4.20 Meta's Muse Spark launch table (Meta reports the better of vendor self-reports… 65.8
    Printed as 65.8%Independent run, measured Apr 2026Configuration: Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
    Introducing Muse Spark: Scaling Towards Personal Superintelligence launch post, Meta, 8 Apr 2026. Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column Grok 4.2 Reasoning; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'
    MedXpertQA (MM) | Medical Multiple Choice | 78.4 | 64.8 | 81.3 | 77.1 | 65.8 (image table row; columns Muse Spark Thinking, Opus 4.6 Max, Gemini 3.1 Pro High, GPT 5.4 Xhigh, Grok 4.2 Reasoning)
    Every result from this document
  14. 14Kimi K2.5 Qwen-run comparison in the Qwen3.5-397B-A17B model card (K2.5-1T-A32B column) 65.3
    Printed as 65.3Independent runConfiguration: Qwen-run comparison in the Qwen3.5-397B-A17B model card (K2.5-1T-A32B column)
    Qwen/Qwen3.5-397B-A17B model card model card, Alibaba / Qwen, 16 Feb 2026. Benchmark Results > Vision Language > Medical VQA, row MedXpertQA-MM; columns GPT5.2 / Claude 4.5 Opus / Gemini-3 Pro / Qwen3-VL-235B-A22B / K2.5-1T-A32B / Qwen3.5-397B-A17B
    MedXpertQA-MM 73.3 63.6 76.0 47.6 65.3 70.0
    Every result from this document
  15. 15Claude Opus 4.6 Meta's Muse Spark launch table (Meta reports the better of vendor self-reports… 64.8
    Printed as 64.8%Independent run, measured Apr 2026Configuration: Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
    Introducing Muse Spark: Scaling Towards Personal Superintelligence launch post, Meta, 8 Apr 2026. Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column Opus 4.6 Max; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'
    MedXpertQA (MM) | Medical Multiple Choice | 78.4 | 64.8 | 81.3 | 77.1 | 65.8 (image table row; columns Muse Spark Thinking, Opus 4.6 Max, Gemini 3.1 Pro High, GPT 5.4 Xhigh, Grok 4.2 Reasoning)
    Every result from this document
  16. 16Claude 4.5 Opus Qwen3.5 model card comparison; source-specific evaluation, not harmonized acros… 63.6
    Printed as 63.6Independent run, measured Feb 2026Configuration: Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors
    Qwen/Qwen3.5-397B-A17B model card model card, Alibaba / Qwen, 16 Feb 2026. Benchmark Results > Vision Language > Medical VQA; MedXpertQA-MM row, Claude 4.5 Opus column.
    | GPT5.2 | Claude 4.5 Opus | Gemini-3 Pro | Qwen3-VL-235B-A22B | K2.5-1T-A32B | Qwen3.5-397B-A17B MedXpertQA-MM | 73.3 | 63.6 | 76.0 | 47.6 | 65.3 | 70.0
    Every result from this document
  17. 17Gemma 4 31B Google model card; MedXPertQA MM row; vendor-reported; protocol differs from ot… 61.3
    Printed as 61.3%Vendor-reported, measured Apr 2026Configuration: Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families
    Gemma 4 model card model card, Google, 2 Apr 2026. Evaluation Results, MedXPertQA MM row, Gemma 4 31B column.
    | Gemma 4 31B | Gemma 4 26B A4B | Gemma 4 12B Unified | Gemma 4 E4B | Gemma 4 E2B | Gemma 3 27B (no think) MedXPertQA MM | 61.3% | 58.1% | 48.7% | 28.7% | 23.5% | -
    Every result from this document
  18. 18Gemma 4 26B A4B Google model card; MedXPertQA MM row; vendor-reported; protocol differs from ot… 58.1
    Printed as 58.1%Vendor-reported, measured Apr 2026Configuration: Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families
    Gemma 4 model card model card, Google, 2 Apr 2026. Evaluation Results, MedXPertQA MM row, Gemma 4 26B A4B column.
    | Gemma 4 31B | Gemma 4 26B A4B | Gemma 4 12B Unified | Gemma 4 E4B | Gemma 4 E2B | Gemma 3 27B (no think) MedXPertQA MM | 61.3% | 58.1% | 48.7% | 28.7% | 23.5% | -
    Every result from this document
  19. 19Gemma 4 12B Google's Gemma 4 model card, Unified 12B; protocol not stated 48.7
    Printed as 48.7%Vendor-reported, measured Apr 2026Configuration: Google's Gemma 4 model card, Unified 12B; protocol not stated
    Gemma 4 model card model card, Google, 2 Apr 2026. Benchmark Results table, Vision section, row MedXPertQA MM, column Gemma 4 12B Unified; header row: | | Gemma 4 31B | Gemma 4 26B A4B | Gemma 4 12B Unified | Gemma 4 E4B | Gemma 4 E2B | Gemma 3 27B (no think) |
    | MedXPertQA MM | 61.3% | 58.1% | 48.7% | 28.7% | 23.5% | \- |
    Every result from this document
  20. 20Qwen3-VL-235B-A22B Qwen3.5 model card comparison; source-specific evaluation, not harmonized acros… 47.6
    Printed as 47.6Vendor-reported, measured Feb 2026Configuration: Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors
    Qwen/Qwen3.5-397B-A17B model card model card, Alibaba / Qwen, 16 Feb 2026. Benchmark Results > Vision Language > Medical VQA; MedXpertQA-MM row, Qwen3-VL-235B-A22B column.
    | GPT5.2 | Claude 4.5 Opus | Gemini-3 Pro | Qwen3-VL-235B-A22B | K2.5-1T-A32B | Qwen3.5-397B-A17B MedXpertQA-MM | 73.3 | 63.6 | 76.0 | 47.6 | 65.3 | 70.0
    Every result from this document
  21. 21Gemma 4 E4B Google model card; MedXPertQA MM row; vendor-reported; protocol differs from ot… 28.7
    Printed as 28.7%Vendor-reported, measured Apr 2026Configuration: Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families
    Gemma 4 model card model card, Google, 2 Apr 2026. Evaluation Results, MedXPertQA MM row, Gemma 4 E4B column.
    | Gemma 4 31B | Gemma 4 26B A4B | Gemma 4 12B Unified | Gemma 4 E4B | Gemma 4 E2B | Gemma 3 27B (no think) MedXPertQA MM | 61.3% | 58.1% | 48.7% | 28.7% | 23.5% | -
    Every result from this document
  22. 22Gemma 4 E2B Google model card; MedXPertQA MM row; vendor-reported; protocol differs from ot… 23.5
    Printed as 23.5%Vendor-reported, measured Apr 2026Configuration: Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families
    Gemma 4 model card model card, Google, 2 Apr 2026. Evaluation Results, MedXPertQA MM row, Gemma 4 E2B column.
    | Gemma 4 31B | Gemma 4 26B A4B | Gemma 4 12B Unified | Gemma 4 E4B | Gemma 4 E2B | Gemma 3 27B (no think) MedXPertQA MM | 61.3% | 58.1% | 48.7% | 28.7% | 23.5% | -
    Every result from this document

Documents

8