MedXpertQA (MM)
Expert-level multiple-choice questions over clinical images across 17 specialties.
More about this boardLess
Expert-level multimodal medical multiple-choice QA covering clinical images (X-ray, histology, dermatology, charts) across 17 specialties; MM subset of the 4,460-question MedXpertQA benchmark.
Assembled table, not a single board. Rows come from several vendors' own launch documents: Meta's Muse Spark evaluation (whose values live only in a rendered table image and cover Meta's own model plus four competitors it re-ran or quoted), Alibaba's Qwen3.5, 3.7 and 3.8 posts (Alibaba's own runs of its models and of competitors), and Google's Gemma 4 model card. Protocols are not known to match across rows; each row's source line says which document it came from and who ran it. benchlm.ai presents a subset as one mirrored view while citing only Meta's methodology.
Published by TsinghuaC3I (Tsinghua University), released Jan 2025. 2,000 multimodal questions (MM subset). Percentage accuracy 0-100, higher better.
- Accuracy. Same 2025 paper; multimodal test under its original prompt and answer-extraction protocol. Unit: %.
- Board last updated
- 8 Apr 2026
- Paper version
- ICML 2025 / arXiv v3
- Rows
- 22
- Models
- 22 of 5on the board
- Labs
- 7
- Last measured
- Aug 2026
- Leader
- 81.5GPT-5.6 Sol
Ranking
Percentage accuracy 0-100, higher better
- Anthropic
- OpenAI
- Alibaba
- Meta
- Moonshot AI
- Other labs
20 of 22 rows
Rows and sources
Open a row for the quote, the page and the document.
1GPT-5.6 Sol Qwen-run comparison in the Qwen3.8-Max launch post 81.5
Printed as 81.5Independent runConfiguration: Qwen-run comparison in the Qwen3.8-Max launch postQwen3.8-Max: A New Bar for Coding and Cowork launch post, Alibaba, 2 Aug 2026. Full Benchmark Table, second (multimodal) table, row MedXpertQA-MM; columns Opus4.8 / Fable5 / Gemini3.1-Pro / GPT5.6-Sol / Qwen3.7-Plus / Qwen3.8-Max| MedXpertQA-MM | 71.7 | 80.0 | 80.7 | 81.5 | 71.0 | 80.4 |
Every result from this document2Gemini 3.1 Pro Meta's Muse Spark launch table (Meta reports the better of vendor self-reports… 81.3
Printed as 81.3%Independent run, measured Apr 2026Configuration: Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)Introducing Muse Spark: Scaling Towards Personal Superintelligence launch post, Meta, 8 Apr 2026. Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column Gemini 3.1 Pro High; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'MedXpertQA (MM) | Medical Multiple Choice | 78.4 | 64.8 | 81.3 | 77.1 | 65.8 (image table row; columns Muse Spark Thinking, Opus 4.6 Max, Gemini 3.1 Pro High, GPT 5.4 Xhigh, Grok 4.2 Reasoning)
Every result from this document- Also printed in Muse Spark Eval Methodology: 81.3, p. 5, benchmark table image, HEALTH section, row MedXpertQA (MM), column Gemini 3.1 Pro High
- Mirrored by MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai: 81.3%, BENCHMARK SCORE TABLE (8 MODELS), rank 1; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology
3Qwen3.8 Max Alibaba's own Qwen3.8 launch table; protocol not stated 80.4
Printed as 80.4%Vendor-reported, measured Aug 2026Configuration: Alibaba's own Qwen3.8 launch table; protocol not statedQwen3.8-Max: A New Bar for Coding and Cowork launch post, Alibaba, 2 Aug 2026. Multimodal Benchmarks table, Multimodal Reasoning section, row MedXpertQA-MM, column Qwen3.8-Max; header row: | | Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus | Qwen3.8-Max || MedXpertQA-MM | 71.7 | 80.0 | 80.7 | 81.5 | 71.0 | 80.4 |
Every result from this document- Mirrored by MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai: 80.4%, BENCHMARK SCORE TABLE (8 MODELS), rank 2; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology
4Claude Fable 5 Qwen-run comparison in the Qwen3.8-Max launch post 80.0
Printed as 80.0Independent runConfiguration: Qwen-run comparison in the Qwen3.8-Max launch postQwen3.8-Max: A New Bar for Coding and Cowork launch post, Alibaba, 2 Aug 2026. Full Benchmark Table, second (multimodal) table, row MedXpertQA-MM; columns Opus4.8 / Fable5 / Gemini3.1-Pro / GPT5.6-Sol / Qwen3.7-Plus / Qwen3.8-Max| MedXpertQA-MM | 71.7 | 80.0 | 80.7 | 81.5 | 71.0 | 80.4 |
Every result from this document5Muse Spark Meta's Muse Spark launch table (Meta reports the better of vendor self-reports… 78.4
Printed as 78.4%Vendor-reported, measured Apr 2026Configuration: Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)Introducing Muse Spark: Scaling Towards Personal Superintelligence launch post, Meta, 8 Apr 2026. Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column Muse Spark Thinking; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'MedXpertQA (MM) | Medical Multiple Choice | 78.4 | 64.8 | 81.3 | 77.1 | 65.8 (image table row; columns Muse Spark Thinking, Opus 4.6 Max, Gemini 3.1 Pro High, GPT 5.4 Xhigh, Grok 4.2 Reasoning)
Every result from this document- Also printed in Muse Spark Eval Methodology: 78.4, p. 5, benchmark table image, HEALTH section, row MedXpertQA (MM), column Muse Spark Thinking
- Mirrored by MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai: 78.4%, BENCHMARK SCORE TABLE (8 MODELS), rank 3; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology
6GPT-5.4 Meta's Muse Spark launch table (Meta reports the better of vendor self-reports… 77.1
Printed as 77.1%Independent run, measured Apr 2026Configuration: Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)Introducing Muse Spark: Scaling Towards Personal Superintelligence launch post, Meta, 8 Apr 2026. Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column GPT 5.4 Xhigh; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'MedXpertQA (MM) | Medical Multiple Choice | 78.4 | 64.8 | 81.3 | 77.1 | 65.8 (image table row; columns Muse Spark Thinking, Opus 4.6 Max, Gemini 3.1 Pro High, GPT 5.4 Xhigh, Grok 4.2 Reasoning)
Every result from this document- Also printed in Muse Spark Eval Methodology: 77.1, p. 5, benchmark table image, HEALTH section, row MedXpertQA (MM), column GPT 5.4 Xhigh
- Mirrored by MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai: 77.1%, BENCHMARK SCORE TABLE (8 MODELS), rank 4; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology
7Gemini-3 Pro Qwen3.5 model card comparison; source-specific evaluation, not harmonized acros… 76.0
Printed as 76.0Independent run, measured Feb 2026Configuration: Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendorsQwen/Qwen3.5-397B-A17B model card model card, Alibaba / Qwen, 16 Feb 2026. Benchmark Results > Vision Language > Medical VQA; MedXpertQA-MM row, Gemini-3 Pro column.| GPT5.2 | Claude 4.5 Opus | Gemini-3 Pro | Qwen3-VL-235B-A22B | K2.5-1T-A32B | Qwen3.5-397B-A17B MedXpertQA-MM | 73.3 | 63.6 | 76.0 | 47.6 | 65.3 | 70.0
Every result from this document8GPT-5.2 Qwen-run comparison in the Qwen3.5-397B-A17B model card 73.3
Printed as 73.3Independent runConfiguration: Qwen-run comparison in the Qwen3.5-397B-A17B model cardQwen/Qwen3.5-397B-A17B model card model card, Alibaba / Qwen, 16 Feb 2026. Benchmark Results > Vision Language > Medical VQA, row MedXpertQA-MM; columns GPT5.2 / Claude 4.5 Opus / Gemini-3 Pro / Qwen3-VL-235B-A22B / K2.5-1T-A32B / Qwen3.5-397B-A17BMedXpertQA-MM 73.3 63.6 76.0 47.6 65.3 70.0
Every result from this document9Claude Opus 4.8 Qwen-run comparison in the Qwen3.8-Max launch post 71.7
Printed as 71.7Independent runConfiguration: Qwen-run comparison in the Qwen3.8-Max launch postQwen3.8-Max: A New Bar for Coding and Cowork launch post, Alibaba, 2 Aug 2026. Full Benchmark Table, second (multimodal) table, row MedXpertQA-MM; columns Opus4.8 / Fable5 / Gemini3.1-Pro / GPT5.6-Sol / Qwen3.7-Plus / Qwen3.8-Max| MedXpertQA-MM | 71.7 | 80.0 | 80.7 | 81.5 | 71.0 | 80.4 |
Every result from this document10Qwen3.7 Plus Alibaba's own Qwen3.7 Plus launch table; protocol not stated 71.0
Printed as 71.0%Vendor-reported, measured May 2026Configuration: Alibaba's own Qwen3.7 Plus launch table; protocol not statedQwen3.7-Plus: Multimodal Agent Intelligence launch post, Alibaba, 31 May 2026. Multimodal Benchmarks table, Multimodal Reasoning section, row MedXpertQA-MM, column Qwen3.7-Plus; header row: | | GPT-5.4 (xhigh) | Opus-4.6 Max | Gemini-3.1 Pro | Qwen3.6-Plus | Qwen3.7-Plus || MedXpertQA-MM | 77.3 | 64.4 | 80.7 | 68.7 | 71.0 |
Every result from this document- Also printed in Qwen3.8-Max: A New Bar for Coding and Cowork: 71.0, Qwen3.8 launch post, Multimodal Benchmarks table, row MedXpertQA-MM, column Qwen3.7-Plus (same value carried forward)
- Mirrored by MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai: 71.0%, BENCHMARK SCORE TABLE (8 MODELS), rank 5; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology
11Qwen3.5 397B A17B self-reported in the Qwen3.5-397B-A17B model card 70.0
Printed as 70.0Vendor-reportedConfiguration: self-reported in the Qwen3.5-397B-A17B model cardQwen/Qwen3.5-397B-A17B model card model card, Alibaba / Qwen, 16 Feb 2026. Benchmark Results > Vision Language > Medical VQA, row MedXpertQA-MM; columns GPT5.2 / Claude 4.5 Opus / Gemini-3 Pro / Qwen3-VL-235B-A22B / K2.5-1T-A32B / Qwen3.5-397B-A17BMedXpertQA-MM 73.3 63.6 76.0 47.6 65.3 70.0
Every result from this document12Qwen3.6 Plus Qwen-run comparison in the Qwen3.7-Plus launch post 68.7
Printed as 68.7Vendor-reportedConfiguration: Qwen-run comparison in the Qwen3.7-Plus launch postQwen3.7-Plus: Multimodal Agent Intelligence launch post, Alibaba, 31 May 2026. Multimodal Benchmarks table, row MedXpertQA-MM; columns GPT-5.4 (xhigh) / Opus-4.6 Max / Gemini-3.1 Pro / Qwen3.6-Plus / Qwen3.7-Plus| MedXpertQA-MM | 77.3 | 64.4 | 80.7 | 68.7 | 71.0 |
Every result from this document13Grok 4.20 Meta's Muse Spark launch table (Meta reports the better of vendor self-reports… 65.8
Printed as 65.8%Independent run, measured Apr 2026Configuration: Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)Introducing Muse Spark: Scaling Towards Personal Superintelligence launch post, Meta, 8 Apr 2026. Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column Grok 4.2 Reasoning; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'MedXpertQA (MM) | Medical Multiple Choice | 78.4 | 64.8 | 81.3 | 77.1 | 65.8 (image table row; columns Muse Spark Thinking, Opus 4.6 Max, Gemini 3.1 Pro High, GPT 5.4 Xhigh, Grok 4.2 Reasoning)
Every result from this document- Also printed in Muse Spark Eval Methodology: 65.8, p. 5, benchmark table image, HEALTH section, row MedXpertQA (MM), column Grok 4.2 Reasoning
- Mirrored by MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai: 65.8%, BENCHMARK SCORE TABLE (8 MODELS), rank 6; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology
14Kimi K2.5 Qwen-run comparison in the Qwen3.5-397B-A17B model card (K2.5-1T-A32B column) 65.3
Printed as 65.3Independent runConfiguration: Qwen-run comparison in the Qwen3.5-397B-A17B model card (K2.5-1T-A32B column)Qwen/Qwen3.5-397B-A17B model card model card, Alibaba / Qwen, 16 Feb 2026. Benchmark Results > Vision Language > Medical VQA, row MedXpertQA-MM; columns GPT5.2 / Claude 4.5 Opus / Gemini-3 Pro / Qwen3-VL-235B-A22B / K2.5-1T-A32B / Qwen3.5-397B-A17BMedXpertQA-MM 73.3 63.6 76.0 47.6 65.3 70.0
Every result from this document15Claude Opus 4.6 Meta's Muse Spark launch table (Meta reports the better of vendor self-reports… 64.8
Printed as 64.8%Independent run, measured Apr 2026Configuration: Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)Introducing Muse Spark: Scaling Towards Personal Superintelligence launch post, Meta, 8 Apr 2026. Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column Opus 4.6 Max; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'MedXpertQA (MM) | Medical Multiple Choice | 78.4 | 64.8 | 81.3 | 77.1 | 65.8 (image table row; columns Muse Spark Thinking, Opus 4.6 Max, Gemini 3.1 Pro High, GPT 5.4 Xhigh, Grok 4.2 Reasoning)
Every result from this document- Also printed in Muse Spark Eval Methodology: 64.8, p. 5, benchmark table image, HEALTH section, row MedXpertQA (MM), column Opus 4.6 Max
- Mirrored by MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai: 64.8%, BENCHMARK SCORE TABLE (8 MODELS), rank 7; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology
16Claude 4.5 Opus Qwen3.5 model card comparison; source-specific evaluation, not harmonized acros… 63.6
Printed as 63.6Independent run, measured Feb 2026Configuration: Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendorsQwen/Qwen3.5-397B-A17B model card model card, Alibaba / Qwen, 16 Feb 2026. Benchmark Results > Vision Language > Medical VQA; MedXpertQA-MM row, Claude 4.5 Opus column.| GPT5.2 | Claude 4.5 Opus | Gemini-3 Pro | Qwen3-VL-235B-A22B | K2.5-1T-A32B | Qwen3.5-397B-A17B MedXpertQA-MM | 73.3 | 63.6 | 76.0 | 47.6 | 65.3 | 70.0
Every result from this document17Gemma 4 31B Google model card; MedXPertQA MM row; vendor-reported; protocol differs from ot… 61.3
Printed as 61.3%Vendor-reported, measured Apr 2026Configuration: Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source familiesGemma 4 model card model card, Google, 2 Apr 2026. Evaluation Results, MedXPertQA MM row, Gemma 4 31B column.| Gemma 4 31B | Gemma 4 26B A4B | Gemma 4 12B Unified | Gemma 4 E4B | Gemma 4 E2B | Gemma 3 27B (no think) MedXPertQA MM | 61.3% | 58.1% | 48.7% | 28.7% | 23.5% | -
Every result from this document18Gemma 4 26B A4B Google model card; MedXPertQA MM row; vendor-reported; protocol differs from ot… 58.1
Printed as 58.1%Vendor-reported, measured Apr 2026Configuration: Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source familiesGemma 4 model card model card, Google, 2 Apr 2026. Evaluation Results, MedXPertQA MM row, Gemma 4 26B A4B column.| Gemma 4 31B | Gemma 4 26B A4B | Gemma 4 12B Unified | Gemma 4 E4B | Gemma 4 E2B | Gemma 3 27B (no think) MedXPertQA MM | 61.3% | 58.1% | 48.7% | 28.7% | 23.5% | -
Every result from this document19Gemma 4 12B Google's Gemma 4 model card, Unified 12B; protocol not stated 48.7
Printed as 48.7%Vendor-reported, measured Apr 2026Configuration: Google's Gemma 4 model card, Unified 12B; protocol not statedGemma 4 model card model card, Google, 2 Apr 2026. Benchmark Results table, Vision section, row MedXPertQA MM, column Gemma 4 12B Unified; header row: | | Gemma 4 31B | Gemma 4 26B A4B | Gemma 4 12B Unified | Gemma 4 E4B | Gemma 4 E2B | Gemma 3 27B (no think) || MedXPertQA MM | 61.3% | 58.1% | 48.7% | 28.7% | 23.5% | \- |
Every result from this document- Also printed in Gemma 4 Technical Report: 48.7, Gemma 4 Technical Report p. 6, Table 6 (vision benchmarks, thinking), row MedXPertQA MM, column 12B
- Mirrored by MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai: 48.7%, BENCHMARK SCORE TABLE (8 MODELS), rank 8; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology
20Qwen3-VL-235B-A22B Qwen3.5 model card comparison; source-specific evaluation, not harmonized acros… 47.6
Printed as 47.6Vendor-reported, measured Feb 2026Configuration: Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendorsQwen/Qwen3.5-397B-A17B model card model card, Alibaba / Qwen, 16 Feb 2026. Benchmark Results > Vision Language > Medical VQA; MedXpertQA-MM row, Qwen3-VL-235B-A22B column.| GPT5.2 | Claude 4.5 Opus | Gemini-3 Pro | Qwen3-VL-235B-A22B | K2.5-1T-A32B | Qwen3.5-397B-A17B MedXpertQA-MM | 73.3 | 63.6 | 76.0 | 47.6 | 65.3 | 70.0
Every result from this document21Gemma 4 E4B Google model card; MedXPertQA MM row; vendor-reported; protocol differs from ot… 28.7
Printed as 28.7%Vendor-reported, measured Apr 2026Configuration: Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source familiesGemma 4 model card model card, Google, 2 Apr 2026. Evaluation Results, MedXPertQA MM row, Gemma 4 E4B column.| Gemma 4 31B | Gemma 4 26B A4B | Gemma 4 12B Unified | Gemma 4 E4B | Gemma 4 E2B | Gemma 3 27B (no think) MedXPertQA MM | 61.3% | 58.1% | 48.7% | 28.7% | 23.5% | -
Every result from this document22Gemma 4 E2B Google model card; MedXPertQA MM row; vendor-reported; protocol differs from ot… 23.5
Printed as 23.5%Vendor-reported, measured Apr 2026Configuration: Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source familiesGemma 4 model card model card, Google, 2 Apr 2026. Evaluation Results, MedXPertQA MM row, Gemma 4 E2B column.| Gemma 4 31B | Gemma 4 26B A4B | Gemma 4 12B Unified | Gemma 4 E4B | Gemma 4 E2B | Gemma 3 27B (no think) MedXPertQA MM | 61.3% | 58.1% | 48.7% | 28.7% | 23.5% | -
Every result from this document
Documents
8
- Gemma 4 model cardmodel card, Google, 2 Apr 2026Results it supports
- Gemma 4 Technical Reportpaper, Google DeepMindResults it supports
- Introducing Muse Spark: Scaling Towards Personal Superintelligencelaunch post, Meta, 8 Apr 2026Results it supports
- MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.aimirror, benchlm.aiResults it supports
- Muse Spark Eval Methodologymodel card, Meta, 8 Apr 2026Results it supports
- Qwen/Qwen3.5-397B-A17B model cardmodel card, Alibaba / Qwen, 16 Feb 2026Results it supports
- Qwen3.7-Plus: Multimodal Agent Intelligencelaunch post, Alibaba, 31 May 2026Results it supports
- Qwen3.8-Max: A New Bar for Coding and Coworklaunch post, Alibaba, 2 Aug 2026Results it supports