Clinical Benchmarks

MedXpertQA (MM): current results

TsinghuaC3I (Tsinghua University) · 2,000 multimodal questions (MM subset) · index updated August 16, 2026

Gemini 3.1 Pro holds the top current result on MedXpertQA (MM), 81.3% as of 2026-08, per benchlm.ai mirror of Meta's Muse Spark evaluation. Expert-level multimodal medical multiple-choice QA covering clinical images (X-ray, histology, dermatology, charts) across 17 specialties; MM subset of the 4,460-question MedXpertQA benchmark.

Current results

Result detail

#modelscoreas of
1Google logoGemini 3.1 Pro Google
as reported in Meta's Muse Spark eval, mirrored by benchlm
81.3%2026-08
2AQwen3.8 Max Alibaba
same source
80.4%2026-08
3Meta logoMuse Spark Meta
same source
78.4%2026-08
4OpenAI logoGPT-5.4 OpenAI
same source
77.1%2026-08
5AQwen3.7 Plus Alibaba
same source
71.0%2026-08
6xAI logoGrok 4.20 xAI
same source
65.8%2026-08
7Anthropic logoClaude Opus 4.6 Anthropic
same source
64.8%2026-08
8Google logoGemma 4 12B Google
same source
48.7%2026-08

Scores appear exactly as benchlm.ai mirror of Meta's Muse Spark evaluation publishes them (vendor-reported scores). The current frontier table is vendor-reported, from Meta's Muse Spark launch evaluation, mirrored display-only on benchlm.ai. The Text subset has no comparable current table.

About the benchmark

publisherTsinghuaC3I (Tsinghua University)
categoryknowledge and exam benchmarks
released2025-01
size2,000 multimodal questions (MM subset)
scalepercentage accuracy 0-100, higher better
result basisvendor-reported scores
sourcebenchlm.ai mirror of Meta's Muse Spark evaluation
last frontier result2026-08

What is MedXpertQA (MM)?

MedXpertQA (MM) is a knowledge and exam benchmark from Tsinghua University, released 2025-01: 2,000 multimodal questions (MM subset), scored on a percentage accuracy 0-100 scale. Expert-level multimodal medical multiple-choice QA covering clinical images (X-ray, histology, dermatology, charts) across 17 specialties; MM subset of the 4,460-question MedXpertQA benchmark.

Which model leads MedXpertQA (MM)?

Gemini 3.1 Pro (Google) holds the top current result on MedXpertQA (MM) at 81.3%, per benchlm.ai mirror of Meta's Muse Spark evaluation, as of 2026-08.

Where do the MedXpertQA (MM) numbers come from?

From benchlm.ai mirror of Meta's Muse Spark evaluation (vendor-reported scores). The current frontier table is vendor-reported, from Meta's Muse Spark launch evaluation, mirrored display-only on benchlm.ai. The Text subset has no comparable current table.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.