Knowledge and exam benchmarks
1 tracked · updated August 16, 2026
Exam-style medical QA is mostly a solved category: MedQA and its siblings saturated above 95 percent and were retired from tracking. What remains current is the expert multimodal tier, where questions carry clinical images across specialties and frontier models still leave headroom.
MedXpertQA (MM)
Tsinghua University · 2,000 multimodal questions- 1
Gemini 3.1 Pro81.3%
- 2AQwen3.8 Max80.4%
- 3
Muse Spark78.4%
- 4
GPT-5.477.1%
- 5AQwen3.7 Plus71.0%
Which knowledge and exam benchmarks have current frontier-model results?
1 as of August 16, 2026: MedXpertQA (MM) (Gemini 3.1 Pro leads at 81.3%).
The other categories sit on the index: rubric-graded benchmarks, agentic and workflow benchmarks, documentation and coding benchmarks, safety benchmarks, composite indices.