Internal Medicine · 2 h ago
MedGemma models improve medical image and text benchmark performance
A Nature Medicine benchmark study evaluated MedGemma, a collection of medically tuned vision-language models based on Gemma 3. The models improved performance on medical image question answering and chest X-ray classification compared with base models, but these results do not establish clinical effectiveness.
- MedGemma was evaluated across medical text and imaging benchmarks.
- Out-of-distribution chest X-ray classification improved 15.5–18.1% over base models.
- Medical fine-tuning showed advantages when training data were limited.
- Benchmark results do not establish clinical effectiveness.
Researchers reporting in Nature Medicine developed and evaluated MedGemma, a collection of open medical artificial intelligence models based on Gemma 3. The collection includes 4-billion- and 27-billion-parameter multimodal models and a 27-billion-parameter text-only model. Evaluations covered medical question answering, image classification, clinical reasoning, report generation and agentic tasks, using both held-out data from training-related datasets and benchmarks not used for training or tuning.
On out-of-distribution tasks, the authors reported improvements over base models of 2.6–10% for medical image question answering, 15.5–18.1% for chest X-ray finding classification and 10.8% for agentic evaluations. MedGemma achieved a score of 86.2 on MedQA, compared with 78.2 for OpenBioLLM 70B and 91.0 for DeepSeek-R1. Fine-tuning MedGemma could outperform fine-tuning base Gemma 3 for medical tasks, particularly with limited training data.
The researchers also introduced MedSigLIP, a 400-million-parameter medical image encoder trained using millions of image–text pairs. These models may support development of specialized medical applications, but the reported findings concern benchmark performance rather than patient outcomes or prospective clinical deployment. Some evaluations used datasets represented during training or tuning, and the authors note that larger frontier models still outperform smaller open models on some tasks.
Is this summary clinically accurate?
Help fellow clinicians: your rating sends inaccurate summaries straight to our editors.
Sign in to rate this summary →Source
Nature Medicine: An open vision-language model for diverse medical applications ↗This is an automated AI-condensed summary that has not yet been reviewed by an editor. Always consult the full item at the original source.
