
Real-World Study Finds Experienced Dermatologists Outperform AI in Skin Cancer Diagnosis
Key Takeaways
- Experts led multiclass classification across nine categories (74.2%), exceeding the physician mean (65.9%), with image-only PanDerm comparable to mid-career performance and superior to novices.
- Unimodal PanDerm achieved the best benign–malignant balanced accuracy (0.82) and high specificity (94%), whereas clinicians traded specificity for sensitivity to minimize missed malignancies.
A JAMA Dermatology study found expert dermatologists achieved the highest multiclass diagnostic accuracy when tested against AI across a realistic spectrum of skin lesions.
Expert dermatologists with more than 10 years' experience outperformed 3
AI models have previously
The TDIV dataset paired clinical and dermoscopic images with patient history, demographics, and risk factors across 1,117 cases, deliberately including
Experienced Dermatologists Achieved the Highest Multiclass Accuracy
Experts remained the highest-performing group, achieving 74.2% multiclass accuracy.1 The unimodal, image-only configuration of PanDerm reached 72.2% accuracy, statistically comparable to dermatologists with 3 to 10 years' experience and higher than readers with less than 1 year of experience, who scored 59.1%.1 The CNN trailed every physician group at 56.7% accuracy, the lowest of any system evaluated. Overall, the 652 participating physicians achieved a mean multiclass accuracy of 65.9%, below the top-performing AI configuration but above both the CNN and the multimodal PanDerm model.3
AI Excelled in Binary Classification but Showed Limitations
For the simpler task of distinguishing benign from malignant lesions, the unimodal PanDerm model achieved the highest balanced accuracy, at 0.82, compared with 0.65 for physicians overall.1 This advantage was driven largely by specificity: the unimodal model correctly classified benign lesions in 94% of cases, while the multimodal version, which incorporated
Adding clinical context reduced rather than improved PanDerm's multiclass performance, with the multimodal configuration achieving 66.3% accuracy compared with 72.2% for the unimodal version.12 The authors attributed part of this decline to a distribution shift between the model's training images and the more complex clinical photographs included in the TDIV dataset, along with an apparent underrepresentation of acral
Authors Highlight AI's Role as Clinical Decision Support
The investigators emphasized AI's role as a clinical support tool rather than a replacement for physician expertise. Anriot and Thomas wrote, "AI systems demonstrate strong potential as diagnostic support tools, particularly for early-career clinicians."1
The authors suggested
No regulatory action accompanies this diagnostic study, and the findings remain investigational pending validation in larger, more diverse patient populations.2
REFERENCES:
- Anriot, J., Yan, S., Coste, C., Tschandl, P., Verlingue, L., Andremasse, C., Amini-Adle, M., Perrot, J. L., Ge, Z., Kittler, H., & Thomas, L. (2026). Limits of Artificial Intelligence Models for Skin Cancer Diagnosis in Realistic Settings. JAMA dermatology, 162(7), 701–708.
https://doi.org/10.1001/jamadermatol.2026.149 - Analysis: AI model outperforms early-career physicians in skin lesion diagnosis. Practical Dermatology. June 4, 2026.
https://practicaldermatology.com/news/analysis-ai-model-outperforms-early-career-physicians-in-skin-lesion-diagnosis/2487332/ - Yan, S., Yu, Z., Primiero, C. et al. A multimodal vision foundation model for clinical dermatology. Nat Med 31, 2691–2702 (2025). https://doi.org/10.1038/s41591-025-03747-y
https://www.nature.com/articles/s41591-025-03747-y#citeas












