Medical AI Improved Diagnosis, but Its Benefits Varied by User Expertise
An MIT-led study found that AI assistance generally improved skin-disease diagnosis for both non-experts and clinicians, but the groups used the tools differently. Non-experts often deferred to AI-generated explanations—even when they were wrong—while clinicians were more likely to identify incorrect assistance.

AI assistance improved performance on skin-disease diagnosis tasks for both non-experts and clinicians in a new study, but the benefits depended strongly on users’ medical expertise and how explanations were presented. The research found that non-experts frequently trusted AI recommendations even when they were incorrect, while clinicians were more likely to catch the system’s mistakes.
The study, described by MIT News and published August 4 in Nature Medicine, examined how different explainable-AI approaches affected primary care providers and people without medical expertise. Participants reviewed medical images alongside an AI prediction and, depending on the test condition, received a confidence level without an explanation, similar images, a heat map highlighting image regions, or a plain-language explanation generated by a large language model.
Non-experts were asked to determine whether skin moles were cancerous. Clinicians faced a more demanding task: producing a differential diagnosis for dermatological disease. All of the explainability approaches improved non-experts’ accuracy, with much of the improvement coming from better identification of noncancerous moles. A fairness-constrained model designed to address bias related to darker skin tones also significantly improved accuracy and reduced diagnostic disparities based on skin tone in the study.
The researchers said, however, that higher accuracy among non-experts was largely explained by their willingness to follow the model. When the AI was correct, that reliance helped. When it was wrong, it harmed performance more substantially. The effect was strongest with large-language-model explanations, which also made non-experts more confident in incorrect answers. The study reported that vague or generic explanations could appear more convincing to these users.
Clinicians responded differently. They were not substantially derailed by incorrect AI assistance, and they performed best when given the model’s prediction without an accompanying explanation. The researchers said clinicians could compare the AI output with a diagnosis formed from their own training, whereas non-experts could use the explanation as the starting point for their judgment.
Timing also influenced reliance. Participants became more deferential to the AI when its explanation appeared before they had made an independent diagnosis. Users who relied most heavily on AI were also the weakest performers when completing the task without assistance. The researchers further found that AI systems performed better than humans when disease presentations were subtle, while humans performed better when images contained atypical symptoms or unrelated features.
The findings suggest that explainability is not automatically beneficial simply because it makes an AI recommendation easier to understand. The same explanation may support a trained clinician but mislead a beginner, particularly when the system is wrong. The researchers proposed that future interfaces could first require users to offer their own diagnostic hypothesis, then use AI to surface alternative conditions or overlooked possibilities rather than simply supplying a persuasive rationale.
The study’s focus was dermatological diagnosis, so the results do not establish how the same interface designs would perform in other medical settings. Its broader contribution is to show that the value and risks of medical AI assistance depend not only on model accuracy, but also on users’ existing expertise and their interaction with the system.
Reporting Note
Biohack Report distinguishes preliminary findings, clinical evidence and commercial claims whenever the available reporting supports that distinction. Coverage is informational and is not medical advice.
