🤖 AI Summary
Clinical grading of atopic dermatitis (AD) suffers from high subjectivity and low inter-rater agreement among dermatologists. Method: This study systematically evaluates, for the first time, the applicability of seven state-of-the-art vision-language models (VLMs)—including CLIP, Flamingo, and Kosmos-2—to automated, objective AD severity quantification. Leveraging medical domain–specific prompt engineering and zero-shot/few-shot inference, we establish an interpretable, multimodal assessment framework. Results: Multiple VLMs demonstrate superior cross-image stability and clinical plausibility compared to conventional CNNs on public AD datasets. Their severity scores achieve substantial inter-rater agreement with dermatologist panels (Cohen’s κ = 0.72–0.81), validating VLMs as promising tools for objective dermatological quantification. This work provides both methodological innovation and empirical evidence supporting AI-assisted diagnosis and management of skin diseases.
📝 Abstract
The task of grading atopic dermatitis (or AD, a form of eczema) from patient images is difficult even for trained dermatologists. Research on automating this task has progressed in recent years with the development of deep learning solutions; however, the rapid evolution of multimodal models and more specifically vision-language models (VLMs) opens the door to new possibilities in terms of explainable assessment of medical images, including dermatology. This report describes experiments carried out to evaluate the ability of seven VLMs to assess the severity of AD on a set of test images.