TrueSkin: Towards Fair and Accurate Skin Tone Recognition and Generation

๐Ÿ“… 2025-09-13
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Skin tone recognition and generation face core challenges including data scarcity, insufficient model robustness, and algorithmic biasโ€”particularly systematic misclassification of intermediate skin tones and interference from irrelevant attributes (e.g., hairstyle, background). To address these, we introduce TrueSkin, the first systematically annotated, multi-condition skin tone benchmark comprising 7,299 real-world images across six tone categories, captured under diverse lighting conditions, viewpoints, and imaging devices. Leveraging TrueSkin, we develop a supervised classification framework and a fine-tuning pipeline for generative models. Experiments demonstrate that our recognition model achieves over 20% higher accuracy than state-of-the-art methods; our generation framework significantly improves target skin tone fidelity while mitigating bias induced by confounding attributes. Critically, TrueSkin reveals, for the first time, systematic bias of large foundation models toward intermediate skin tones. It establishes a reproducible benchmark and methodological foundation for fair, robust skin tone perception and synthesis.

Technology Category

Application Category

๐Ÿ“ Abstract
Skin tone recognition and generation play important roles in model fairness, healthcare, and generative AI, yet they remain challenging due to the lack of comprehensive datasets and robust methodologies. Compared to other human image analysis tasks, state-of-the-art large multimodal models (LMMs) and image generation models struggle to recognize and synthesize skin tones accurately. To address this, we introduce TrueSkin, a dataset with 7299 images systematically categorized into 6 classes, collected under diverse lighting conditions, camera angles, and capture settings. Using TrueSkin, we benchmark existing recognition and generation approaches, revealing substantial biases: LMMs tend to misclassify intermediate skin tones as lighter ones, whereas generative models struggle to accurately produce specified skin tones when influenced by inherent biases from unrelated attributes in the prompts, such as hairstyle or environmental context. We further demonstrate that training a recognition model on TrueSkin improves classification accuracy by more than 20% compared to LMMs and conventional approaches, and fine-tuning with TrueSkin significantly improves skin tone fidelity in image generation models. Our findings highlight the need for comprehensive datasets like TrueSkin, which not only serves as a benchmark for evaluating existing models but also provides a valuable training resource to enhance fairness and accuracy in skin tone recognition and generation tasks.
Problem

Research questions and friction points this paper is trying to address.

Addressing skin tone recognition and generation fairness gaps
Overcoming dataset limitations for accurate skin tone analysis
Mitigating biases in multimodal and generative AI models
Innovation

Methods, ideas, or system contributions that make the work stand out.

TrueSkin dataset with 7299 categorized images
Benchmarking recognition and generation model biases
Training improves accuracy by 20% and fidelity
๐Ÿ”Ž Similar Papers
No similar papers found.
H
Haoming Lu
Topaz Labs, Dallas, United States