Doomed to Re-Annotate, Forever: The ImageNet Story

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses label noise, ambiguous definitions, and missing multi-label annotations in ImageNet-1K by proposing a human-AI collaborative iterative refinement paradigm to construct the ReImageNet dataset. Through joint human and large language model efforts in correcting labels and supplementing semantic attributes, this work reveals an approximate 12% error rate in the original data, demonstrating the limitations of single-pass annotation. Experiments indicate that this refined benchmark improves supervised model Top-1 accuracy by 1.2% and multimodal large language models by 5–6%, significantly enhancing evaluation reliability. By open-sourcing the complete dataset and codebase, this project establishes a new standard for quality in large-scale visual annotation.
📝 Abstract
Top-1 accuracy on ImageNet-1k remains the most commonly reported metric in visual recognition. Quality issues with the dataset have been repeatedly reported, yet the original 2012 noisy labels are still predominantly used. The paper presents a comprehensive effort, which goes well beyond prior correction attempts, towards obtaining accurate and complete ImageNet-1k validation set annotations. The result, ReImageNet, includes multilabel correction, object localization, revised class definitions, and semantic attributes (text-recognition, rendition, reflection, crowd, dominant). The reannotation reveals that approximately 12% of the original ImageNet-1k labels are incorrect, 33.3% of images are multilabel and 3.8% contain no object from an ImageNet-1k class. With the new labels, top-1 accuracy increases by up to 1.2% for supervised models and by 5-6% for MLLMs. We argue that annotation at ImageNet scale cannot realistically be completed in one pass, as errors and definitional issues are discovered only through annotating, and we build our pipeline around repeated refinement and error checking. We observed that human and LLM collaboration with appropriate tooling represents the current quality ceiling for annotation at this scale. ImageNet-1k issues propagate into its derivative test sets, indicating that the problem is structural rather than specific to any single benchmark. All annotations, class definitions, guidelines, and analysis code have been publicly released. Project page: https://vrg.fel.cvut.cz/reimagenet Annotations: https://huggingface.co/datasets/vrg-prague/ReImageNet Code: https://github.com/klarajanouskova/ImageNet
Problem

Research questions and friction points this paper is trying to address.

ImageNet-1k
Label Noise
Data Quality
Re-annotation
Visual Recognition Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

ReImageNet
Human-LLM Collaboration
Iterative Refinement
Label Correction
Semantic Attributes
💼 Related Jobs
No related jobs found.
Illia Volkov
Illia Volkov
Unknown affiliation
N
Nikita Kisel
Visual Recognition Group, Faculty of Electrical Engineering, Czech Technical University in Prague
T
Tetiana Mishkina
Visual Recognition Group, Faculty of Electrical Engineering, Czech Technical University in Prague
K
Klara Janouskova
Visual Recognition Group, Faculty of Electrical Engineering, Czech Technical University in Prague
Jiri Matas
Jiri Matas
Professor, Czech Technical University
computer visionimage processingpattern recognitionmachine learning