Hide and Seek in Noise Labels: Noise-Robust Collaborative Active Learning with LLMs-Powered Assistance

πŸ“… 2025-04-03
πŸ›οΈ Annual Meeting of the Association for Computational Linguistics
πŸ“ˆ Citations: 5
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
To address the challenge of accurately identifying and correcting mislabeled samples in learning with noisy labels, this paper proposes NoiseALβ€”a novel framework that achieves fine-grained separation of clean and noisy samples via dual lightweight model co-prediction and a dynamically adjusted confidence threshold. It further introduces an LLM-driven active labeling mechanism for semantic-level correction of noisy labels. Innovatively, we establish a hierarchical collaborative learning paradigm for noisy data and design subset-specific multi-objective optimization (employing CE, GCE, and SCE losses) tailored to varying sample quality. Extensive experiments on both synthetic and real-world noisy benchmarks demonstrate that NoiseAL improves noise robustness by 12.7% over state-of-the-art methods and reduces human annotation cost by over 40%, thereby overcoming the coarse-grained partitioning limitation inherent in conventional label-noise learning approaches.

Technology Category

Application Category

πŸ“ Abstract
Learning from noisy labels (LNL) is a challenge that arises in many real-world scenarios where collected training data can contain incorrect or corrupted labels. Most existing solutions identify noisy labels and adopt active learning to query human experts on them for denoising. In the era of large language models (LLMs), although we can reduce the human effort to improve these methods, their performances are still subject to accurately separating the clean and noisy samples from noisy data. In this paper, we propose an innovative collaborative learning framework NoiseAL based on active learning to combine LLMs and small models (SMs) for learning from noisy labels. During collaborative training, we first adopt two SMs to form a co-prediction network and propose a dynamic-enhanced threshold strategy to divide the noisy data into different subsets, then select the clean and noisy samples from these subsets to feed the active annotator LLMs to rectify noisy samples. Finally, we employ different optimization objectives to conquer subsets with different degrees of label noises. Extensive experiments on synthetic and real-world noise datasets further demonstrate the superiority of our framework over state-of-the-art baselines.
Problem

Research questions and friction points this paper is trying to address.

Develop a noise-robust collaborative learning framework using LLMs and small models
Separate clean and noisy samples dynamically from noisy label datasets
Optimize learning objectives for subsets with varying noise levels
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-powered active learning for noise reduction
Dynamic-enhanced threshold for data subset division
Dual optimization for varying noise levels
πŸ”Ž Similar Papers
No similar papers found.
Bo Yuan
Bo Yuan
PhD Student in Machine Learning, Georgia Institute of Technology
Markov chain Monte CarloLarge Language Model
Y
Yulin Chen
Zhejiang University, Hangzhou, China
Y
Yin Zhang
Zhejiang University, Hangzhou, China
W
Wei Jiang
Ant Group, Hangzhou, China