FairFlow: Mitigating Dataset Biases through Undecided Learning for Natural Language Understanding
Language models are vulnerable to dataset biases—such as shortcut learning and spurious correlations—leading to degraded cross-domain generalization and compromised fairness. To address this, we propose *Undecided Learning*, a novel paradigm that actively rejects high-confidence yet potentially biased predictions via an uncertainty-driven prediction suppression mechanism. Methodologically, we design a bias-aware multi-view generative and contrastive learning framework, integrating both data- and model-level perturbations to jointly model and suppress both known and unknown biases in a unified manner. Empirically, our approach significantly outperforms existing debiasing methods on cross-domain transfer and challenging sample benchmarks, while preserving source-domain performance. Notably, it achieves the first demonstrated robust mitigation of *unknown* biases—without requiring prior knowledge or bias annotations—thereby offering a principled pathway toward improved model generalization and fairness.