OrdRankBen: A Novel Ranking Benchmark for Ordinal Relevance in NLP
NLP ranking evaluation has long suffered from coarse-grained relevance labeling: binary labels fail to distinguish degrees of relevance, while continuous scores lack explicit ordinal structure, hindering fine-grained discrimination. To address this, we introduce OrdRankBen—the first NLP ranking benchmark explicitly designed for ordinal relevance. Its core innovation lies in structured ordinal relevance labels (e.g., “strongly relevant > moderately relevant > weakly relevant > irrelevant”) and two real-world datasets capturing diverse ordinal label distributions. We employ a hybrid construction methodology combining human annotation with controlled distribution sampling, enabling unified evaluation of ranking-specific LMs, general-purpose LLMs, and dedicated ranking LLMs. Experiments demonstrate that ordinal modeling significantly enhances model sensitivity to subtle relevance distinctions, yielding more precise and robust ranking performance characterization across diverse model architectures.