Batch Size or Negatives? A Selection Rule for Memory-Constrained Recommender Training

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unresolved trade-off between batch size and the number of negative samples under fixed memory budgets in training memory-constrained recommender systems. It theoretically and empirically demonstrates, for the first time, that when using sampled Softmax, prioritizing larger batch sizes over a greater number of negative samples yields faster convergence and superior recommendation quality within the same memory constraints. The proposed configuration principle is validated across four real-world sequential recommendation benchmarks—including MovieLens-20M—as well as synthetic data, offering clear guidance for efficient training in resource-limited scenarios.
📝 Abstract
Large-scale neural recommender systems are typically trained with a softmax cross-entropy objective over the full item vocabulary. For a typical large number of possible items $K$, the final classification layer dominates memory, requiring $O(nK)$ logits and gradients to materialize for a batch of $n$ examples. Sampled softmax reduces this cost by restricting the objective to only $k \ll K$ candidate negative items, resulting in an $O(nk)$ memory. However, for a fixed budget $B = n k$, it remains unclear whether one should prioritize larger batches or the inclusion of more negative items. We address this question by analyzing sampled-softmax training under a fixed memory constraint. Under standard smoothness and variance assumptions, our theoretical evidence suggests that the fastest convergence arises from an $ n \sim B, k \sim 1$ allocation. So, an actionable rule is to include as many objects as possible given computational constraints. Our theory is supported by controlled synthetic and synthetic and four real sequential recommendation benchmarks, including MovieLens-20M. The suggested configuration achieve faster convergence and better final recommendation quality than imbalanced alternatives within the same memory constraint. These findings provide a theoretical and empirical foundation for configuring memory during the training of recommender systems. Code, reproducibility materials, and all scripts for generating figures are available at https://anonymous.4open.science/r/LimitedMemoryRule-BBFB
Problem

Research questions and friction points this paper is trying to address.

recommender systems
sampled softmax
memory constraint
batch size
negative sampling
Innovation

Methods, ideas, or system contributions that make the work stand out.

sampled softmax
memory-constrained training
batch size vs negatives
convergence analysis
recommender systems