Back to All-Entity Ranking: Sampler-Dependent Evaluation in Continuous-Time Dynamic Graphs

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the instability in evaluating next-destination prediction models on continuous-time dynamic graphs, where negative sampling introduces variance and biases due to its dependence on sampling distribution and candidate set size. To mitigate this, the authors propose replacing negative sampling with full-entity ranking, thereby eliminating sampling-induced degrees of freedom and enabling more reliable performance comparisons within a fixed target space. Through theoretical analysis, they reveal how negative sampling perturbs Bayesian optimal ranking and systematically investigate interactions between repeated versus novel positive samples and seen versus unseen negative samples using factorized evaluation, a minimalistic history-pair scoring module, and controlled representation interventions. Experiments across LastFM, MOOC, Reddit, and Wikipedia datasets show that in at least three datasets, model rankings reverse with changes in candidate sets, and the direction and magnitude of modular effects shift significantly.
📝 Abstract
Next-destination prediction in continuous-time dynamic graphs (CTDGs) commonly ranks an observed interaction against sampled negative destinations. The resulting score is conditional on both the negative distribution and the number of candidates chosen by the researcher. We show that a non-uniform negative distribution changes the Bayes-optimal ranking, while even a finite candidate set drawn uniformly can destabilize model rankings and measured module effects. Time-varying source-destination history membership and model operations that use this information directly transmit the sampler's influence to the evaluation score. We examine this mechanism using a factorial evaluation of repeated and new positives against seen and unseen negatives, a minimal scorer based solely on pair-history membership, and controlled representation interventions. Across six models on LastFM, MOOC, Reddit, and Wikipedia, at least one model pair changes relative order between the expected Uniform-20 metric and the full catalog on three of the four datasets. The measured effect of the same module also changes in magnitude and direction with the candidate-set size and training objective. These results establish that model-superiority and ablation conclusions from sampled-negative benchmarks are conditional on the stated candidate configuration. All-entity ranking evaluates every destination in a fixed catalog, eliminating negative-selection freedom and sampling variation while retaining the original CTDG scorer. We therefore recommend all-entity ranking as the primary evidence for architecture comparisons on CTDG benchmarks with an enumerable, fixed destination catalog.
Problem

Research questions and friction points this paper is trying to address.

continuous-time dynamic graphs
next-destination prediction
negative sampling
evaluation bias
all-entity ranking
Innovation

Methods, ideas, or system contributions that make the work stand out.

all-entity ranking
continuous-time dynamic graphs
negative sampling bias
next-destination prediction
evaluation robustness
🔎 Similar Papers
No similar papers found.