Evaluating for the long term: Learnings from industry
Short-term online experiments often fail to accurately predict long-term business outcomes, potentially leading to decisions misaligned with strategic objectives. Drawing on insights from industry expert workshops, this work proposes a design principle for surrogate metrics centered on decision utility rather than solely on unbiasedness, advocating that interpretable, experiment-driven simple surrogates outperform complex black-box models. Through expert consensus synthesis, surrogate metric analysis, and comparative evaluation of experimental versus observational data, the study systematically outlines a methodology for constructing effective surrogates and uncovers a critical relationship between the stability of long-term effects and surrogate validity. While underscoring the irreplaceable value of high-quality long-term experimentation, the research also delineates core challenges and practical guidelines for surrogate learning in real-world settings.