🤖 AI Summary
Short-term online experiments often fail to accurately predict long-term business outcomes, potentially leading to decisions misaligned with strategic objectives. Drawing on insights from industry expert workshops, this work proposes a design principle for surrogate metrics centered on decision utility rather than solely on unbiasedness, advocating that interpretable, experiment-driven simple surrogates outperform complex black-box models. Through expert consensus synthesis, surrogate metric analysis, and comparative evaluation of experimental versus observational data, the study systematically outlines a methodology for constructing effective surrogates and uncovers a critical relationship between the stability of long-term effects and surrogate validity. While underscoring the irreplaceable value of high-quality long-term experimentation, the research also delineates core challenges and practical guidelines for surrogate learning in real-world settings.
📝 Abstract
Online platforms prioritize long-term business outcomes, yet typical experiments are far too short to measure these outcomes directly. Our goal in this paper is to collect and share industry knowledge on how to make decisions from short-term experiments that are better aligned with long-term outcomes. Based on a daylong workshop with 26 experts from 15 online platforms and 4 universities, we formulate a series of propositions that reflect current industry knowledge.
Participants largely agreed that reversals of sign from short-run to long-run treatment effects are rare, with reversals concentrating in specific cases such as treatments involving content quality signals, hyper-monetization, and pricing. Although the magnitude of treatment effects can shift over time, a "univariate autosurrogate", corresponding to the short-run treatment effect on the long-run metric of interest, is often hard to beat. A recurring theme was the importance of surrogates that are not only (or even primarily) unbiased for true long-run outcomes, but that improve decision-making. Thus, participants generally agreed that simple, interpretable surrogates were generally preferable to elaborate but hard-to-explain surrogate indices. Participants also agreed that, due to concerns about confounding and transportability, experimentally-learned surrogates are generally preferable to observationally-learned surrogates. However, the drawback is that learning good surrogates from experiments typically requires a large, representative portfolio of long-run experiments that few platforms possess.
We conclude that there is no substitute for a well-run long-term experiment, whether for learning surrogates or validating them, and we highlight open challenges including evolving treatments, persistent treatments not fully mediated by short-term proxies, and mismatch between experimental samples and the target population.