Learning the Lake: Reliable Experience for Adaptive Data Product Discovery

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过建立Evolving Discovery Memory和SafeLake方法,利用可靠的经验数据减少数据产品发现过程中的搜索工作量,从而提高效率。
📝 Abstract
Data-product discovery searches a full lake even when workloads revisit related products and regions. Repetition permits contracted search, but similarity cannot justify a route because one omitted asset invalidates a conjunctive product. We study when serving experience can safely reduce this work. Evolving Discovery Memory records source-labelled query--product--region evidence above a fixed regional index. SafeLake separates operational familiarity, which determines how much to search, from independently calibrated product evidence, which determines where to search. The fixed-probe comparison holds the adaptive budget constant between SafeLake and Familiarity-only. On TAT-QA, product steering raises Product Recall by 0.072; ConvFinQA shows no resolved map gain, while the HybridQA sensitivity favors Familiarity-only in Full R@100. Trace-only, missing, and false feedback expose boundaries on map steering, while scope-audit agreement cannot certify the source. Across clean confirmed-feedback streams under the frozen transductive protocol, the formal controller saves 49.5--82.7% of cumulative asset exposure. Experience determines when to contract; reliable evidence determines where to contract.
Problem

Research questions and friction points this paper is trying to address.

Data-product discovery
Evolving Discovery Memory
SafeLake
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evolving Discovery Memory
SafeLake
Operational Familiarity
Product Evidence
Cumulative Asset Exposure
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yixi Zhou
Hong Kong Baptist University
F
Fan Zhang
The University of Tokyo
S
Sikun Wang
Tokyo University of Science
Y
Yingfan Xu
ShanghaiTech University
Haipeng Zhang
Haipeng Zhang
ShanghaiTech University
Data MiningGeo-spatial Data MiningWeb MiningSocial ComputingFintech