PinSieve: Production Selective VLM Serving and a Governed Memory Flywheel for Enterprise Content-Quality Triage

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决企业内容质量筛选问题,提出PinSieve系统,通过选择性视觉-语言模型服务代理和管理记忆飞轮方法提高审查效率并降低成本。
📝 Abstract
Enterprise AI agents in production often need to be bounded, stateful, observable, and governable rather than fully autonomous. We present PinSieve, a production case study in a large-scale content-quality pipeline. Its deployed component is a selective vision-language-model (VLM) Serving Agent that operates only on the grey-zone slice left unresolved by lightweight upstream models, exposes a scalar routing score online, and preserves controlled human escalation. On this slice, the deployed system filters 2.05x more non-actionable items than the previous production module while slightly reducing estimated miss rate; after promotion, it improves review productivity by 25.7%, reduces normalized operating cost by 16.2%, and moves signal delivery from next-day to same-day. We then study maintenance through a governed memory flywheel under selective feedback, where escalated items are reviewed by default and auto-passed items are labeled mainly through audit sampling. Feedback Memory records routing traces, observation paths, audit propensities, and replay metadata for evaluation and debugging. The Data Curation Agent uses a bounded proposal-verifier loop over representative, uncertainty, recency, and fresh-review replay, with positive-rate and score-bin guardrails before batch acceptance. In chained monthly refresh over six months of production data, this design reduces average FNR@50% from 17.73% under representative random replay to 13.29%. A Reasoning Review Agent audits teacher-generated rationales and supports keep/repair/drop decisions. Production claims are attributed only to the deployed Serving Agent; replay and rationale-review results are offline or sampled-governance evidence. The same serving-agent recipe has been adopted to several additional internal signals, suggesting transferability beyond one task.
Problem

Research questions and friction points this paper is trying to address.

Enterprise AI agents
content-quality pipeline
selective VLM serving
governed memory flywheel
Innovation

Methods, ideas, or system contributions that make the work stand out.

Selective VLM Serving
Governed Memory Flywheel
Feedback Memory
Data Curation Agent
💼 Related Jobs
No related jobs found.
C
Chuqing Gao
Pinterest, New York, NY, USA
Y
Yuanfang Song
Pinterest, Palo Alto, CA, USA
J
Jonathan Zhang
Pinterest, New York, NY, USA
Y
Yifan Wu
Pinterest, Palo Alto, CA, USA
V
Vishwakarma Singh
Pinterest, Palo Alto, CA, USA
Q
Qinglong Zeng
Pinterest, San Francisco, CA, USA
A
Andrey Gusev
Pinterest, San Francisco, CA, USA