Institution profile

Snap Inc.

Industry researchnorthamerica · us
Official website
Research library156linked papers
Opportunities21open roles
Selected work

Representative Papers

Why Keep Your Doubts to Yourself? Trading Visual Uncertainties in Multi-Agent Bandit Systems

Jan 26, 2026

This work addresses the high coordination costs and low efficiency commonly encountered by multi-agent vision systems under information asymmetry, problems exacerbated by existing approaches that overlook the structural nature of uncertainty and lack economic sustainability. To this end, we propose Agora, a novel framework that formalizes epistemic uncertainty as a tradable asset and establishes a decentralized uncertainty market, wherein agents are incentivized through economic mechanisms to exchange uncertainty at the levels of perception, semantics, and reasoning. Agora integrates vision-language models with multi-armed bandits and Thompson Sampling to devise market-aware brokerage strategies. Experiments demonstrate that Agora significantly outperforms current methods across five multimodal benchmarks, achieving an 8.5% accuracy gain on MMMU while reducing coordination costs by more than threefold.

2 citationsRead paper

MemRec: Collaborative Memory-Augmented Agentic Recommender System

Jan 13, 2026

This work addresses the limitations of existing agent-based recommender systems, which suffer from isolated memory and struggle to effectively leverage user collaborative signals and graph-structured context. To overcome this, we propose MemRec, a novel framework that decouples reasoning from memory management for the first time. MemRec employs a lightweight language model (LM_Mem) to dynamically maintain a collaborative memory graph and supplies high-signal contextual information to a large recommendation model (LLM_Rec). By integrating efficient retrieval with an asynchronous graph propagation algorithm, our approach enables local, open-source deployment while preserving user privacy and minimizing computational cost. Extensive experiments on four benchmark datasets demonstrate state-of-the-art performance, establishing a new Pareto frontier that balances recommendation quality, efficiency, and privacy preservation.

2 citationsRead paper

Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling

Jan 17, 2026

This work addresses the vulnerability of Softmax attention to irrelevant tokens—so-called “attention sinks”—and the dispersion of attention probabilities over long sequences, which degrade model performance. To mitigate these issues, the authors propose Thresholded Differentiable Attention (TDA), a novel mechanism that integrates row-wise extreme-value thresholding, a length-dependent gating function, and inhibitory attention views. TDA achieves ultra-sparse, sink-free, and non-diffuse attention distributions without incurring additional projection costs. It generates over 99% exact-zero attention weights, confines spurious activations to a constant level, and ensures that cross-view false matches asymptotically vanish as context length increases. Empirical results demonstrate that TDA maintains competitive performance on both standard and long-context benchmarks.

1 citationsRead paper

SnapGen++: Unleashing Diffusion Transformers for Efficient High-Fidelity Image Generation on Edge Devices

Jan 13, 2026

Although Diffusion Transformers (DiTs) achieve impressive performance in image generation, their high computational and memory demands hinder deployment on edge devices. To address this challenge, this work proposes an efficient DiT framework featuring three key innovations: an adaptive global-local sparse attention mechanism to reduce computational complexity, an elastic training strategy within a unified hypernetwork enabling dynamic model scaling, and a four-step generative approach—KG-DMD—that integrates distribution matching with knowledge distillation. The resulting framework enables high-fidelity image generation in just four steps across diverse edge hardware platforms, significantly improving inference efficiency while maintaining visual quality and achieving a favorable balance between real-time performance and generation fidelity.

1 citationsRead paper
Recent publications

Latest Papers