Institution profile

Duke University

Academic institutionnorthamerica · us
Official website
Research library1,252linked papers
Opportunities0open roles
Selected work

Representative Papers

Budget Pacing in Repeated Auctions: Regret and Efficiency without Convergence

May 18, 2022Information Technology Convergence and Services

This paper investigates the impact of dynamic bidding pacing algorithms on group liquid welfare and individual dynamic regret in repeated ad auctions under budget constraints. To overcome the limitation of prior work—reliance on convergence assumptions about algorithmic dynamics—we propose a novel theoretical framework that makes no such assumptions. First, we establish that liquid welfare is guaranteed to be at least 50% of the optimal expected value, irrespective of convergence. Second, we derive an upper bound on dynamic regret tailored to time-varying budgets. Third, we design a gradient-based linear pacing algorithm within the core auction framework, integrating monotonic return-on-spend modeling and dynamic regret analysis to ensure broad applicability across first-price, second-price, and generalized second-price auctions. Empirical validation on Bing Ads data confirms the theoretical guarantees.

36 citations1 influentialRead paper

Assessing Omitted Variable Bias when the Controls are Endogenous

Jun 06, 2022

This paper addresses sensitivity analysis for omitted-variable bias in causal inference, focusing on the critical yet overlooked scenario where omitted variables are endogenous with respect to included controls—a setting neglected by existing methods. Conventional residualization-based approaches suffer from theoretical deficiencies under endogeneity, leading to erroneous robustness assessments; meanwhile, prevailing sensitivity analyses either rely on strong independence assumptions or lack comparable calibration. We formally prove the failure mechanism of residualization and propose a novel sensitivity analysis framework that explicitly accommodates correlation between omitted and observed covariates. Our approach introduces a standardized sensitivity parameter enabling comparable calibration of observable and unobservable selection strength. Theoretical derivation, implementation via a Stata module (regsensitivity), and empirical validation—using historical frontier settlement to instrument cultural beliefs—demonstrate that the framework rectifies fundamental theoretical shortcomings of mainstream methods and delivers a ready-to-use tool for robust causal inference.

17 citations2 influentialRead paper

Federated Large Language Models: Current Progress and Future Directions

Sep 24, 2024arXiv.org

To address the convergence difficulties and high communication overhead of large language models (LLMs) in federated learning (FL) caused by data heterogeneity, this paper introduces FedLLM—the first unified analytical framework for LLMs in FL. It systematically surveys two dominant paradigms: federated fine-tuning and federated prompt learning, while rigorously analyzing core challenges including data heterogeneity, communication efficiency, and privacy preservation. The work identifies promising future directions—namely, federated pre-training and LLM-augmented FL—and fills a critical gap in systematic literature review. A multidimensional taxonomy and evaluation framework is established to clarify key technical bottlenecks. Integrating insights from FL, LLM adaptation, prompt engineering, distributed optimization, and privacy-preserving computation, this study delivers a practical, robust, and privacy-aware methodology for deploying LLMs in real-world federated settings. (149 words)

16 citations1 influentialRead paper

An Optimal and Scalable Matrix Mechanism for Noisy Marginals under Convex Loss Functions

May 14, 2023Neural Information Processing Systems

Existing methods (e.g., HDMM) for releasing marginal queries over high-dimensional data (up to 100 dimensions) under differential privacy—especially for composite workloads combining marginals with range or prefix-sum queries—suffer from memory explosion and computational intractability. Method: We propose an efficient, unbiased Gaussian noise matrix mechanism that enables global optimization of arbitrary loss objectives expressible as convex functions of marginal variances. Our approach integrates residual planning (ResidualPlanner), convex optimization modeling, sparse linear algebra acceleration, and analytical computation of variance–covariance matrices. Results: Experiments demonstrate scalability: optimizing tens of thousands of marginals completes in seconds; hundred-attribute datasets are processed within two minutes. Memory consumption is reduced by one to two orders of magnitude. Crucially, our method is the first to support exact per-marginal variance and covariance output at scale—enabling principled downstream analysis and adaptive query answering in large-scale differentially private data release.

4 citationsRead paper

Sample Complexity of Distributionally Robust Off-Dynamics Reinforcement Learning with Online Interaction

Nov 07, 2025

This paper addresses online robust reinforcement learning under dynamics mismatch between training and deployment environments, focusing on exploration challenges induced by dynamic uncertainty. We introduce the *supremal visitation ratio* to quantify discrepancies in environment dynamics and, within a distributionally robust MDP framework, propose the first efficient online algorithm achieving sublinear regret under an *f*-divergence ambiguity set—attaining optimal dependence in its regret bound. Theoretically, we establish matching upper and lower bounds on regret. Empirically, the algorithm demonstrates significant improvements over baseline methods across diverse dynamic shift scenarios, exhibiting both strong robustness and high sample efficiency.

3 citations1 influentialRead paper
Recent publications

Latest Papers