Institution profile

Meta

Industry researchnorthamerica · us
Official website
Research library1,350linked papers
Opportunities0open roles
Selected work

Representative Papers

Treatment Allocation under Uncertain Costs

Mar 20, 2021

This paper addresses the problem of optimal treatment allocation under a budget constraint when treatment costs vary heterogeneously with covariates. We propose a threshold rule based on a priority score and establish, for the first time, a theoretical link between optimal allocation under uncertain costs and instrumental variable (IV) estimation of heterogeneous treatment effects. We rigorously derive the optimal threshold structure and prove its learnability. Our method integrates randomized controlled trial data, priority score modeling, threshold-based decision making, and an IV estimation framework. Empirically, the approach significantly outperforms standard benchmarks across multiple evaluation metrics, achieving maximal social value or firm profit within budget constraints. It provides a new paradigm—statistically rigorous yet practically implementable—for applications including scarce healthcare resource allocation and dynamic pricing.

14 citations1 influentialRead paper

AdvPrefix: An Objective for Nuanced LLM Jailbreaks

Dec 13, 2024arXiv.org

Existing LLM jailbreaking attacks suffer from weak controllability, incomplete response generation, and rigid, non-adaptive optimization formats. To address these limitations, we propose AdvPrefix—a novel prefix-based objective function designed for fine-grained jailbreaking. AdvPrefix introduces the first model-adaptive prefix selection mechanism, which automatically identifies high-quality prefixes during the prefilling stage using a dual criterion: attack success rate and negative log-likelihood. It further enables multi-prefix collaborative optimization, departing from conventional fixed-prefix paradigms and exposing alignment models’ generalization vulnerabilities to unseen prefixes. AdvPrefix is fully compatible with mainstream optimization frameworks (e.g., GCG) and requires no model modification or retraining. Evaluated on Llama-3, AdvPrefix boosts GCG’s fine-grained jailbreaking success rate from 14% to 80%, demonstrating the critical impact of objective function design on jailbreaking efficacy.

10 citations3 influentialRead paper

Agentic Reasoning for Large Language Models

Jan 18, 2026

This work addresses the limited capacity of large language models (LLMs) to plan, act, and learn through sustained interaction in open, dynamic environments. To overcome this, the authors propose a three-tiered reasoning framework that treats LLMs as autonomous agents, unifying single-agent foundational reasoning, self-evolution, and multi-agent collaboration within a coherent paradigm. The framework orchestrates structured interactions, incorporates memory mechanisms, enables tool use, and integrates both reinforcement learning and supervised fine-tuning. It explicitly distinguishes between in-context reasoning and post-training optimization pathways, systematically coupling cognition with action. Empirical validation across diverse domains—including scientific discovery, robotics, healthcare, autonomous research, and mathematics—demonstrates its effectiveness and highlights promising future directions such as personalized interaction, long-horizon engagement, world modeling, and scalable multi-agent training.

7 citations1 influentialRead paper

Learning Latent Action World Models In The Wild

Jan 08, 2026arXiv.org

This work addresses the challenge of lacking explicit action labels in in-the-wild videos by proposing an unsupervised learning framework to construct generalizable world models for agent reasoning and planning. The approach introduces spatially localized, continuously constrained latent action representations, enabling self-supervised learning of action-state dynamics from diverse real-world videos without requiring a unified embodied structure. By integrating continuous latent action modeling, a controller mapping mechanism, and tailored architectural design, this framework is the first to successfully extend latent-action world models to complex in-the-wild scenarios. Experiments demonstrate that the model achieves performance on par with baselines using ground-truth action labels in cross-video action transfer and planning tasks, validating the effectiveness and scalability of latent actions as a universal interface.

4 citations1 influentialRead paper

Machine Learning-Guided Memory Optimization for DLRM Inference on Tiered Memory

Mar 01, 2025International Symposium on High-Performance Computer Architecture

Deep learning recommendation models (DLRMs) require terabyte-scale embedding memory, and while hierarchical memory offers cost efficiency, its irregular access patterns severely degrade embedding placement and cache efficiency. To address this, we propose RecMG—a novel system that decouples cache admission decisions from prefetching prediction into two independently trainable models for the first time. We introduce a differentiable loss function explicitly modeling long reuse distances and infrequent embedding accesses, significantly improving prefetching accuracy. RecMG integrates vectorized access pattern learning, hierarchical-memory-aware scheduling, and dynamic prefetching policies. Experimental results show that RecMG reduces on-demand embedding loads by 1.5–2.8× over state-of-the-art baselines and achieves up to a 43% reduction in end-to-end inference latency in industrial-scale deployments.

4 citationsRead paper
Recent publications

Latest Papers