Institution profile

Uber

Industry researchnorthamerica · us
Official website
Research library13linked papers
Opportunities37open roles
Selected work

Representative Papers

Governance at the Boundary: How Agent Decomposition Degrades Policy Compliance

Aug 16, 2026

This study addresses the compliance degradation arising from agent component decomposition by introducing Fiducia-bench, the first governance benchmark for financial agents. Through KYC/AML tasks and comparative experiments across multiple architectures, this work reveals a novel mechanism wherein boundary information attenuation within orchestration frameworks leads to compliance failures. Notably, factual decay rates reach 85% and are significantly modulated by model capabilities. Furthermore, the research elucidates the correlation between architectural complexity and governance costs, providing critical mechanistic insights into agent compliance risks. By open-sourcing both the dataset and validation framework, this work fills a significant gap in evaluating financial agent governance, establishing a foundational resource for future research on compliant autonomous systems in regulated domains.

0 citationsRead paper

Evaluating for the long term: Learnings from industry

Aug 08, 2026

Short-term online experiments often fail to accurately predict long-term business outcomes, potentially leading to decisions misaligned with strategic objectives. Drawing on insights from industry expert workshops, this work proposes a design principle for surrogate metrics centered on decision utility rather than solely on unbiasedness, advocating that interpretable, experiment-driven simple surrogates outperform complex black-box models. Through expert consensus synthesis, surrogate metric analysis, and comparative evaluation of experimental versus observational data, the study systematically outlines a methodology for constructing effective surrogates and uncovers a critical relationship between the stability of long-term effects and surrogate validity. While underscoring the irreplaceable value of high-quality long-term experimentation, the research also delineates core challenges and practical guidelines for surrogate learning in real-world settings.

0 citationsRead paper

Beyond Static Evaluation: Building Simulation Environments for Scalable Agentic Reinforcement Learning

Jul 06, 2026

Existing static evaluation methods struggle to effectively assess the agentic capabilities of large language models in multi-step decision-making tasks. To address this limitation, this work proposes AgenticAI-Supervisor, the first reinforcement learning simulation framework for agentic AI that supports closed-loop feedback. By decoupling environment construction from scalable execution, the framework integrates API- and UI-driven Gym environments, generates high-fidelity execution trajectories, employs multi-dimensional reward shaping, and incorporates internal state validation mechanisms to mitigate reward hacking. Demonstrated in a customer service agent case study, the framework enables stable closed-loop feedback and significantly enhances model optimization outcomes.

0 citationsRead paper

How Much Coordination Gain Is Real? A Paired Noise-Floor Protocol for Multi-Agent LLM Benchmarks

Jun 15, 2026

This study addresses the common tendency in existing research to attribute minor performance differences among multi-agent large language model (LLM) coordination architectures to their inherent merits, while overlooking the impact of random noise. To rigorously isolate architectural effects, we introduce a paired experimental protocol that strictly controls model and task configurations, enabling the first quantification of a noise baseline in coordination mechanisms. We further propose a pass^k reporting standard that scores only when coordination logic is actively engaged. Leveraging SHA-256 auditing, Wilson confidence intervals, and coordination activation detection hooks, our experiments on Claude Haiku 4.5 reveal performance fluctuations due to noise ranging from –3 to +18 percentage points (with a combined upper bound of approximately 15 pp). Notably, most recently reported gains fall below this noise floor, and their coordination mechanisms show no significant improvement in task recovery rates on first attempts.

0 citationsRead paper
Recent publications

Latest Papers

Governance at the Boundary: How Agent Decomposition Degrades Policy Compliance

Aug 16, 2026

This study addresses the compliance degradation arising from agent component decomposition by introducing Fiducia-bench, the first governance benchmark for financial agents. Through KYC/AML tasks and comparative experiments across multiple architectures, this work reveals a novel mechanism wherein boundary information attenuation within orchestration frameworks leads to compliance failures. Notably, factual decay rates reach 85% and are significantly modulated by model capabilities. Furthermore, the research elucidates the correlation between architectural complexity and governance costs, providing critical mechanistic insights into agent compliance risks. By open-sourcing both the dataset and validation framework, this work fills a significant gap in evaluating financial agent governance, establishing a foundational resource for future research on compliant autonomous systems in regulated domains.

0 citationsRead paper

Evaluating for the long term: Learnings from industry

Aug 08, 2026

Short-term online experiments often fail to accurately predict long-term business outcomes, potentially leading to decisions misaligned with strategic objectives. Drawing on insights from industry expert workshops, this work proposes a design principle for surrogate metrics centered on decision utility rather than solely on unbiasedness, advocating that interpretable, experiment-driven simple surrogates outperform complex black-box models. Through expert consensus synthesis, surrogate metric analysis, and comparative evaluation of experimental versus observational data, the study systematically outlines a methodology for constructing effective surrogates and uncovers a critical relationship between the stability of long-term effects and surrogate validity. While underscoring the irreplaceable value of high-quality long-term experimentation, the research also delineates core challenges and practical guidelines for surrogate learning in real-world settings.

0 citationsRead paper

Beyond Static Evaluation: Building Simulation Environments for Scalable Agentic Reinforcement Learning

Jul 06, 2026

Existing static evaluation methods struggle to effectively assess the agentic capabilities of large language models in multi-step decision-making tasks. To address this limitation, this work proposes AgenticAI-Supervisor, the first reinforcement learning simulation framework for agentic AI that supports closed-loop feedback. By decoupling environment construction from scalable execution, the framework integrates API- and UI-driven Gym environments, generates high-fidelity execution trajectories, employs multi-dimensional reward shaping, and incorporates internal state validation mechanisms to mitigate reward hacking. Demonstrated in a customer service agent case study, the framework enables stable closed-loop feedback and significantly enhances model optimization outcomes.

0 citationsRead paper

How Much Coordination Gain Is Real? A Paired Noise-Floor Protocol for Multi-Agent LLM Benchmarks

Jun 15, 2026

This study addresses the common tendency in existing research to attribute minor performance differences among multi-agent large language model (LLM) coordination architectures to their inherent merits, while overlooking the impact of random noise. To rigorously isolate architectural effects, we introduce a paired experimental protocol that strictly controls model and task configurations, enabling the first quantification of a noise baseline in coordination mechanisms. We further propose a pass^k reporting standard that scores only when coordination logic is actively engaged. Leveraging SHA-256 auditing, Wilson confidence intervals, and coordination activation detection hooks, our experiments on Claude Haiku 4.5 reveal performance fluctuations due to noise ranging from –3 to +18 percentage points (with a combined upper bound of approximately 15 pp). Notably, most recently reported gains fall below this noise floor, and their coordination mechanisms show no significant improvement in task recovery rates on first attempts.

0 citationsRead paper