Institution profile

Columbia Business School

Academic institutionnorthamerica · us
Official website
Research library8linked papers
Opportunities0open roles
Selected work

Representative Papers

WorkstreamBench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance

May 21, 2026

Existing benchmarks inadequately assess the ability of large language model (LLM) agents to generate complete spreadsheets end-to-end in financial contexts—such as financial modeling or scenario analysis—being largely confined to question answering or single-formula editing. This work proposes the first end-to-end spreadsheet generation evaluation framework tailored to real-world financial workflows, evaluating performance along three core dimensions: accuracy, formula correctness, and formatting compliance, while also introducing professional criteria such as readability and modifiability for the first time. The framework integrates both human and automated evaluation methods for comprehensive assessment. Experimental results demonstrate that Claude-family models achieve the strongest performance among current LLMs, yet still fall significantly short of expert human practitioners on complex tasks, revealing critical limitations of contemporary LLM agents in authentic financial settings.

0 citationsRead paper

Empirical Likelihood for Nonsmooth Functionals

Mar 29, 2026

This work addresses the limitations of conventional empirical likelihood methods, which rely on smoothness assumptions and fail to provide reliable inference for nonsmooth functionals—such as optimal value functionals in policy evaluation—particularly when the optimum is non-unique. To overcome this, the authors propose a geometric bootstrap empirical likelihood approach that reformulates the profile likelihood as the distance from the mean of estimating scores to a specific level set. By circumventing Taylor expansions and directly leveraging the convex optimization structure inherent in nonsmooth functionals, this method integrates empirical likelihood, geometric analysis, and an adaptive multiplier bootstrap. The resulting framework enables valid statistical inference for nonsmooth functionals, yielding accurately calibrated confidence intervals even in complex scenarios and substantially improving inferential reliability.

0 citationsRead paper

Calibrated Mechanism Design

Dec 19, 2025

This paper studies how to sustain incentive compatibility in dynamic environments where agents learn underlying states from allocation outcomes, under repeated use of a fixed mechanism. We propose “calibrated mechanism design”—a novel framework that decouples information disclosure from allocation by splitting the mechanism into two stages: first, a signal structure reveals state information; second, a state-independent static allocation rule is applied. We establish the theoretical foundation of this framework and prove that, in the single-agent case, its implementable set coincides exactly with the set of all incentive-compatible mechanisms. We show that full transparency is optimal under private values, while standard surplus extraction fails. The framework provides a rigorous microfoundation for infinite-horizon repeated interactions. By integrating information design, Bayesian mechanism design, and convex optimization, we derive necessary and sufficient conditions characterizing calibrated mechanisms. Finally, we demonstrate that history-dependent mechanisms expand feasibility only in non-quasilinear settings.

0 citationsRead paper

Leveraging LLMs to Improve Experimental Design: A Generative Stratification Approach

Sep 29, 2025

In high-dimensional covariate settings, conventional experimental stratification designs suffer from low efficiency, while traditional variable selection and weighting methods struggle to simultaneously ensure covariate balance and interpretability. Method: We propose a generative stratification framework that—novelly—integrates large language models (LLMs) into the pre-experimental design stage. Leveraging generative modeling, it automatically fuses heterogeneous covariate information to construct semantically informed strata, without requiring manual specification of variable importance or functional forms. Contribution/Results: Theoretically and empirically, our method reduces the variance of treatment effect estimation by 10–50% relative to simple randomization; further gains accrue when combined with classical stratification. Crucially, it extends the application paradigm of LLMs to causal inference design—offering a scalable, interpretable solution for high-dimensional experimental design.

0 citationsRead paper
Recent publications

Latest Papers

WorkstreamBench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance

May 21, 2026

Existing benchmarks inadequately assess the ability of large language model (LLM) agents to generate complete spreadsheets end-to-end in financial contexts—such as financial modeling or scenario analysis—being largely confined to question answering or single-formula editing. This work proposes the first end-to-end spreadsheet generation evaluation framework tailored to real-world financial workflows, evaluating performance along three core dimensions: accuracy, formula correctness, and formatting compliance, while also introducing professional criteria such as readability and modifiability for the first time. The framework integrates both human and automated evaluation methods for comprehensive assessment. Experimental results demonstrate that Claude-family models achieve the strongest performance among current LLMs, yet still fall significantly short of expert human practitioners on complex tasks, revealing critical limitations of contemporary LLM agents in authentic financial settings.

0 citationsRead paper

Empirical Likelihood for Nonsmooth Functionals

Mar 29, 2026

This work addresses the limitations of conventional empirical likelihood methods, which rely on smoothness assumptions and fail to provide reliable inference for nonsmooth functionals—such as optimal value functionals in policy evaluation—particularly when the optimum is non-unique. To overcome this, the authors propose a geometric bootstrap empirical likelihood approach that reformulates the profile likelihood as the distance from the mean of estimating scores to a specific level set. By circumventing Taylor expansions and directly leveraging the convex optimization structure inherent in nonsmooth functionals, this method integrates empirical likelihood, geometric analysis, and an adaptive multiplier bootstrap. The resulting framework enables valid statistical inference for nonsmooth functionals, yielding accurately calibrated confidence intervals even in complex scenarios and substantially improving inferential reliability.

0 citationsRead paper

Calibrated Mechanism Design

Dec 19, 2025

This paper studies how to sustain incentive compatibility in dynamic environments where agents learn underlying states from allocation outcomes, under repeated use of a fixed mechanism. We propose “calibrated mechanism design”—a novel framework that decouples information disclosure from allocation by splitting the mechanism into two stages: first, a signal structure reveals state information; second, a state-independent static allocation rule is applied. We establish the theoretical foundation of this framework and prove that, in the single-agent case, its implementable set coincides exactly with the set of all incentive-compatible mechanisms. We show that full transparency is optimal under private values, while standard surplus extraction fails. The framework provides a rigorous microfoundation for infinite-horizon repeated interactions. By integrating information design, Bayesian mechanism design, and convex optimization, we derive necessary and sufficient conditions characterizing calibrated mechanisms. Finally, we demonstrate that history-dependent mechanisms expand feasibility only in non-quasilinear settings.

0 citationsRead paper

Leveraging LLMs to Improve Experimental Design: A Generative Stratification Approach

Sep 29, 2025

In high-dimensional covariate settings, conventional experimental stratification designs suffer from low efficiency, while traditional variable selection and weighting methods struggle to simultaneously ensure covariate balance and interpretability. Method: We propose a generative stratification framework that—novelly—integrates large language models (LLMs) into the pre-experimental design stage. Leveraging generative modeling, it automatically fuses heterogeneous covariate information to construct semantically informed strata, without requiring manual specification of variable importance or functional forms. Contribution/Results: Theoretically and empirically, our method reduces the variance of treatment effect estimation by 10–50% relative to simple randomization; further gains accrue when combined with classical stratification. Crucially, it extends the application paradigm of LLMs to causal inference design—offering a scalable, interpretable solution for high-dimensional experimental design.

0 citationsRead paper