Synthetic Data in Marketing Research: How to Evaluate and When to Trust
本文探讨了合成数据在市场研究中的应用,通过分类合成数据类型和准确性度量方法,并提出了一种无需真实数据的诊断方法来提高合成数据的可靠性。
本文探讨了合成数据在市场研究中的应用,通过分类合成数据类型和准确性度量方法,并提出了一种无需真实数据的诊断方法来提高合成数据的可靠性。
Existing benchmarks inadequately assess the ability of large language model (LLM) agents to generate complete spreadsheets end-to-end in financial contexts—such as financial modeling or scenario analysis—being largely confined to question answering or single-formula editing. This work proposes the first end-to-end spreadsheet generation evaluation framework tailored to real-world financial workflows, evaluating performance along three core dimensions: accuracy, formula correctness, and formatting compliance, while also introducing professional criteria such as readability and modifiability for the first time. The framework integrates both human and automated evaluation methods for comprehensive assessment. Experimental results demonstrate that Claude-family models achieve the strongest performance among current LLMs, yet still fall significantly short of expert human practitioners on complex tasks, revealing critical limitations of contemporary LLM agents in authentic financial settings.
This work addresses the limitations of conventional empirical likelihood methods, which rely on smoothness assumptions and fail to provide reliable inference for nonsmooth functionals—such as optimal value functionals in policy evaluation—particularly when the optimum is non-unique. To overcome this, the authors propose a geometric bootstrap empirical likelihood approach that reformulates the profile likelihood as the distance from the mean of estimating scores to a specific level set. By circumventing Taylor expansions and directly leveraging the convex optimization structure inherent in nonsmooth functionals, this method integrates empirical likelihood, geometric analysis, and an adaptive multiplier bootstrap. The resulting framework enables valid statistical inference for nonsmooth functionals, yielding accurately calibrated confidence intervals even in complex scenarios and substantially improving inferential reliability.
This paper studies how to sustain incentive compatibility in dynamic environments where agents learn underlying states from allocation outcomes, under repeated use of a fixed mechanism. We propose “calibrated mechanism design”—a novel framework that decouples information disclosure from allocation by splitting the mechanism into two stages: first, a signal structure reveals state information; second, a state-independent static allocation rule is applied. We establish the theoretical foundation of this framework and prove that, in the single-agent case, its implementable set coincides exactly with the set of all incentive-compatible mechanisms. We show that full transparency is optimal under private values, while standard surplus extraction fails. The framework provides a rigorous microfoundation for infinite-horizon repeated interactions. By integrating information design, Bayesian mechanism design, and convex optimization, we derive necessary and sufficient conditions characterizing calibrated mechanisms. Finally, we demonstrate that history-dependent mechanisms expand feasibility only in non-quasilinear settings.
In high-dimensional covariate settings, conventional experimental stratification designs suffer from low efficiency, while traditional variable selection and weighting methods struggle to simultaneously ensure covariate balance and interpretability. Method: We propose a generative stratification framework that—novelly—integrates large language models (LLMs) into the pre-experimental design stage. Leveraging generative modeling, it automatically fuses heterogeneous covariate information to construct semantically informed strata, without requiring manual specification of variable importance or functional forms. Contribution/Results: Theoretically and empirically, our method reduces the variance of treatment effect estimation by 10–50% relative to simple randomization; further gains accrue when combined with classical stratification. Crucially, it extends the application paradigm of LLMs to causal inference design—offering a scalable, interpretable solution for high-dimensional experimental design.
本文探讨了合成数据在市场研究中的应用,通过分类合成数据类型和准确性度量方法,并提出了一种无需真实数据的诊断方法来提高合成数据的可靠性。
Existing benchmarks inadequately assess the ability of large language model (LLM) agents to generate complete spreadsheets end-to-end in financial contexts—such as financial modeling or scenario analysis—being largely confined to question answering or single-formula editing. This work proposes the first end-to-end spreadsheet generation evaluation framework tailored to real-world financial workflows, evaluating performance along three core dimensions: accuracy, formula correctness, and formatting compliance, while also introducing professional criteria such as readability and modifiability for the first time. The framework integrates both human and automated evaluation methods for comprehensive assessment. Experimental results demonstrate that Claude-family models achieve the strongest performance among current LLMs, yet still fall significantly short of expert human practitioners on complex tasks, revealing critical limitations of contemporary LLM agents in authentic financial settings.
This work addresses the limitations of conventional empirical likelihood methods, which rely on smoothness assumptions and fail to provide reliable inference for nonsmooth functionals—such as optimal value functionals in policy evaluation—particularly when the optimum is non-unique. To overcome this, the authors propose a geometric bootstrap empirical likelihood approach that reformulates the profile likelihood as the distance from the mean of estimating scores to a specific level set. By circumventing Taylor expansions and directly leveraging the convex optimization structure inherent in nonsmooth functionals, this method integrates empirical likelihood, geometric analysis, and an adaptive multiplier bootstrap. The resulting framework enables valid statistical inference for nonsmooth functionals, yielding accurately calibrated confidence intervals even in complex scenarios and substantially improving inferential reliability.
This paper studies how to sustain incentive compatibility in dynamic environments where agents learn underlying states from allocation outcomes, under repeated use of a fixed mechanism. We propose “calibrated mechanism design”—a novel framework that decouples information disclosure from allocation by splitting the mechanism into two stages: first, a signal structure reveals state information; second, a state-independent static allocation rule is applied. We establish the theoretical foundation of this framework and prove that, in the single-agent case, its implementable set coincides exactly with the set of all incentive-compatible mechanisms. We show that full transparency is optimal under private values, while standard surplus extraction fails. The framework provides a rigorous microfoundation for infinite-horizon repeated interactions. By integrating information design, Bayesian mechanism design, and convex optimization, we derive necessary and sufficient conditions characterizing calibrated mechanisms. Finally, we demonstrate that history-dependent mechanisms expand feasibility only in non-quasilinear settings.
In high-dimensional covariate settings, conventional experimental stratification designs suffer from low efficiency, while traditional variable selection and weighting methods struggle to simultaneously ensure covariate balance and interpretability. Method: We propose a generative stratification framework that—novelly—integrates large language models (LLMs) into the pre-experimental design stage. Leveraging generative modeling, it automatically fuses heterogeneous covariate information to construct semantically informed strata, without requiring manual specification of variable importance or functional forms. Contribution/Results: Theoretically and empirically, our method reduces the variance of treatment effect estimation by 10–50% relative to simple randomization; further gains accrue when combined with classical stratification. Crucially, it extends the application paradigm of LLMs to causal inference design—offering a scalable, interpretable solution for high-dimensional experimental design.