When Can We Work in Embedding Space? What Text Embeddings Preserve

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了在何种情况下可以使用文本嵌入进行实证分析,通过假设文档是潜在主题的混合,并利用该方法在美国363个大都市区的应用中验证其有效性。
📝 Abstract
When do text embeddings work as inputs to empirical analysis? Their use rests on an assumption: that we can trade text for its low-dimensional embedding, and lose little in doing so. I make that assumption precise under a generative model in which documents are mixtures of latent topics. I study two uses---clustering units in embedding space and controlling for high-dimensional text. A cluster of embeddings is a set of documents with similar topic mixtures; controlling for the embedding is equivalent to controlling for the topic mixture, so validity reduces to whether that mixture captures the confounding. In an application to 363 U.S. metropolitan areas, embedding-based clusters of LLM-generated economic descriptions recover interpretable economic archetypes and separate local employment dynamics more sharply than clustering on model residuals, or on a curated set of industry and demographic covariates.
Problem

Research questions and friction points this paper is trying to address.

text embeddings
empirical analysis
low-dimensional embedding
latent topics
clustering
Innovation

Methods, ideas, or system contributions that make the work stand out.

text embeddings
generative model
latent topics
clustering in embedding space
controlling for high-dimensional text