Columbo: Expanding Abbreviated Column Names for Tabular Data Using Large Language Models

📅 2025-08-12

📈 Citations: 0

✨ Influential: 0

career value

144K/year

🤖 AI Summary

This work addresses the semantic restoration of table column name abbreviations (e.g., “esal”)—a critical challenge for cross-domain data understanding and downstream task performance. Methodologically, we construct four high-quality benchmark datasets featuring real-world abbreviations; design synonym-aware, fine-grained evaluation metrics; and propose a novel LLM-based framework integrating context awareness, rule-based constraints, and chain-of-thought reasoning to enable token-level semantic modeling. Compared to the state-of-the-art NameGuess, our approach achieves 4–29% absolute accuracy gains across five benchmarks and is successfully deployed in the Environmental Data Initiative (EDI), a large-scale environmental science data platform. Key contributions include: (1) the first curated dataset of real-world column name abbreviations; (2) a new evaluation paradigm emphasizing semantic fidelity and synonym sensitivity; and (3) an interpretable, robust, end-to-end column name expansion framework.

Technology Category

Application Category

📝 Abstract

Expanding the abbreviated column names of tables, such as ``esal'' to ``employee salary'', is critical for numerous downstream data tasks. This problem arises in enterprises, domain sciences, government agencies, and more. In this paper we make three contributions that significantly advances the state of the art. First, we show that synthetic public data used by prior work has major limitations, and we introduce 4 new datasets in enterprise/science domains, with real-world abbreviations. Second, we show that accuracy measures used by prior work seriously undercount correct expansions, and we propose new synonym-aware measures that capture accuracy much more accurately. Finally, we develop Columbo, a powerful LLM-based solution that exploits context, rules, chain-of-thought reasoning, and token-level analysis. Extensive experiments show that Columbo significantly outperforms NameGuess, the current most advanced solution, by 4-29%, over 5 datasets. Columbo has been used in production on EDI, a major data portal for environmental sciences.

Problem

Research questions and friction points this paper is trying to address.

Expanding abbreviated column names in tables

Improving accuracy measures for name expansions

Developing an LLM-based solution for name expansion

Innovation

Methods, ideas, or system contributions that make the work stand out.

Introduces 4 real-world datasets for abbreviation expansion

Proposes synonym-aware accuracy measures for better evaluation

Develops Columbo using LLMs with context and token analysis

🔎 Similar Papers

On The Role of Prompt Construction In Enhancing Efficacy and Efficiency of LLM-Based Tabular Data Generation