Columbo: Expanding Abbreviated Column Names for Tabular Data Using Large Language Models

Ting Cai; Stephen Sheen; AnHai Doan

doi:10.18653/v1/2025.findings-emnlp.1348

Columbo: Expanding Abbreviated Column Names for Tabular Data Using Large Language Models

Abstract

Expanding the abbreviated column names of tables, such as “esal” to “employee salary”, is critical for many downstream NLP tasks for tabular data, such as NL2SQL, table QA, and keyword search. This problem arises in enterprises, domain sciences, government agencies, and more. In this paper, we make three contributions that significantly advance the state of the art. First, we show that the synthetic public data used by prior work has major limitations, and we introduce four new datasets in enterprise/science domains, with real-world abbreviations. Second, we show that accuracy measures used by prior work seriously undercount correct expansions, and we propose new synonym-aware measures that capture accuracy much more accurately. Finally, we develop Columbo, a powerful LLM-based solution that exploits context, rules, chain-of-thought reasoning, and token-level analysis. Extensive experiments show that Columbo significantly outperforms NameGuess, the current most advanced solution, by 4-29%, over five datasets. Columbo has been used in production on EDI, a major data lake for environmental sciences.

Anthology ID:: 2025.findings-emnlp.1348
Volume:: Findings of the Association for Computational Linguistics: EMNLP 2025
Month:: November
Year:: 2025
Address:: Suzhou, China
Editors:: Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, Violet Peng
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 24774–24792
Language:
URL:: https://preview.aclanthology.org/name-variant-enfa-fane/2025.findings-emnlp.1348/
DOI:: 10.18653/v1/2025.findings-emnlp.1348
Bibkey:
Cite (ACL):: Ting Cai, Stephen Sheen, and AnHai Doan. 2025. Columbo: Expanding Abbreviated Column Names for Tabular Data Using Large Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 24774–24792, Suzhou, China. Association for Computational Linguistics.
Cite (Informal):: Columbo: Expanding Abbreviated Column Names for Tabular Data Using Large Language Models (Cai et al., Findings 2025)
Copy Citation:
PDF:: https://preview.aclanthology.org/name-variant-enfa-fane/2025.findings-emnlp.1348.pdf
Checklist:: 2025.findings-emnlp.1348.checklist.pdf

PDF Cite Search Checklist Fix data