Lexicons and grammars for language processing: industrial or handcrafted products?

📅 2026-06-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the fundamental question of whether linguistic resources—such as dictionaries and grammars—should be constructed through meticulous manual curation or via scalable, automated methods. By systematically comparing representative resources like WordNet, FrameNet, and TAG, and integrating insights from both theoretical linguistics and natural language processing practice, the paper evaluates the trade-offs between these approaches in terms of semantic richness, construction efficiency, and downstream application performance. The findings indicate that manually crafted resources offer fine-grained semantic detail but incur high development costs, whereas automatically generated ones provide strong scalability at the expense of informational depth. A hybrid strategy combining both paradigms emerges as the most viable path forward. This work thus proposes a new paradigm for linguistic resource development that balances quality and efficiency, fostering synergistic advancement between linguistic theory and language technology.
📝 Abstract
During the recent years, the use of linguistic data for language processing increased progressively. Such data are now commonly called language resources. Most of the language resources used for this purpose are collections of texts as the Brown Corpus and the Penn Treebank, but electronic lexicons (WordNet, FrameNet, VerbNet, ComLex, Lexicon-Grammar...) and formal grammars (TAG...) developed recently. Most processes of construction of lexicons and grammars are manual, whereas the construction of corpora has always been highly automated. However, more and more specialists of language processing realize that the information content of lexicons and grammars is richer than that of corpora, and hence the former make more elaborate processing possible. The difference in construction time is likely to be connected with the difference in information content: the handcrafting of lexicons and grammars by linguists would make them more informative than automatically generated data. This situation can evolve into two directions: either specialists of language technology get progressively used to handling manually constructed resources, which are more informative and more complex, or the process of construction of lexicons and grammars is automated and industrialized, which is the mainstream perspective. Both evolutions are already in progress, and a tension exists between them. The relation between linguists and computer scientists depends on the future of these evolutions, since the first implies training and hiring numerous linguists, whereas the other depends essentially on solutions elaborated by computer engineers. The aim of this article is to analyse practical examples of the language resources in question, and to discuss about which of the two trends, handcrafting or generating industrially, or a combination of both, can give the best results or is the most realistic.
Problem

Research questions and friction points this paper is trying to address.

lexicons
grammars
language resources
handcrafting
industrialization
Innovation

Methods, ideas, or system contributions that make the work stand out.

language resources
handcrafted lexicons
industrial NLP
grammar engineering
corpus linguistics
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
É
Éric Laporte
Université Paris-Est - Laboratoire d'informatique de l'Institut Gaspard-Monge (IGM-Labinfo) - 5, bd Descartes - F77454 Marne-la-Vallée CEDEX 2 - France