Institution profile

Université Paris-Est

Academic institutioneurope · fr
Official website
Research library11linked papers
Opportunities0open roles
Selected work

Representative Papers

A French Corpus Annotated for Multiword Expressions with Adverbial Function

Jun 03, 2026

This study addresses a critical gap in natural language processing research: the scarcity of corpora specifically annotated for adverbial multiword expressions (MWEs). Focusing on French, the work provides the first systematic typology of MWEs that function adverbially and constructs a high-quality annotated corpus by integrating linguistic rules with manual validation, leveraging existing language resources. The resulting dataset, released under the LGPLLR license, constitutes the first publicly available resource dedicated to adverbial MWEs in French. By filling this longstanding void, the corpus offers foundational, reusable data to support downstream tasks such as information retrieval, information extraction, and syntactic parsing.

0 citationsRead paper

Lexicons and grammars for language processing: industrial or handcrafted products?

Jun 02, 2026

This study addresses the fundamental question of whether linguistic resources—such as dictionaries and grammars—should be constructed through meticulous manual curation or via scalable, automated methods. By systematically comparing representative resources like WordNet, FrameNet, and TAG, and integrating insights from both theoretical linguistics and natural language processing practice, the paper evaluates the trade-offs between these approaches in terms of semantic richness, construction efficiency, and downstream application performance. The findings indicate that manually crafted resources offer fine-grained semantic detail but incur high development costs, whereas automatically generated ones provide strong scalability at the expense of informational depth. A hybrid strategy combining both paradigms emerges as the most viable path forward. This work thus proposes a new paradigm for linguistic resource development that balances quality and efficiency, fostering synergistic advancement between linguistic theory and language technology.

0 citationsRead paper

French parsing enhanced with a word clustering method based on a syntactic lexicon

May 30, 2026

This study addresses the challenge of verb handling in French syntactic parsing under data sparsity by introducing, for the first time, a Lexicon-Grammar dictionary-driven verb clustering approach into a probabilistic context-free grammar (PCFG) parsing framework. By integrating lexically guided verb class information from the French Treebank, the proposed method effectively mitigates data sparsity and significantly improves parsing accuracy for French. Experimental results demonstrate that linguistically informed verb clustering enhances the performance of probabilistic parsers, offering a promising direction for syntactic modeling in low-resource languages.

0 citationsRead paper

Classification of non-analyzable word types in web documents to implement an effective Korean e-learning system

May 28, 2026

This study addresses the challenge that existing Korean e-learning systems struggle to effectively handle the abundant nonstandard and non-analyzable linguistic expressions prevalent in web-based texts. To tackle this issue, the research introduces, for the first time, the Local Grammar Graphs (LGG) model to model and classify nonstandard Korean expressions. By constructing and systematically comparing corpora of formal and informal Korean texts, the study reveals salient differences in their linguistic features. Experimental results demonstrate that LGG is highly effective in identifying and classifying non-analyzable lexical items, substantially enhancing the system’s coverage of authentic language phenomena. This advancement provides crucial technical support for developing Korean e-learning systems that better reflect real-world language use.

0 citationsRead paper

A new semantically annotated corpus with syntactic-semantic and cross-lingual senses

May 27, 2026

This study addresses the scarcity of fine-grained, cross-lingual annotated resources that integrate syntactic and semantic information for word sense disambiguation of French polysemous verbs. To bridge this gap, the authors construct a novel corpus covering 20 high-frequency polysemous verbs. By aligning actual translations from English parallel texts with lexical entries from the French Lexicon-Grammar dictionary, they propose and implement a tripartite joint annotation framework—comprising translation alignment labels, lexicon-grammar entry labels, and derived fine-grained semantic labels—for the first time. This resource substantially enhances data support and annotation granularity for word sense disambiguation, offering an innovative foundation for multilingual, multidimensional lexical semantic representation.

0 citationsRead paper
Recent publications

Latest Papers

A French Corpus Annotated for Multiword Expressions with Adverbial Function

Jun 03, 2026

This study addresses a critical gap in natural language processing research: the scarcity of corpora specifically annotated for adverbial multiword expressions (MWEs). Focusing on French, the work provides the first systematic typology of MWEs that function adverbially and constructs a high-quality annotated corpus by integrating linguistic rules with manual validation, leveraging existing language resources. The resulting dataset, released under the LGPLLR license, constitutes the first publicly available resource dedicated to adverbial MWEs in French. By filling this longstanding void, the corpus offers foundational, reusable data to support downstream tasks such as information retrieval, information extraction, and syntactic parsing.

0 citationsRead paper

Lexicons and grammars for language processing: industrial or handcrafted products?

Jun 02, 2026

This study addresses the fundamental question of whether linguistic resources—such as dictionaries and grammars—should be constructed through meticulous manual curation or via scalable, automated methods. By systematically comparing representative resources like WordNet, FrameNet, and TAG, and integrating insights from both theoretical linguistics and natural language processing practice, the paper evaluates the trade-offs between these approaches in terms of semantic richness, construction efficiency, and downstream application performance. The findings indicate that manually crafted resources offer fine-grained semantic detail but incur high development costs, whereas automatically generated ones provide strong scalability at the expense of informational depth. A hybrid strategy combining both paradigms emerges as the most viable path forward. This work thus proposes a new paradigm for linguistic resource development that balances quality and efficiency, fostering synergistic advancement between linguistic theory and language technology.

0 citationsRead paper

French parsing enhanced with a word clustering method based on a syntactic lexicon

May 30, 2026

This study addresses the challenge of verb handling in French syntactic parsing under data sparsity by introducing, for the first time, a Lexicon-Grammar dictionary-driven verb clustering approach into a probabilistic context-free grammar (PCFG) parsing framework. By integrating lexically guided verb class information from the French Treebank, the proposed method effectively mitigates data sparsity and significantly improves parsing accuracy for French. Experimental results demonstrate that linguistically informed verb clustering enhances the performance of probabilistic parsers, offering a promising direction for syntactic modeling in low-resource languages.

0 citationsRead paper

Classification of non-analyzable word types in web documents to implement an effective Korean e-learning system

May 28, 2026

This study addresses the challenge that existing Korean e-learning systems struggle to effectively handle the abundant nonstandard and non-analyzable linguistic expressions prevalent in web-based texts. To tackle this issue, the research introduces, for the first time, the Local Grammar Graphs (LGG) model to model and classify nonstandard Korean expressions. By constructing and systematically comparing corpora of formal and informal Korean texts, the study reveals salient differences in their linguistic features. Experimental results demonstrate that LGG is highly effective in identifying and classifying non-analyzable lexical items, substantially enhancing the system’s coverage of authentic language phenomena. This advancement provides crucial technical support for developing Korean e-learning systems that better reflect real-world language use.

0 citationsRead paper

A new semantically annotated corpus with syntactic-semantic and cross-lingual senses

May 27, 2026

This study addresses the scarcity of fine-grained, cross-lingual annotated resources that integrate syntactic and semantic information for word sense disambiguation of French polysemous verbs. To bridge this gap, the authors construct a novel corpus covering 20 high-frequency polysemous verbs. By aligning actual translations from English parallel texts with lexical entries from the French Lexicon-Grammar dictionary, they propose and implement a tripartite joint annotation framework—comprising translation alignment labels, lexicon-grammar entry labels, and derived fine-grained semantic labels—for the first time. This resource substantially enhances data support and annotation granularity for word sense disambiguation, offering an innovative foundation for multilingual, multidimensional lexical semantic representation.

0 citationsRead paper