Peacemaker at ATE-IT: Automatic term extraction from Italian text for waste management data using encoder model

📅 2026-05-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of scarce annotated data and difficulties in multi-word term extraction under domain shift in Italian waste management texts. To tackle these issues, the authors propose a low-cost, interpretable approach to automatic term extraction based on a fine-tuned encoder-based pretrained language model. The method achieves efficient term identification under limited computational resources and is evaluated using both type-level and token-level metrics. When applied to Task A of the ATE shared task (ATE-IT), the approach yields stable and balanced performance in precision, recall, and F1 score, comparable to that of other participating teams. These results demonstrate its effectiveness and offer a reliable, scalable solution for term extraction in low-resource scenarios.
📝 Abstract
The development of automatic term extraction has become increasingly important in modern technology. Automatic term extraction can be found in virtually every search engine that is currently available to users. Recent advancements have provided promising results for the extraction of automatic terms; however, accurate labeling is difficult because of several factors, such as the limited number of annotated documents available for training and the complexity of extracting multi-word expressions due to shifts in the domain. In this paper, we will present a low-cost and interpretable method of automatic term extraction, developed specifically for Task A of the ATE Shared Task. This new method utilizes fine-tuning extraction strategies that can run on a small amount of computational resources. We evaluated our automated system using both type-level and micro-level measures of precision, recall, and F1-score to measure both complementary aspects of the extraction performance. According to the experimental results, our proposed approach achieves consistent and balanced performance compared to other teams. Even though the technique itself is relatively straightforward, it serves as a good starting point for low-resource models. Overall, the findings point toward the possibility of significant future advancements (in model expansion) with higher-level performance still able to retain their ability to be interpreted.
Problem

Research questions and friction points this paper is trying to address.

automatic term extraction
low-resource
multi-word expressions
domain-specific terminology
Italian text
Innovation

Methods, ideas, or system contributions that make the work stand out.

automatic term extraction
low-resource
encoder model
interpretability
multi-word expressions
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Mahdi Bakhtiyarzadeh
Department of Computer Science, University of Tabriz, 29 Bahman Boulevard, Tabriz 51666-16471, Iran
H
Hadi Bayrami Asl Tekanlou
Department of Computer Science, University of Tabriz, 29 Bahman Boulevard, Tabriz 51666-16471, Iran
Jafar Razmara
Jafar Razmara
Associate Professor, Department of Computer Science, University of Tabriz
Machine learningDeep learningBioinformatics