Fine PT-PT Web: A High-Quality 41 Billion Tokens Data Collection of the European Portuguese Web

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出一种高效管道,通过去除冗余、语言识别和质量分类等方法,从葡萄牙语网页中筛选出高质量的欧洲葡萄牙语文本数据集。
📝 Abstract
Curating Web corpora for regional language variants like European Portuguese (PT-PT) is heavily bottlenecked by dialectal overlap (mainly with PT-BR) and data processing scale. This paper presents an efficient pipeline to curate a production-ready PT-PT corpus from the Portuguese Web, spanning 411 TB of raw data from Arquivo.pt. We introduce a novel post-scraping block that removes boilerplate and line duplicates prior to filtering. This early-stage intervention increases final document yield by 19.04% by rescuing valid text that standard heuristic filters prematurely discard. Integrated with rigorous language identification, weighted fuzzy deduplication, and neural quality classification, our pipeline offers a scalable framework and a clean, representative corpus optimized for LLM pre-training.
Problem

Research questions and friction points this paper is trying to address.

European Portuguese
dialectal overlap
data processing scale
corpus curation
Innovation

Methods, ideas, or system contributions that make the work stand out.

post-scraping block
language identification
fuzzy deduplication
neural quality classification
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
G
Gonçalo Vinagre
NOVA School of Science and Technology, NOVA LINCS
R
Rui Pedro Guerra
NOVA School of Science and Technology, Fundação para a Ciência e Tecnologia
P
Pedro Gomes
Fundação para a Ciência e Tecnologia
M
Miguel Moura Ramos
Instituto Superior Técnico, Universidade de Lisboa, Instituto de Telecomunicações
D
Duarte Miguel Alves
Instituto Superior Técnico, Universidade de Lisboa, Instituto de Telecomunicações
A
Afonso Simplício
NOVA School of Science and Technology, NOVA LINCS
Diogo Tavares
Diogo Tavares
NOVA School of Science and Technology
David Semedo
David Semedo
Universidade NOVA de Lisboa
Vision and LanguageDeep Learning for MultimediaConversational AI
D
Daniel Gomes
Fundação para a Ciência e Tecnologia
J
João Magalhães
NOVA School of Science and Technology, NOVA LINCS