The WebKurator.de Platform: Combined Regional and Topical Web Curation

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决网络内容系统性保存问题,本文提出WebKurator.de平台,采用结合地域与主题的双维度策展模型,并利用LLM技术进行分类和地址提取。
📝 Abstract
The systematic curation of the Web remains a central challenge for national libraries and memory institutions that aim to preserve culturally and regionally relevant content. Existing directory-based approaches such as Curlie implement a predominantly topic-centric, one-dimensional hierarchy, where geographic aspects are intertwined with topical and linguistic categories. To address this limitation, we present WebKurator.de, a collaborative platform for combined regional and topical web curation, initially focused on the German web. WebKurator introduces a two-dimensional curation model that explicitly separates topical categorization and geographic annotation. The system integrates LLM-based topic classification and imprint-based address extraction with geocoding, and supports user suggestions together with moderated review. The platform is bootstrapped from the German Imprints Dataset, a large-scale collection of 5.54 million websites. Among them, 3.14 million contain imprint pages, for which we successfully extracted and geocoded postal addresses. Of these, 2.58 million (85.17%) are located in Germany and also have an assigned topic label. These websites form the initial foundation of WebKurator.de and can be continuously extended through user suggestions.
Problem

Research questions and friction points this paper is trying to address.

Web Curation
National Libraries
Memory Institutions
Topical Hierarchy
Geographic Aspects
Innovation

Methods, ideas, or system contributions that make the work stand out.

combined regional and topical curation
two-dimensional curation model
LLM-based topic classification
imprint-based address extraction
geocoding
🔎 Similar Papers