Extracting Dataset Mentions in Forced Displacement and FCV Documents: A Weakly Supervised Framework with LLM-Based Label Refinement

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种弱监督框架,利用轻量级模型和大型语言模型结合的方法,从非结构化文本中提取数据集引用,解决了系统识别数据集引用困难的问题。
📝 Abstract
Development and humanitarian organizations produce and support surveys, administrative registries, and other data resources to inform research, policy, and operations, yet systematically identifying where these datasets are referenced remains difficult. Such references are dispersed across research papers, project documents, humanitarian reports, and other unstructured text, limiting both the ability to trace data use and to identify potential gaps in data availability or dissemination. We present a weakly supervised framework for adapting dataset extraction to forced displacement and Fragile, Conflict, and Violence (FCV) documents without first constructing a large manually labeled training corpus. A lightweight model trained on general research literature generates candidate dataset mentions from unlabeled domain documents, which a frontier large language model (LLM) reviews in context, validating or rejecting candidates and correcting their extraction boundaries. The resulting annotations are supplemented with targeted synthetic and contrastive examples and used to fine-tune the lightweight model for large-scale extraction. We evaluate the resulting model on an independent gold-standard benchmark of 1,706 text passages spanning research, humanitarian, and operational documents. Across the full benchmark, the model achieves 74.1\% precision and 70.5\% recall at the mention level; among passages containing dataset references, precision reaches 89.5\%. At the passage level, the model achieves 88.2\% accuracy and 88.6\% specificity in distinguishing passages with dataset references from those without them. These results demonstrate a practical approach for constructing domain-specific supervision when labeled data are limited, and provide a technical foundation for larger-scale analysis of data use and potential gaps in the displacement data landscape.
Problem

Research questions and friction points this paper is trying to address.

dataset mentions
forced displacement
Fragile, Conflict, and Violence (FCV)
unstructured text
data use
Innovation

Methods, ideas, or system contributions that make the work stand out.

weakly supervised framework
LLM-based label refinement
dataset mention extraction
forced displacement and FCV documents
synthetic and contrastive examples
R
Rafael Macalaba
Development Data Group, Office of the World Bank Group Chief Statistician, The World Bank
A
Aivin V. Solatorio
Development Data Group, Office of the World Bank Group Chief Statistician, The World Bank
P
Patrick Michael Brock
World Bank UNHCR Joint Data Center, UN City, Marmorvej 51, 2100 Copenhagen, Denmark
O
Olivier Dupriez
Development Data Group, Office of the World Bank Group Chief Statistician, The World Bank