Verifiable Disaster Storylines and Causal Knowledge Graphs: A Citation-Grounded Pipeline from Heterogeneous Humanitarian Sources
为解决灾害初期信息合成难题,本文提出一种结合多种数据源的管线方法,利用检索增强生成技术提取结构化故事线和因果知识图谱,以支持应急响应。
为解决灾害初期信息合成难题,本文提出一种结合多种数据源的管线方法,利用检索增强生成技术提取结构化故事线和因果知识图谱,以支持应急响应。
The widespread adoption of remote work has exacerbated intra-urban inequalities in health risks, social interaction, and economic opportunity. Leveraging high-resolution hourly human mobility data and firm registration records, this study exploits variations in pandemic-related mobility restrictions as a quasi-natural experiment to investigate the drivers of workplace dependence through multidimensional regression and spatial heterogeneity models. The analysis reveals that industry type and firm productivity are key determinants. Moreover, income and gender effects are significantly moderated by distance from the city center, giving rise to a “service trap” in core urban areas: neighborhoods characterized by female-dominated employment and diverse income sources exhibit heightened reliance on in-person attendance, thereby extending remote-work disparities from the individual level to the broader urban ecosystem.
This study addresses the challenge of generating fine-grained subnational inferences in humanitarian contexts, where sparse survey data often prove insufficient. The authors propose a context-conditional normalizing flow generative model that integrates multisource geospatial and socioeconomic covariates as external context to learn full conditional distributions—rather than point estimates—of population characteristics. By leveraging rich contextual information, the model effectively enhances local population distribution estimates even under extreme data scarcity. Experiments across eight household survey datasets from six low- and middle-income countries demonstrate that the approach substantially improves subnational estimation accuracy, with performance systematically increasing as the richness of contextual information grows.
This study addresses geographic and socioeconomic biases in existing automated systems for extracting locations from humanitarian texts, which result in uneven visibility of crisis-affected regions. To mitigate this, the authors propose a two-stage framework: first employing a few-shot large language model (LLM) for named entity recognition, followed by an agent-based, context-aware geocoding module for precise toponym disambiguation. This approach represents the first integration of LLMs with fairness principles in humanitarian geospatial analysis. Evaluated on an expanded HumSet dataset, the method significantly outperforms current rule-based and pretrained systems, achieving higher overall accuracy while notably improving location recognition in underrepresented regions, thereby advancing more inclusive and equitable humanitarian response efforts.
Humanitarian decision-making urgently requires timely, accurate, and verifiable situational reports, yet current practices rely heavily on manual processes—resulting in low efficiency and inconsistent quality. This paper introduces the first end-to-end large language model (LLM) framework for fully automating the transformation of heterogeneous, multi-source humanitarian documents into structured, verifiable, and action-oriented reports. Our method innovatively integrates semantic clustering, evidence-grounded question generation, and a multi-level expert-simulation evaluation paradigm. It ensures explainability, verifiability, and operational utility across key stages: event aggregation, question generation, retrieval-augmented answer extraction, multi-granularity summarization, and executive summary generation. Evaluated on 13 real-world humanitarian incidents, our framework achieves 84.7% and 86.3% relevance scores for generated questions and answers, respectively; citation precision and recall both exceed 76%; and human-AI collaborative evaluation yields an F1-score >0.80—significantly outperforming all baselines.
为解决灾害初期信息合成难题,本文提出一种结合多种数据源的管线方法,利用检索增强生成技术提取结构化故事线和因果知识图谱,以支持应急响应。
The widespread adoption of remote work has exacerbated intra-urban inequalities in health risks, social interaction, and economic opportunity. Leveraging high-resolution hourly human mobility data and firm registration records, this study exploits variations in pandemic-related mobility restrictions as a quasi-natural experiment to investigate the drivers of workplace dependence through multidimensional regression and spatial heterogeneity models. The analysis reveals that industry type and firm productivity are key determinants. Moreover, income and gender effects are significantly moderated by distance from the city center, giving rise to a “service trap” in core urban areas: neighborhoods characterized by female-dominated employment and diverse income sources exhibit heightened reliance on in-person attendance, thereby extending remote-work disparities from the individual level to the broader urban ecosystem.
This study addresses the challenge of generating fine-grained subnational inferences in humanitarian contexts, where sparse survey data often prove insufficient. The authors propose a context-conditional normalizing flow generative model that integrates multisource geospatial and socioeconomic covariates as external context to learn full conditional distributions—rather than point estimates—of population characteristics. By leveraging rich contextual information, the model effectively enhances local population distribution estimates even under extreme data scarcity. Experiments across eight household survey datasets from six low- and middle-income countries demonstrate that the approach substantially improves subnational estimation accuracy, with performance systematically increasing as the richness of contextual information grows.
This study addresses geographic and socioeconomic biases in existing automated systems for extracting locations from humanitarian texts, which result in uneven visibility of crisis-affected regions. To mitigate this, the authors propose a two-stage framework: first employing a few-shot large language model (LLM) for named entity recognition, followed by an agent-based, context-aware geocoding module for precise toponym disambiguation. This approach represents the first integration of LLMs with fairness principles in humanitarian geospatial analysis. Evaluated on an expanded HumSet dataset, the method significantly outperforms current rule-based and pretrained systems, achieving higher overall accuracy while notably improving location recognition in underrepresented regions, thereby advancing more inclusive and equitable humanitarian response efforts.
Humanitarian decision-making urgently requires timely, accurate, and verifiable situational reports, yet current practices rely heavily on manual processes—resulting in low efficiency and inconsistent quality. This paper introduces the first end-to-end large language model (LLM) framework for fully automating the transformation of heterogeneous, multi-source humanitarian documents into structured, verifiable, and action-oriented reports. Our method innovatively integrates semantic clustering, evidence-grounded question generation, and a multi-level expert-simulation evaluation paradigm. It ensures explainability, verifiability, and operational utility across key stages: event aggregation, question generation, retrieval-augmented answer extraction, multi-granularity summarization, and executive summary generation. Evaluated on 13 real-world humanitarian incidents, our framework achieves 84.7% and 86.3% relevance scores for generated questions and answers, respectively; citation precision and recall both exceed 76%; and human-AI collaborative evaluation yields an F1-score >0.80—significantly outperforming all baselines.