Transfer Learning for Socioeconomic Estimation in Forced-Displacement Settings
研究利用迁移学习和地球观测数据,通过预训练的多模态时空视觉转换器,为受冲突影响地区的社会经济状况提供频繁更新的空间详细补充证据。
研究利用迁移学习和地球观测数据,通过预训练的多模态时空视觉转换器,为受冲突影响地区的社会经济状况提供频繁更新的空间详细补充证据。
本文提出了一种弱监督框架,利用轻量级模型和大型语言模型结合的方法,从非结构化文本中提取数据集引用,解决了系统识别数据集引用困难的问题。
研究通过直接提问和媒体生成两种方式评估了大语言模型在不同国家语境下的性别偏见问题,发现模型在本地语言生成中偏向男性。
This study addresses the sensitivity of structured metadata retrieval models to field ordering, which causes overreliance on positional cues rather than semantic field labels, thereby impairing discoverability in cross-lingual low-resource settings. To mitigate this issue, the authors propose Permutation-Invariant Fine-Tuning (PI-FT), a lightweight approach that randomizes field order and stochastically drops fields during data loading, encouraging the model to attend to semantic labels instead of positional patterns. Implemented with only two lines of code modification in the data loader, PI-FT enables a 118M-parameter CPU-based model to achieve an nDCG@10 of 0.707 on nearly 10,000 development statistics—outperforming all zero-shot baselines, including text-embedding-3-large—and reduces performance degradation under field-order perturbations from 7.4 to just 0.2 points, substantially enhancing robustness and generalization.
This study addresses a critical limitation in existing document layout analysis methods, which treat figures and tables as generic objects and thus fail to identify semantically valuable, reusable analytical visual content—referred to as “data snapshots”—in institutional documents. The work introduces the novel task of data snapshot extraction, presents a benchmark dataset comprising humanitarian reports and World Bank policy papers, and proposes an evaluation framework that integrates spatial localization with semantic annotation. Systematic evaluation of multiple open-source layout models reveals consistent shortcomings in handling institutional documents, including confusion between analytical and non-analytical content, fragmentation of composite charts, and lack of contextual awareness. By exposing the generalization bottlenecks of current models in operational documents, this research provides a foundation for future advancements through the public release of its dataset and codebase.
研究利用迁移学习和地球观测数据,通过预训练的多模态时空视觉转换器,为受冲突影响地区的社会经济状况提供频繁更新的空间详细补充证据。
本文提出了一种弱监督框架,利用轻量级模型和大型语言模型结合的方法,从非结构化文本中提取数据集引用,解决了系统识别数据集引用困难的问题。
研究通过直接提问和媒体生成两种方式评估了大语言模型在不同国家语境下的性别偏见问题,发现模型在本地语言生成中偏向男性。
This study addresses the sensitivity of structured metadata retrieval models to field ordering, which causes overreliance on positional cues rather than semantic field labels, thereby impairing discoverability in cross-lingual low-resource settings. To mitigate this issue, the authors propose Permutation-Invariant Fine-Tuning (PI-FT), a lightweight approach that randomizes field order and stochastically drops fields during data loading, encouraging the model to attend to semantic labels instead of positional patterns. Implemented with only two lines of code modification in the data loader, PI-FT enables a 118M-parameter CPU-based model to achieve an nDCG@10 of 0.707 on nearly 10,000 development statistics—outperforming all zero-shot baselines, including text-embedding-3-large—and reduces performance degradation under field-order perturbations from 7.4 to just 0.2 points, substantially enhancing robustness and generalization.
This study addresses a critical limitation in existing document layout analysis methods, which treat figures and tables as generic objects and thus fail to identify semantically valuable, reusable analytical visual content—referred to as “data snapshots”—in institutional documents. The work introduces the novel task of data snapshot extraction, presents a benchmark dataset comprising humanitarian reports and World Bank policy papers, and proposes an evaluation framework that integrates spatial localization with semantic annotation. Systematic evaluation of multiple open-source layout models reveals consistent shortcomings in handling institutional documents, including confusion between analytical and non-analytical content, fragmentation of composite charts, and lack of contextual awareness. By exposing the generalization bottlenecks of current models in operational documents, this research provides a foundation for future advancements through the public release of its dataset and codebase.