Institution profile

World Bank

Academic institutionnorthamerica · us
Official website
Research library14linked papers
Opportunities0open roles
Selected work

Representative Papers

Field Order Should Not Matter: Permutation-Invariant Embedding Model Fine-Tuning for Structured Metadata Retrieval

Jun 29, 2026

This study addresses the sensitivity of structured metadata retrieval models to field ordering, which causes overreliance on positional cues rather than semantic field labels, thereby impairing discoverability in cross-lingual low-resource settings. To mitigate this issue, the authors propose Permutation-Invariant Fine-Tuning (PI-FT), a lightweight approach that randomizes field order and stochastically drops fields during data loading, encouraging the model to attend to semantic labels instead of positional patterns. Implemented with only two lines of code modification in the data loader, PI-FT enables a 118M-parameter CPU-based model to achieve an nDCG@10 of 0.707 on nearly 10,000 development statistics—outperforming all zero-shot baselines, including text-embedding-3-large—and reduces performance degradation under field-order perturbations from 7.4 to just 0.2 points, substantially enhancing robustness and generalization.

0 citationsRead paper

Benchmarking Open-Source Layout Detection Models for Data Snapshot Extraction from Institutional Documents

Jun 04, 2026

This study addresses a critical limitation in existing document layout analysis methods, which treat figures and tables as generic objects and thus fail to identify semantically valuable, reusable analytical visual content—referred to as “data snapshots”—in institutional documents. The work introduces the novel task of data snapshot extraction, presents a benchmark dataset comprising humanitarian reports and World Bank policy papers, and proposes an evaluation framework that integrates spatial localization with semantic annotation. Systematic evaluation of multiple open-source layout models reveals consistent shortcomings in handling institutional documents, including confusion between analytical and non-analytical content, fragmentation of composite charts, and lack of contextual awareness. By exposing the generalization bottlenecks of current models in operational documents, this research provides a foundation for future advancements through the public release of its dataset and codebase.

0 citationsRead paper
Recent publications

Latest Papers

Field Order Should Not Matter: Permutation-Invariant Embedding Model Fine-Tuning for Structured Metadata Retrieval

Jun 29, 2026

This study addresses the sensitivity of structured metadata retrieval models to field ordering, which causes overreliance on positional cues rather than semantic field labels, thereby impairing discoverability in cross-lingual low-resource settings. To mitigate this issue, the authors propose Permutation-Invariant Fine-Tuning (PI-FT), a lightweight approach that randomizes field order and stochastically drops fields during data loading, encouraging the model to attend to semantic labels instead of positional patterns. Implemented with only two lines of code modification in the data loader, PI-FT enables a 118M-parameter CPU-based model to achieve an nDCG@10 of 0.707 on nearly 10,000 development statistics—outperforming all zero-shot baselines, including text-embedding-3-large—and reduces performance degradation under field-order perturbations from 7.4 to just 0.2 points, substantially enhancing robustness and generalization.

0 citationsRead paper

Benchmarking Open-Source Layout Detection Models for Data Snapshot Extraction from Institutional Documents

Jun 04, 2026

This study addresses a critical limitation in existing document layout analysis methods, which treat figures and tables as generic objects and thus fail to identify semantically valuable, reusable analytical visual content—referred to as “data snapshots”—in institutional documents. The work introduces the novel task of data snapshot extraction, presents a benchmark dataset comprising humanitarian reports and World Bank policy papers, and proposes an evaluation framework that integrates spatial localization with semantic annotation. Systematic evaluation of multiple open-source layout models reveals consistent shortcomings in handling institutional documents, including confusion between analytical and non-analytical content, fragmentation of composite charts, and lack of contextual awareness. By exposing the generalization bottlenecks of current models in operational documents, this research provides a foundation for future advancements through the public release of its dataset and codebase.

0 citationsRead paper