Foundation Model-driven Key Anatomy Frame Selection for Blind-sweep Ultrasound Fetal Birth Weight Estimation
This study addresses the challenge of operator-independent fetal birth weight estimation from unguided (“blind-scan”) ultrasound videos in resource-limited settings. The authors propose a vision–language foundation model-based framework that automatically selects key anatomical frames from videos acquired within 48 hours before delivery and incorporates a redundancy-aware feature compression module to preserve task-relevant anatomical information for end-to-end weight regression. As the first method to estimate fetal weight directly from blind-scan videos, this approach achieves a mean absolute error of 161.3 grams on a prospective cohort of 839 cases, with 90.23% and 100% of estimates falling within 10% and 15% absolute percentage error, respectively—significantly outperforming the conventional Hadlock method and current strong baselines.