A two-stream network with global-local feature fusion for bone age assessment

📅 2025-12-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing bone age assessment methods struggle to simultaneously capture global skeletal structure and fine-grained epiphyseal details, limiting accuracy. To address this, we propose BoNet+, a novel dual-stream deep network that integrates Transformer-based global modeling with RFAConv-generated multi-scale adaptive local attention maps; feature fusion is further optimized via fine-tuned Inception-V3. This architecture overcomes the limitations of conventional single-stream CNNs by jointly modeling hierarchical anatomical cues. Evaluated on the RSNA and RHPE benchmarks, BoNet+ achieves state-of-the-art mean absolute errors (MAE) of 3.81 months and 5.65 months, respectively. The framework enables fully automated, high-accuracy, and clinically deployable bone age estimation—significantly reducing reliance on expert interpretation and advancing the practical deployment of intelligent bone age analysis in real-world clinical settings.

Technology Category

Application Category

📝 Abstract
Bone Age Assessment (BAA) is a widely used clinical technique that can accurately reflect an individual's growth and development level, as well as maturity. In recent years, although deep learning has advanced the field of bone age assessment, existing methods face challenges in efficiently balancing global features and local skeletal details. This study aims to develop an automated bone age assessment system based on a two-stream deep learning architecture to achieve higher accuracy in bone age assessment. We propose the BoNet+ model incorporating global and local feature extraction channels. A Transformer module is introduced into the global feature extraction channel to enhance the ability in extracting global features through multi-head self-attention mechanism. A RFAConv module is incorporated into the local feature extraction channel to generate adaptive attention maps within multiscale receptive fields, enhancing local feature extraction capabilities. Global and local features are concatenated along the channel dimension and optimized by an Inception-V3 network. The proposed method has been validated on the Radiological Society of North America (RSNA) and Radiological Hand Pose Estimation (RHPE) test datasets, achieving mean absolute errors (MAEs) of 3.81 and 5.65 months, respectively. These results are comparable to the state-of-the-art. The BoNet+ model reduces the clinical workload and achieves automatic, high-precision, and more objective bone age assessment.
Problem

Research questions and friction points this paper is trying to address.

Automates bone age assessment using two-stream deep learning
Balances global skeletal features with local anatomical details
Reduces clinical workload with high-precision objective evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Two-stream network with global-local feature fusion
Transformer module for global feature extraction
RFAConv module for local feature extraction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Q
Qiong Lou
School of Science, Zhejiang University of Science and Technology, Hangzhou, China
H
Han Yang
School of Science, Zhejiang University of Science and Technology, Hangzhou, China
F
Fang Lu
School of Science, Zhejiang University of Science and Technology, Hangzhou, China