Towards Investigating Residual Hearing Loss: Quantification of Fibrosis in a Novel Cochlear OCT Dataset
研究使用计算机视觉方法,特别是改进的UNET架构(2D-OCT-UNET),来量化豚鼠耳蜗植入后产生的纤维化情况,以期减少纤维化负担并改善人工耳蜗患者的结果。
研究使用计算机视觉方法,特别是改进的UNET架构(2D-OCT-UNET),来量化豚鼠耳蜗植入后产生的纤维化情况,以期减少纤维化负担并改善人工耳蜗患者的结果。
This study addresses the challenge of artifacts in optical coherence tomography angiography (OCTA) that severely compromise accurate quantification of retinal blood flow and non-perfused areas. Existing approaches are often limited to two-dimensional processing or fail to fully exploit three-dimensional vascular architecture. To overcome these limitations, this work proposes an end-to-end deep learning algorithm that, for the first time in OCTA, integrates 3D vascular structure by reconstructing a central B-scan from three adjacent input B-scans, thereby preserving spatial resolution while substantially enhancing microvascular fidelity. The model employs an EfficientNet-B5 encoder, a decoder augmented with spatial-channel parallel squeeze-and-excitation modules, and skip connections, trained under supervision using ground truth derived from multi-scan averaging. Experimental results demonstrate significant improvements: PSNR increased to 26.16 ± 1.26 (from 22.23 ± 0.78), SSIM reached 0.91 ± 0.02 (from 0.72 ± 0.03), and 2D and 3D Dice coefficients improved by at least 3.8% and 51.2%, respectively (both p < 0.001).
This study addresses the challenge of automated, precise staging diagnosis of age-related macular degeneration (AMD) by systematically comparing the performance of deep learning models based on three distinct input modalities: biomarker maps, 2D en face projections, and 3D OCT/OCTA volumetric data. Utilizing an EfficientNet architecture combined with normalization, data augmentation, and five-fold cross-validation, the study evaluates these inputs on a four-stage AMD classification task. Results demonstrate strong agreement between all models and expert annotations (QWK ≥ 0.83), with the biomarker-based model achieving the highest overall and most balanced performance (QWK = 0.85 ± 0.03) and an F1-score of 0.59 ± 0.14 for early AMD detection. The 2D model excels in identifying non-AMD cases with the highest precision (0.79 ± 0.06). This work provides critical guidance for selecting input strategies in automated AMD staging.
This work addresses the challenge of batch effects in histopathology images—arising from variations in staining protocols and scanning devices—that severely hinder model generalization across clinical sites. To this end, the authors propose Latent Manifold Compression (LMC), an unsupervised representation learning framework that explicitly compresses stain-induced latent manifolds within a single source domain to construct a batch-invariant embedding space. Notably, LMC enables cross-batch image normalization without requiring any target-domain data, thereby eliminating the need for paired or multi-domain training samples. Evaluated on three public and internal benchmarks, the method significantly reduces inter-batch separation and consistently outperforms state-of-the-art approaches in both cross-batch classification and detection tasks, substantially improving model generalization.
Existing metrics struggle to distinguish geometric transformations that affect data cluster fusion from those—such as global scaling or sampling layout changes—that do not, leading to inaccurate quantification of fusion and separability in representation spaces. To address this, this work proposes the Cross-Fusion Distance (CFD), which, grounded in geometric invariance theory, theoretically disentangles and isolates fusion-relevant from fusion-irrelevant geometric factors for the first time, enabling precise measurement of fusion extent. CFD exhibits linear computational complexity and, in synthetic experiments, demonstrates high sensitivity to fusion-altering transformations while remaining invariant to irrelevant ones. Moreover, on real-world domain-shift datasets, CFD shows stronger correlation with downstream task generalization performance than existing metrics.
研究使用计算机视觉方法,特别是改进的UNET架构(2D-OCT-UNET),来量化豚鼠耳蜗植入后产生的纤维化情况,以期减少纤维化负担并改善人工耳蜗患者的结果。
This study addresses the challenge of artifacts in optical coherence tomography angiography (OCTA) that severely compromise accurate quantification of retinal blood flow and non-perfused areas. Existing approaches are often limited to two-dimensional processing or fail to fully exploit three-dimensional vascular architecture. To overcome these limitations, this work proposes an end-to-end deep learning algorithm that, for the first time in OCTA, integrates 3D vascular structure by reconstructing a central B-scan from three adjacent input B-scans, thereby preserving spatial resolution while substantially enhancing microvascular fidelity. The model employs an EfficientNet-B5 encoder, a decoder augmented with spatial-channel parallel squeeze-and-excitation modules, and skip connections, trained under supervision using ground truth derived from multi-scan averaging. Experimental results demonstrate significant improvements: PSNR increased to 26.16 ± 1.26 (from 22.23 ± 0.78), SSIM reached 0.91 ± 0.02 (from 0.72 ± 0.03), and 2D and 3D Dice coefficients improved by at least 3.8% and 51.2%, respectively (both p < 0.001).
This study addresses the challenge of automated, precise staging diagnosis of age-related macular degeneration (AMD) by systematically comparing the performance of deep learning models based on three distinct input modalities: biomarker maps, 2D en face projections, and 3D OCT/OCTA volumetric data. Utilizing an EfficientNet architecture combined with normalization, data augmentation, and five-fold cross-validation, the study evaluates these inputs on a four-stage AMD classification task. Results demonstrate strong agreement between all models and expert annotations (QWK ≥ 0.83), with the biomarker-based model achieving the highest overall and most balanced performance (QWK = 0.85 ± 0.03) and an F1-score of 0.59 ± 0.14 for early AMD detection. The 2D model excels in identifying non-AMD cases with the highest precision (0.79 ± 0.06). This work provides critical guidance for selecting input strategies in automated AMD staging.
This work addresses the challenge of batch effects in histopathology images—arising from variations in staining protocols and scanning devices—that severely hinder model generalization across clinical sites. To this end, the authors propose Latent Manifold Compression (LMC), an unsupervised representation learning framework that explicitly compresses stain-induced latent manifolds within a single source domain to construct a batch-invariant embedding space. Notably, LMC enables cross-batch image normalization without requiring any target-domain data, thereby eliminating the need for paired or multi-domain training samples. Evaluated on three public and internal benchmarks, the method significantly reduces inter-batch separation and consistently outperforms state-of-the-art approaches in both cross-batch classification and detection tasks, substantially improving model generalization.
Existing metrics struggle to distinguish geometric transformations that affect data cluster fusion from those—such as global scaling or sampling layout changes—that do not, leading to inaccurate quantification of fusion and separability in representation spaces. To address this, this work proposes the Cross-Fusion Distance (CFD), which, grounded in geometric invariance theory, theoretically disentangles and isolates fusion-relevant from fusion-irrelevant geometric factors for the first time, enabling precise measurement of fusion extent. CFD exhibits linear computational complexity and, in synthetic experiments, demonstrates high sensitivity to fusion-altering transformations while remaining invariant to irrelevant ones. Moreover, on real-world domain-shift datasets, CFD shows stronger correlation with downstream task generalization performance than existing metrics.