AI-CNet3D: An Anatomically-Informed Cross-Attention Network with Multi-Task Consistency Fine-tuning for 3D Glaucoma Classification

📅 2025-10-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Structural details in 3D OCT volumes—critical for glaucoma classification—are often lost during 2D projection-based compression. Method: We propose a lightweight, anatomy-aware 3D classification framework featuring (i) Channel-wise Anatomy-aware Representation (CARE) modules that explicitly encode retinal and optic nerve head anatomical priors; (ii) cross-attention mechanisms to model inter-regional structural asymmetry; and (iii) multi-task consistency fine-tuning coupled with Grad-CAM-guided interpretability optimization. Contribution/Results: To our knowledge, this is the first work to jointly embed anatomical constraints and attention mechanisms into 3D OCT analysis. Evaluated on two large public datasets, our method surpasses state-of-the-art models with significant accuracy gains (e.g., +3.2% average accuracy), reduces parameter count by two orders of magnitude, and achieves efficient inference without sacrificing performance. Moreover, it enhances alignment between model decisions and clinically relevant anatomical features, substantially improving interpretability and clinical trustworthiness.

Technology Category

Application Category

📝 Abstract
Glaucoma is a progressive eye disease that leads to optic nerve damage, causing irreversible vision loss if left untreated. Optical coherence tomography (OCT) has become a crucial tool for glaucoma diagnosis, offering high-resolution 3D scans of the retina and optic nerve. However, the conventional practice of condensing information from 3D OCT volumes into 2D reports often results in the loss of key structural details. To address this, we propose a novel hybrid deep learning model that integrates cross-attention mechanisms into a 3D convolutional neural network (CNN), enabling the extraction of critical features from the superior and inferior hemiretinas, as well as from the optic nerve head (ONH) and macula, within OCT volumes. We introduce Channel Attention REpresentations (CAREs) to visualize cross-attention outputs and leverage them for consistency-based multi-task fine-tuning, aligning them with Gradient-Weighted Class Activation Maps (Grad-CAMs) from the CNN's final convolutional layer to enhance performance, interpretability, and anatomical coherence. We have named this model AI-CNet3D (AI-`See'-Net3D) to reflect its design as an Anatomically-Informed Cross-attention Network operating on 3D data. By dividing the volume along two axes and applying cross-attention, our model enhances glaucoma classification by capturing asymmetries between the hemiretinal regions while integrating information from the optic nerve head and macula. We validate our approach on two large datasets, showing that it outperforms state-of-the-art attention and convolutional models across all key metrics. Finally, our model is computationally efficient, reducing the parameter count by one-hundred--fold compared to other attention mechanisms while maintaining high diagnostic performance and comparable GFLOPS.
Problem

Research questions and friction points this paper is trying to address.

Classifying glaucoma using 3D OCT scans by preserving structural details
Addressing information loss from condensing 3D volumes into 2D reports
Capturing retinal asymmetries and integrating optic nerve/macula information
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-attention 3D CNN integrates optic nerve and macula features
Channel Attention Representations enable interpretable multi-task fine-tuning
Anatomically-informed model reduces parameters hundredfold while maintaining performance
💼 Related Jobs
No related jobs found.
R
Roshan Kenia
Department of Computer Science, Columbia University, New York, NY, USA
A
Anfei Li
Department of Ophthalmology, Columbia University Irving Medical Center, New York, NY, USA
R
Rishabh Srivastava
Department of Computer Science, Columbia University, New York, NY, USA
Kaveri A. Thakoor
Kaveri A. Thakoor
Columbia University