CARDIAG: A Dense Segment Classification Benchmark of Deep Learning Architectures for Coronary Angiography

📅 2026-07-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of a standardized benchmark for pixel-level SYNTAX category segmentation in coronary angiography, which has hindered fair comparison of deep learning models. To this end, we introduce CARDIAG, a multicenter, multilabel dense classification benchmark that establishes the first standardized evaluation framework tailored for SYNTAX scoring. It encompasses 24 state-of-the-art architectures—including ConvNeXt V2, Mamba U-Net, DeepLabV3+, and FPN—and provides a high-quality dataset annotated with SYNTAX labels, segmentation masks, uncertainty estimates, and DICOM metadata. Through rigorous data partitioning and multidimensional metrics such as diameter error and centerline quality, we systematically assess model calibration, generalization, and data efficiency. The best single model (ConvNeXt V2 + DeepLabV3+) achieves a macro F₁ score of 0.456, improving to 0.479 with ensembling, while all models demonstrate strong calibration and robust cross-center performance.
📝 Abstract
Accurate pixel-level classification of coronary angiograms is critical for cardiovascular disease assessment, yet the field lacks standardized evaluation protocols. In this work we demonstrate a new benchmark for the assessment of deep learning models which densely classify pixels of coronary angiograms to one of SYNTAX classes (or background). The evaluation covers 24 distinct architectures starting with classic convnets to recent state-space-based vision algorithms. We release CARDIAG - a multi-center, multi-label dataset which we carefully split to reliably compute metrics, accounting for diameter error, overlap, centerline quality and calibration. The data contains SYNTAX labels, binary, uncertainty and segmentation masks as well as intermediate frames together with the selected non-sensitive DICOM metadata. From the multitude of algorithms, we nominate ConvNeXt V2 encoder with DeepLab V3 Plus decoder as the best performing, achieving macro $F_1=0.456$, which we then ensemble with Mamba U-Net and Feature Pyramid Network, for an increased $F_1=0.479$. We demonstrate all the architectures to be well calibrated and determine the generalization of the top 5 methods, together with the data efficiency of these architectures. We highlight the importance of both high-resolution and low-resolution features in encoding. We also demonstrate the model correctness in the context of patient demographic, vessel sides and projection angle configurations. Overall the released benchmark allows for future studies to robustly and rigorously assess the proposals, not only for SYNTAX segmentation, but lesion detection and many more.
Problem

Research questions and friction points this paper is trying to address.

coronary angiography
pixel-level classification
benchmark
SYNTAX scoring
dense segmentation
Innovation

Methods, ideas, or system contributions that make the work stand out.

dense segmentation
coronary angiography
deep learning benchmark
SYNTAX classification
multi-center dataset
🔎 Similar Papers
No similar papers found.
D
Dominik Bernard Lau
Gdańsk University of Technology
H
Hubert Malinowski
Gdańsk University of Technology
J
Jerzy Szyjut
Gdańsk University of Technology
A
Adam Brzeski
Gdańsk University of Technology
Tomasz Dziubich
Tomasz Dziubich
Politechnika Gdańska
przetwarzanie wszechobecne
R
Radosław Targoński
Medical University of Gdańsk
T
Tomasz Figatowski
Medical University of Gdańsk
N
Natalia Zielińska
Medical University of Gdańsk