Vision Transformer-Based Multi-Level Feature Fusion for Multi-Label Sewer Defect Classification

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究针对大规模多标签场景下排水管道缺陷自动分类问题,提出基于视觉Transformer的多层次特征融合方法及两种轻量级架构,有效平衡了分类精度与计算复杂度。
📝 Abstract
Automated classification of sewer defects is essential for infrastructure condition assessment and maintenance decision-making, but existing deep learning methods struggle to balance classification accuracy and computational complexity in large-scale multi-label scenarios. This study develops Sewer-Transformer-ML, a hierarchical vision Transformer with multi-level feature fusion, together with two lightweight architectures, Sewer-MobileNet-ML and Sewer-Mobile-TransNet, for resource-constrained inspection scenarios. On the Sewer-ML test set, Sewer-Transformer-ML-Base achieved an $F2_{\text{CIW}}$ of 65.68% and an $F1_{\text{Normal}}$ of 92.68%, ranking first on the public leaderboard and exceeding the second-ranked method by 7.6 percentage points in $F2_{\text{CIW}}$. Sewer-MobileNet-ML achieved an $F2_{\text{CIW}}$ of 65.73% with only 17 M parameters, representing an approximately 95% parameter reduction relative to the base model. Under the standard Sewer-Capsule data split, Sewer-Mobile-TransNet achieved 96.43% classification accuracy. When the training set was reduced to 1,177 images, pretraining on Sewer-ML consistently improved model performance. Ablation experiments further showed that direct concatenation was more effective for Transformer features, whereas attention-based fusion better supported multiscale CNN features. These findings provide a computational basis for automated sewer inspection, lightweight model design, and adaptation across civil infrastructure inspection platforms.
Problem

Research questions and friction points this paper is trying to address.

sewer defects
classification accuracy
computational complexity
multi-label scenarios
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision Transformer
Multi-Level Feature Fusion
Lightweight Architecture
Resource-Constrained Scenarios
Attention-Based Fusion
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xu Fang
Shenzhen Polytechnic University, Shenzhen, China
Zhuoran Wang
Zhuoran Wang
professor of University of Electronic Science and Technology of China
photonicselectronics
Qing Li
Qing Li
Pengcheng Laboratory
Indoor localizationdepth predictioncamera relocalizationneural reonstruction
S
Shengyu Zhang
Shenzhen Polytechnic University, Shenzhen, China
G
Guanzhi Deng
City University of Hong Kong, Hong Kong, China
J
Jianbiao He
Shenzhen Polytechnic University, Shenzhen, China
Q
Qingquan Li
Shenzhen University, Shenzhen, China