Vision Transformer-Based Multi-Level Feature Fusion for Multi-Label Sewer Defect Classification
研究针对大规模多标签场景下排水管道缺陷自动分类问题,提出基于视觉Transformer的多层次特征融合方法及两种轻量级架构,有效平衡了分类精度与计算复杂度。
研究针对大规模多标签场景下排水管道缺陷自动分类问题,提出基于视觉Transformer的多层次特征融合方法及两种轻量级架构,有效平衡了分类精度与计算复杂度。
本文提出LetOccVote方法,通过跨帧投票机制改进几何和语义监督,以解决弱监督3D占据预测中伪标签不可靠的问题。
This study addresses privacy concerns and high user burden in conventional dietary monitoring by proposing an audio-visual-free ear-worn sensing system. Integrating PPG and IMU signals, the framework employs a two-stage event-gating mechanism, five-class behavior classification, and lightweight personalized memory matching to achieve accurate recognition with minimal intrusion. Experimental results demonstrate that the pooled model attains 70.99% accuracy, while multimodal fusion achieves 85.13% accuracy with only 60% calibration data, significantly outperforming unimodal approaches. As the first audio-visual-free paradigm for dietary sensing, this work effectively balances privacy preservation with recognition precision, offering a novel pathway for wearable health interventions.
This study addresses the high computational cost of double machine learning (DML) with large-scale observational data and the inadequacy of simple random subsampling, which often results in insufficient covariate coverage and imbalance between treatment and control groups. The authors propose UD-DML, a novel subsampling approach that incorporates uniform design principles into DML for the first time. The method first applies PCA rotation to covariates, then constructs a low-discrepancy skeleton based on a hybrid discrepancy criterion, and finally uses KD-trees to match nearest-neighbor treated and control units to each skeleton point, yielding a representative and balanced subsample. Theoretical analysis establishes that UD-DML retains √n-asymptotic normality even when the subsample size is much smaller than the full dataset. Experiments demonstrate that, compared to uniform sampling, UD-DML achieves lower RMSE, narrower confidence intervals, and more reliable coverage—particularly under low overlap and model misspecification.
This work addresses the high computational cost and sensitivity to input noise inherent in traditional capsule networks due to iterative dynamic routing. The authors introduce, for the first time, the information bottleneck principle into capsule networks and propose a one-shot variational aggregation mechanism that eliminates iterative routing altogether. By leveraging global context compression and class-specific variational autoencoders, the method directly infers latent capsules in a single pass. This approach substantially enhances both efficiency and robustness: on benchmarks such as MNIST, it achieves an average accuracy improvement of over 14% under noisy conditions, accelerates training by 2.54×, increases inference throughput by 3.64×, reduces parameter count by 4.66%, and maintains high accuracy on clean data.
研究针对大规模多标签场景下排水管道缺陷自动分类问题,提出基于视觉Transformer的多层次特征融合方法及两种轻量级架构,有效平衡了分类精度与计算复杂度。
本文提出LetOccVote方法,通过跨帧投票机制改进几何和语义监督,以解决弱监督3D占据预测中伪标签不可靠的问题。
This study addresses privacy concerns and high user burden in conventional dietary monitoring by proposing an audio-visual-free ear-worn sensing system. Integrating PPG and IMU signals, the framework employs a two-stage event-gating mechanism, five-class behavior classification, and lightweight personalized memory matching to achieve accurate recognition with minimal intrusion. Experimental results demonstrate that the pooled model attains 70.99% accuracy, while multimodal fusion achieves 85.13% accuracy with only 60% calibration data, significantly outperforming unimodal approaches. As the first audio-visual-free paradigm for dietary sensing, this work effectively balances privacy preservation with recognition precision, offering a novel pathway for wearable health interventions.
This study addresses the high computational cost of double machine learning (DML) with large-scale observational data and the inadequacy of simple random subsampling, which often results in insufficient covariate coverage and imbalance between treatment and control groups. The authors propose UD-DML, a novel subsampling approach that incorporates uniform design principles into DML for the first time. The method first applies PCA rotation to covariates, then constructs a low-discrepancy skeleton based on a hybrid discrepancy criterion, and finally uses KD-trees to match nearest-neighbor treated and control units to each skeleton point, yielding a representative and balanced subsample. Theoretical analysis establishes that UD-DML retains √n-asymptotic normality even when the subsample size is much smaller than the full dataset. Experiments demonstrate that, compared to uniform sampling, UD-DML achieves lower RMSE, narrower confidence intervals, and more reliable coverage—particularly under low overlap and model misspecification.
This work addresses the high computational cost and sensitivity to input noise inherent in traditional capsule networks due to iterative dynamic routing. The authors introduce, for the first time, the information bottleneck principle into capsule networks and propose a one-shot variational aggregation mechanism that eliminates iterative routing altogether. By leveraging global context compression and class-specific variational autoencoders, the method directly infers latent capsules in a single pass. This approach substantially enhances both efficiency and robustness: on benchmarks such as MNIST, it achieves an average accuracy improvement of over 14% under noisy conditions, accelerates training by 2.54×, increases inference throughput by 3.64×, reduces parameter count by 4.66%, and maintains high accuracy on clean data.