VT-3DAD: Cross-Category 3D Anomaly Detection via Visual-Text Normal Space Alignment
This work addresses the challenge of cross-category 3D point cloud anomaly detection using only a few normal samples by proposing a training-free framework that, for the first time, introduces vision–text normal space alignment to this task. The method extracts visual features from multi-view depth maps using a frozen CLIP model and constructs semantic normal anchors via depth-aware and 3D-aware textual prompts. Anomaly scores are computed by fusing visual and semantic deviations. Evaluated on ShapeNetPart, the approach achieves an average single-sample AUC-ROC of 94.80%, outperforming a purely visual baseline by 2.31% and reducing the standard deviation from 5.64 to 3.41. These results demonstrate the effectiveness of the proposed method in jointly modeling geometric and semantic normality, as well as its strong cross-category generalization capability.