Multi-View Reflective Surface Inspection via Semantic-Saliency Cross-Verification
为解决从单个固定视角难以检测反射智能手机盖玻璃缺陷的问题,提出了一种多视图检查框架,通过语义显著性交叉验证提高检测准确性。
为解决从单个固定视角难以检测反射智能手机盖玻璃缺陷的问题,提出了一种多视图检查框架,通过语义显著性交叉验证提高检测准确性。
Current multimodal large language models for ECG diagnosis suffer from limited interpretability, susceptibility to hallucinations, and deviations from clinical guidelines, undermining their clinical reliability. To address these issues, this work proposes a knowledge-anchored multimodal framework that, for the first time, distills authoritative ECG guidelines into structured explanatory knowledge offline and integrates this as a fixed module within the prompting pipeline. The approach combines CNN-based ECG feature extraction with Grad-CAM to produce class-specific heatmaps and factual evidence packages, guiding the model to generate structured diagnostic reports aligned with clinical standards. Evaluated on the PTB-XL test set, the method improves BERTScore for the impression section from 0.818 to 0.953, significantly enhancing guideline adherence, semantic quality, and interpretability while maintaining strong classification performance.
To address the high training costs and limited generalization of large models in Earth observation (EO) data mining, this paper proposes a lightweight paradigm based on compositional pre-trained foundation models. Instead of training from scratch, our method fuses domain-specific remote sensing models (e.g., Prithvi) with general-purpose vision models (e.g., Hiera, DOFA) via a feature-level ensemble architecture, and transfers the ensemble knowledge to a compact student model through knowledge distillation. Evaluated across 11 multi-resolution, multi-sensor, and multi-task benchmarks in GEO-Bench, our approach matches or surpasses individual large models in performance while significantly reducing training time and computational overhead. The core contribution is the empirical validation that synergistic small-model ensembles outperform monolithic large models—establishing a scalable, cost-efficient, and highly generalizable pathway for EO AI.
为解决从单个固定视角难以检测反射智能手机盖玻璃缺陷的问题,提出了一种多视图检查框架,通过语义显著性交叉验证提高检测准确性。
Current multimodal large language models for ECG diagnosis suffer from limited interpretability, susceptibility to hallucinations, and deviations from clinical guidelines, undermining their clinical reliability. To address these issues, this work proposes a knowledge-anchored multimodal framework that, for the first time, distills authoritative ECG guidelines into structured explanatory knowledge offline and integrates this as a fixed module within the prompting pipeline. The approach combines CNN-based ECG feature extraction with Grad-CAM to produce class-specific heatmaps and factual evidence packages, guiding the model to generate structured diagnostic reports aligned with clinical standards. Evaluated on the PTB-XL test set, the method improves BERTScore for the impression section from 0.818 to 0.953, significantly enhancing guideline adherence, semantic quality, and interpretability while maintaining strong classification performance.
To address the high training costs and limited generalization of large models in Earth observation (EO) data mining, this paper proposes a lightweight paradigm based on compositional pre-trained foundation models. Instead of training from scratch, our method fuses domain-specific remote sensing models (e.g., Prithvi) with general-purpose vision models (e.g., Hiera, DOFA) via a feature-level ensemble architecture, and transfers the ensemble knowledge to a compact student model through knowledge distillation. Evaluated across 11 multi-resolution, multi-sensor, and multi-task benchmarks in GEO-Bench, our approach matches or surpasses individual large models in performance while significantly reducing training time and computational overhead. The core contribution is the empirical validation that synergistic small-model ensembles outperform monolithic large models—establishing a scalable, cost-efficient, and highly generalizable pathway for EO AI.