EVADE: Evidence-Verified Agentic Diagnosis with Escape
为解决医疗视觉-语言模型的不可靠问题,EVADE通过在不确定时局部放大图像并验证不同视图间的一致性来提高模型的安全性和准确性。
为解决医疗视觉-语言模型的不可靠问题,EVADE通过在不确定时局部放大图像并验证不同视图间的一致性来提高模型的安全性和准确性。
为解决酒店业可持续性报告数据粗糙问题,通过构建基于LoRaWAN的物联网平台LoRIS,实现对资源消耗、环境条件和客人行为的高分辨率监测。
This study addresses the critical need for high accuracy, interpretability, and scalability in online banking fraud detection by leveraging real-world transaction data from India. Through exploratory data analysis and SMOTE-based oversampling to mitigate class imbalance, the authors systematically evaluate five deep learning models: DNN, GRU, LSTM, 1D-CNN, and TabNet. Notably, they harness TabNet’s intrinsic sparse feature selection mechanism to simultaneously enhance model interpretability and generalization. Experimental results demonstrate that TabNet achieves a 97.39% accuracy and a 0.9739 ROC-AUC under three-fold cross-validation, significantly outperforming baseline models. The approach effectively reduces both false positives and false negatives, supports real-time deployment, and satisfies stringent financial regulatory requirements for model transparency.
This study addresses the significant challenges in semantic segmentation for automated disassembly of electrolyzer components, which arise from high visual similarity among materials, spectral overlap, irregular shapes, and severe class imbalance. To overcome these issues, the authors propose HREM-Net, a dual-branch deep network that effectively fuses hyperspectral and RGB imagery through a novel adaptive gated cross-modal fusion mechanism. The architecture integrates efficient channel attention, coordinate attention, Mobile Inverted Bottleneck blocks, and an atrous spatial pyramid pooling module, further enhanced by a composite loss function to strengthen multimodal feature synergy. Evaluated on the Electrolyzers-HSI dataset, the method achieves a mean class accuracy of 91.66% and an mIoU of 0.82, while demonstrating strong generalization on PCB-Vision with 96.91% accuracy and 0.93 mIoU.
Existing evaluation metrics exhibit inconsistent behavior in multimodal machine unlearning tasks, making it difficult to reliably assess unlearning efficacy. This work systematically analyzes the conflicting rankings produced by five widely used metrics across three visual question answering (VQA) benchmarks and proposes a Unified Quality Score (UQS) that achieves more stable performance ranking by weighting each metric according to its distance correlation with an idealized reference model. Empirical evaluation on 36 variants of LLaVA-1.5-7B and BLIP-2 models reveals substantial discrepancies in metric-induced rankings. The proposed UQS demonstrates high stability under 100 random perturbations, achieving a Kendall’s τ of 0.647 ± 0.262. The authors publicly release the benchmark suite, model checkpoints, and an interactive leaderboard to support reproducible research in multimodal unlearning.
为解决医疗视觉-语言模型的不可靠问题,EVADE通过在不确定时局部放大图像并验证不同视图间的一致性来提高模型的安全性和准确性。
为解决酒店业可持续性报告数据粗糙问题,通过构建基于LoRaWAN的物联网平台LoRIS,实现对资源消耗、环境条件和客人行为的高分辨率监测。
This study addresses the critical need for high accuracy, interpretability, and scalability in online banking fraud detection by leveraging real-world transaction data from India. Through exploratory data analysis and SMOTE-based oversampling to mitigate class imbalance, the authors systematically evaluate five deep learning models: DNN, GRU, LSTM, 1D-CNN, and TabNet. Notably, they harness TabNet’s intrinsic sparse feature selection mechanism to simultaneously enhance model interpretability and generalization. Experimental results demonstrate that TabNet achieves a 97.39% accuracy and a 0.9739 ROC-AUC under three-fold cross-validation, significantly outperforming baseline models. The approach effectively reduces both false positives and false negatives, supports real-time deployment, and satisfies stringent financial regulatory requirements for model transparency.
This study addresses the significant challenges in semantic segmentation for automated disassembly of electrolyzer components, which arise from high visual similarity among materials, spectral overlap, irregular shapes, and severe class imbalance. To overcome these issues, the authors propose HREM-Net, a dual-branch deep network that effectively fuses hyperspectral and RGB imagery through a novel adaptive gated cross-modal fusion mechanism. The architecture integrates efficient channel attention, coordinate attention, Mobile Inverted Bottleneck blocks, and an atrous spatial pyramid pooling module, further enhanced by a composite loss function to strengthen multimodal feature synergy. Evaluated on the Electrolyzers-HSI dataset, the method achieves a mean class accuracy of 91.66% and an mIoU of 0.82, while demonstrating strong generalization on PCB-Vision with 96.91% accuracy and 0.93 mIoU.
Existing evaluation metrics exhibit inconsistent behavior in multimodal machine unlearning tasks, making it difficult to reliably assess unlearning efficacy. This work systematically analyzes the conflicting rankings produced by five widely used metrics across three visual question answering (VQA) benchmarks and proposes a Unified Quality Score (UQS) that achieves more stable performance ranking by weighting each metric according to its distance correlation with an idealized reference model. Empirical evaluation on 36 variants of LLaVA-1.5-7B and BLIP-2 models reveals substantial discrepancies in metric-induced rankings. The proposed UQS demonstrates high stability under 100 random perturbations, achieving a Kendall’s τ of 0.647 ± 0.262. The authors publicly release the benchmark suite, model checkpoints, and an interactive leaderboard to support reproducible research in multimodal unlearning.