Streaming Hierarchical Inference with Tabular Foundation Models
为解决高吞吐量数据流中部署TFM的通信开销和延迟问题,提出HINT框架,结合边缘检索与云端TFM推理,平衡预测性能与通信成本。
为解决高吞吐量数据流中部署TFM的通信开销和延迟问题,提出HINT框架,结合边缘检索与云端TFM推理,平衡预测性能与通信成本。
This work addresses the gap in machine unlearning research by introducing XGBoost-Forget, the first unlearning method tailored for XGBoost models in tabular network intrusion detection scenarios. Unlike existing approaches primarily designed for deep learning and image data, XGBoost-Forget efficiently removes specified intrusion data points without requiring full model retraining. The method incorporates a customized unlearning mechanism specifically designed for tabular network traffic data and is validated on real-world datasets such as IoT-23 and GeNIS. Experimental results demonstrate that XGBoost-Forget achieves substantial gains in unlearning efficiency while preserving predictive performance nearly equivalent to that of the original model, thereby offering a practical unlearning solution for security applications driven by tabular data.
This study systematically investigates the sources of vulnerability in Retrieval-Augmented Generation (RAG) systems under poisoning attacks. Through a comprehensive full-factorial experiment encompassing 432 configurations, it evaluates the impact of datasets, retriever types (dense, graph-based, and BM25), retrieval depth, knowledge base composition, chunking strategies, and generation models on system robustness. The findings reveal that RAG’s susceptibility arises from complex interactions among retrieval, generation, and knowledge base configurations rather than from any single component’s deficiency. Dense and graph-based retrievers significantly outperform BM25, while increasing retrieval depth or replicating poisoned content across multiple knowledge sources substantially elevates attack success rates. Conversely, incorporating clean data from diverse sources effectively mitigates such attacks. This work is the first to uncover the key factors governing RAG’s robustness against poisoning and their underlying coupling mechanisms.
This study addresses the challenge that local interpretability methods often produce seemingly plausible yet unfaithful explanations for complex tabular data. To rigorously evaluate explanation fidelity, robustness, and complexity, the authors construct a comprehensive benchmarking framework encompassing multiple models and datasets, incorporating for the first time a prediction-consistency grouping strategy. They systematically assess prominent methods—including LIME, Kernel SHAP, and feature ablation—across 32 tabular datasets. The findings reveal that explanation quality shows no significant correlation with model accuracy; instead, it is predominantly influenced by data complexity and feature distribution, particularly on samples where models consistently err. This work provides a novel perspective and empirical foundation for trustworthy evaluation in explainable AI.
This study addresses the security vulnerabilities of the MQTT protocol in Internet of Things (IoT) environments, where the absence of robust encryption and authentication mechanisms renders it susceptible to eavesdropping, message tampering, denial-of-service attacks, and brute-force exploits. For the first time, this work systematically reproduces and evaluates a range of representative MQTT attacks within a high-fidelity smart home simulation environment. By integrating theoretical protocol analysis with empirical validation, the research uncovers critical implementation-level security flaws inherent in widely deployed MQTT systems. Building on these findings, the paper proposes targeted mitigation strategies, practical security hardening measures, and deployment best practices, thereby offering actionable guidance for securing real-world MQTT deployments.
为解决高吞吐量数据流中部署TFM的通信开销和延迟问题,提出HINT框架,结合边缘检索与云端TFM推理,平衡预测性能与通信成本。
This work addresses the gap in machine unlearning research by introducing XGBoost-Forget, the first unlearning method tailored for XGBoost models in tabular network intrusion detection scenarios. Unlike existing approaches primarily designed for deep learning and image data, XGBoost-Forget efficiently removes specified intrusion data points without requiring full model retraining. The method incorporates a customized unlearning mechanism specifically designed for tabular network traffic data and is validated on real-world datasets such as IoT-23 and GeNIS. Experimental results demonstrate that XGBoost-Forget achieves substantial gains in unlearning efficiency while preserving predictive performance nearly equivalent to that of the original model, thereby offering a practical unlearning solution for security applications driven by tabular data.
This study systematically investigates the sources of vulnerability in Retrieval-Augmented Generation (RAG) systems under poisoning attacks. Through a comprehensive full-factorial experiment encompassing 432 configurations, it evaluates the impact of datasets, retriever types (dense, graph-based, and BM25), retrieval depth, knowledge base composition, chunking strategies, and generation models on system robustness. The findings reveal that RAG’s susceptibility arises from complex interactions among retrieval, generation, and knowledge base configurations rather than from any single component’s deficiency. Dense and graph-based retrievers significantly outperform BM25, while increasing retrieval depth or replicating poisoned content across multiple knowledge sources substantially elevates attack success rates. Conversely, incorporating clean data from diverse sources effectively mitigates such attacks. This work is the first to uncover the key factors governing RAG’s robustness against poisoning and their underlying coupling mechanisms.
This study addresses the challenge that local interpretability methods often produce seemingly plausible yet unfaithful explanations for complex tabular data. To rigorously evaluate explanation fidelity, robustness, and complexity, the authors construct a comprehensive benchmarking framework encompassing multiple models and datasets, incorporating for the first time a prediction-consistency grouping strategy. They systematically assess prominent methods—including LIME, Kernel SHAP, and feature ablation—across 32 tabular datasets. The findings reveal that explanation quality shows no significant correlation with model accuracy; instead, it is predominantly influenced by data complexity and feature distribution, particularly on samples where models consistently err. This work provides a novel perspective and empirical foundation for trustworthy evaluation in explainable AI.
This study addresses the security vulnerabilities of the MQTT protocol in Internet of Things (IoT) environments, where the absence of robust encryption and authentication mechanisms renders it susceptible to eavesdropping, message tampering, denial-of-service attacks, and brute-force exploits. For the first time, this work systematically reproduces and evaluates a range of representative MQTT attacks within a high-fidelity smart home simulation environment. By integrating theoretical protocol analysis with empirical validation, the research uncovers critical implementation-level security flaws inherent in widely deployed MQTT systems. Building on these findings, the paper proposes targeted mitigation strategies, practical security hardening measures, and deployment best practices, thereby offering actionable guidance for securing real-world MQTT deployments.