Agentic Security: A Systematization of Tools, Failure Modes, and Design Laws for LLM-Driven Penetration Testing
本文通过评估十种常用安全工具,系统化分析了基于大语言模型的代理在渗透测试中的失败模式,并提出了解决策略。
本文通过评估十种常用安全工具,系统化分析了基于大语言模型的代理在渗透测试中的失败模式,并提出了解决策略。
This study addresses the dual challenges of benign heterogeneity—arising from diverse operating conditions and failure modes—and adversarial heterogeneity caused by malicious poisoning attacks in federated learning for aircraft engine remaining useful life prediction. To enhance both accuracy and security while preserving the privacy of raw sensor data, the work integrates personalized federated learning with robust aggregation mechanisms. It innovatively introduces a physics-informed sensor backdoor attack and presents the first systematic evaluation of the synergy between shared representation personalization and robust aggregators such as Krum. Experimental results demonstrate that shared representation personalization reduces the performance gap between local and centralized models by 70%, while Krum suppresses attack success rates from 94.9% to 2.8%, achieving high predictive accuracy alongside significantly improved robustness.
This work addresses the challenge of inadequately modeling rare yet high-impact extreme events in time series forecasting, particularly in hydrological streamflow prediction characterized by highly skewed distributions. To this end, the authors propose Exformer, a novel framework that explicitly captures dependencies between normal and extreme events within a Transformer architecture. Central to Exformer is an extreme-adaptive attention mechanism comprising three sparse attention components—Local, Stride, and Extreme—which respectively model short-term dynamics, periodic patterns, and extreme-event-related dependencies. Experimental results on four real-world hydrological datasets demonstrate that Exformer significantly outperforms state-of-the-art models in 3-day-ahead forecasting tasks, effectively enhancing predictive accuracy for extreme events in imbalanced time series.
This work addresses the unreliability of existing medical image diagnosis models that often rely on non-causal or clinically irrelevant visual cues. To enhance trustworthiness, the authors propose a systematic framework that integrates explanation-aware loss directly into the end-to-end training objective by incorporating saliency-based interpretability supervision. A custom explanation loss function jointly optimizes diagnostic accuracy and spatial fidelity of model explanations. The study introduces two quantitative metrics—annotation coverage and saliency precision—to evaluate explanation quality and uncover the trade-off between explanation loss strength and model performance. Experiments on a chest X-ray dataset demonstrate that the proposed method achieves diagnostic accuracy comparable to baseline models while significantly improving spatial alignment between model-generated explanations and clinical annotations.
This work addresses the challenges of deploying large language model–based reinforcement learning agents on resource-constrained edge devices, where memory, computational capacity, and energy consumption pose significant bottlenecks. It presents the first systematic integration of 1-bit quantized language models into reinforcement learning by constructing lightweight decision-making agents based on the BitNet b1.58 architecture. The study introduces a novel theoretical perspective framing quantization as structured parameter perturbation and establishes convergence bounds for quantized policy gradients under a frozen backbone setting, revealing a fundamental trade-off between exploration and stability under extreme quantization. Experiments demonstrate that the proposed approach reduces memory usage by 10–16× and improves energy efficiency by 3–5× compared to full-precision baselines, while retaining 85%–98% of task performance across multiple benchmarks and enabling feasible on-device training and inference on commercial edge hardware.
本文通过评估十种常用安全工具,系统化分析了基于大语言模型的代理在渗透测试中的失败模式,并提出了解决策略。
This study addresses the dual challenges of benign heterogeneity—arising from diverse operating conditions and failure modes—and adversarial heterogeneity caused by malicious poisoning attacks in federated learning for aircraft engine remaining useful life prediction. To enhance both accuracy and security while preserving the privacy of raw sensor data, the work integrates personalized federated learning with robust aggregation mechanisms. It innovatively introduces a physics-informed sensor backdoor attack and presents the first systematic evaluation of the synergy between shared representation personalization and robust aggregators such as Krum. Experimental results demonstrate that shared representation personalization reduces the performance gap between local and centralized models by 70%, while Krum suppresses attack success rates from 94.9% to 2.8%, achieving high predictive accuracy alongside significantly improved robustness.
This work addresses the challenge of inadequately modeling rare yet high-impact extreme events in time series forecasting, particularly in hydrological streamflow prediction characterized by highly skewed distributions. To this end, the authors propose Exformer, a novel framework that explicitly captures dependencies between normal and extreme events within a Transformer architecture. Central to Exformer is an extreme-adaptive attention mechanism comprising three sparse attention components—Local, Stride, and Extreme—which respectively model short-term dynamics, periodic patterns, and extreme-event-related dependencies. Experimental results on four real-world hydrological datasets demonstrate that Exformer significantly outperforms state-of-the-art models in 3-day-ahead forecasting tasks, effectively enhancing predictive accuracy for extreme events in imbalanced time series.
This work addresses the unreliability of existing medical image diagnosis models that often rely on non-causal or clinically irrelevant visual cues. To enhance trustworthiness, the authors propose a systematic framework that integrates explanation-aware loss directly into the end-to-end training objective by incorporating saliency-based interpretability supervision. A custom explanation loss function jointly optimizes diagnostic accuracy and spatial fidelity of model explanations. The study introduces two quantitative metrics—annotation coverage and saliency precision—to evaluate explanation quality and uncover the trade-off between explanation loss strength and model performance. Experiments on a chest X-ray dataset demonstrate that the proposed method achieves diagnostic accuracy comparable to baseline models while significantly improving spatial alignment between model-generated explanations and clinical annotations.
This work addresses the challenges of deploying large language model–based reinforcement learning agents on resource-constrained edge devices, where memory, computational capacity, and energy consumption pose significant bottlenecks. It presents the first systematic integration of 1-bit quantized language models into reinforcement learning by constructing lightweight decision-making agents based on the BitNet b1.58 architecture. The study introduces a novel theoretical perspective framing quantization as structured parameter perturbation and establishes convergence bounds for quantized policy gradients under a frozen backbone setting, revealing a fundamental trade-off between exploration and stability under extreme quantization. Experiments demonstrate that the proposed approach reduces memory usage by 10–16× and improves energy efficiency by 3–5× compared to full-precision baselines, while retaining 85%–98% of task performance across multiple benchmarks and enabling feasible on-device training and inference on commercial edge hardware.