Faithfulness Is Not Free: Auditing Offline KV-Cache Quantization in Retrieval-Augmented Generation
研究探讨了量化KV缓存对检索增强生成系统忠实性的影响,使用Qwen2.5-7B-Instruct模型在不同量化级别下评估准确性和忠实性。
研究探讨了量化KV缓存对检索增强生成系统忠实性的影响,使用Qwen2.5-7B-Instruct模型在不同量化级别下评估准确性和忠实性。
研究提出LAWA架构,通过紧凑的潜在动作表示未来意图,解决WAM中未来观察生成导致的延迟问题,提高效率和泛化能力。
This work addresses the limitations of existing test-time scaling methods, which often suffer from insufficient sampling diversity or reliance on external verifiers. The authors propose a verifier-free breadth–depth optimization framework that, during inference, iteratively performs self-critique and refinement across multiple reasoning trajectories, followed by majority voting to aggregate final answers. By eliminating dependence on external reward models, the approach effectively balances broad exploration with deep reasoning, significantly enhancing both computational efficiency and accuracy. Empirical results demonstrate substantial improvements over greedy decoding, conventional majority voting, and state-of-the-art verifier-based sampling strategies on mathematical reasoning benchmarks such as MATH500 and AMC; for instance, Qwen2.5-1.5B achieves an accuracy of 32.5% on AMC, up from 25.0%.
This work addresses the vulnerability of existing neural network–based intrusion detection systems to gradient-based adversarial attacks and the limitations of conventional adversarial training in interpretability and defense efficacy. The authors propose LARAR, a novel approach that integrates hierarchical vulnerability analysis, adaptive regularization, and an auxiliary classifier to quantify and optimize layer-wise vulnerability during adversarial training. LARAR introduces, for the first time, an interpretable hierarchical vulnerability score that effectively identifies critical vulnerable layers, thereby reducing computational overhead and enabling early detection of adversarial samples. Experimental results on the UNSW-NB15 dataset demonstrate that LARAR achieves a clean accuracy of 95.01% while significantly enhancing robustness against FGSM, PGD, and transfer attacks.
Existing ransomware detection methods struggle to handle the complexity and variability of evolving ransomware families when relying solely on static, heuristic, or behavioral analysis. This work proposes the first multimodal, multi-agent collaborative detection framework that integrates static, dynamic, and network-based features. Specialized agents extract heterogeneous features, which are then fused by a dedicated fusion agent and fed into a Transformer-based classifier for ransomware family identification. The framework incorporates a confidence-aware abstention mechanism and an adaptive feedback loop to iteratively refine feature representations. Experimental results on large-scale datasets demonstrate a Macro-F1 score of 0.936, significantly reduced calibration error, an agent quality improvement exceeding 0.75, and an overall composite score of approximately 0.88—all achieved without fine-tuning any language model.
研究探讨了量化KV缓存对检索增强生成系统忠实性的影响,使用Qwen2.5-7B-Instruct模型在不同量化级别下评估准确性和忠实性。
研究提出LAWA架构,通过紧凑的潜在动作表示未来意图,解决WAM中未来观察生成导致的延迟问题,提高效率和泛化能力。
This work addresses the limitations of existing test-time scaling methods, which often suffer from insufficient sampling diversity or reliance on external verifiers. The authors propose a verifier-free breadth–depth optimization framework that, during inference, iteratively performs self-critique and refinement across multiple reasoning trajectories, followed by majority voting to aggregate final answers. By eliminating dependence on external reward models, the approach effectively balances broad exploration with deep reasoning, significantly enhancing both computational efficiency and accuracy. Empirical results demonstrate substantial improvements over greedy decoding, conventional majority voting, and state-of-the-art verifier-based sampling strategies on mathematical reasoning benchmarks such as MATH500 and AMC; for instance, Qwen2.5-1.5B achieves an accuracy of 32.5% on AMC, up from 25.0%.
This work addresses the vulnerability of existing neural network–based intrusion detection systems to gradient-based adversarial attacks and the limitations of conventional adversarial training in interpretability and defense efficacy. The authors propose LARAR, a novel approach that integrates hierarchical vulnerability analysis, adaptive regularization, and an auxiliary classifier to quantify and optimize layer-wise vulnerability during adversarial training. LARAR introduces, for the first time, an interpretable hierarchical vulnerability score that effectively identifies critical vulnerable layers, thereby reducing computational overhead and enabling early detection of adversarial samples. Experimental results on the UNSW-NB15 dataset demonstrate that LARAR achieves a clean accuracy of 95.01% while significantly enhancing robustness against FGSM, PGD, and transfer attacks.
Existing ransomware detection methods struggle to handle the complexity and variability of evolving ransomware families when relying solely on static, heuristic, or behavioral analysis. This work proposes the first multimodal, multi-agent collaborative detection framework that integrates static, dynamic, and network-based features. Specialized agents extract heterogeneous features, which are then fused by a dedicated fusion agent and fed into a Transformer-based classifier for ransomware family identification. The framework incorporates a confidence-aware abstention mechanism and an adaptive feedback loop to iteratively refine feature representations. Experimental results on large-scale datasets demonstrate a Macro-F1 score of 0.936, significantly reduced calibration error, an agent quality improvement exceeding 0.75, and an overall composite score of approximately 0.88—all achieved without fine-tuning any language model.