The Calls are Coming from Inside the Model: Investigating Probe-based Detection of Tool-Calling Errors in LLMs
研究使用线性探针检测大型语言模型中工具调用错误的有效性,通过18个模型测试,发现该方法能有效识别多种错误。
研究使用线性探针检测大型语言模型中工具调用错误的有效性,通过18个模型测试,发现该方法能有效识别多种错误。
This work addresses the growing risk of synthetic speech misuse by proposing a novel task—“audio avatar fingerprinting”—to verify whether synthesized utterances originate from an authorized identity. Building upon existing speaker verification models, the approach enables authentication of synthetic speech against legitimate user profiles. To support this task, we introduce the first forensic dataset comprising paired authentic and corresponding synthetic speech samples. Using this dataset, we demonstrate the feasibility of reliably distinguishing authorized synthetic speech from unauthorized or spoofed instances. Beyond defining a new research direction at the intersection of audio forensics and trustworthy synthetic media, this study provides a foundational benchmark dataset and evaluation framework to advance future work in secure and accountable text-to-speech applications.
Autonomous network defense agents suffer from poor generalization and limited knowledge transferability under dynamic threats, heterogeneous network environments, and multiple conflicting objectives. Method: This paper proposes a unified modeling framework based on Open-ended Reinforcement Learning (OEL). It jointly encodes attack patterns, network topologies, and defense objectives via task embeddings to establish a consistent state-action-reward space; integrates multi-objective reward shaping with a scalable simulation environment to enable continual learning and cross-task policy transfer. Contribution/Results: Experiments demonstrate significant improvements in agent robustness and adaptability to unseen attack types and network topologies. The framework provides a reusable training paradigm and principled guidelines for benchmark design in AI-driven autonomous network defense.
研究使用线性探针检测大型语言模型中工具调用错误的有效性,通过18个模型测试,发现该方法能有效识别多种错误。
This work addresses the growing risk of synthetic speech misuse by proposing a novel task—“audio avatar fingerprinting”—to verify whether synthesized utterances originate from an authorized identity. Building upon existing speaker verification models, the approach enables authentication of synthetic speech against legitimate user profiles. To support this task, we introduce the first forensic dataset comprising paired authentic and corresponding synthetic speech samples. Using this dataset, we demonstrate the feasibility of reliably distinguishing authorized synthetic speech from unauthorized or spoofed instances. Beyond defining a new research direction at the intersection of audio forensics and trustworthy synthetic media, this study provides a foundational benchmark dataset and evaluation framework to advance future work in secure and accountable text-to-speech applications.
Autonomous network defense agents suffer from poor generalization and limited knowledge transferability under dynamic threats, heterogeneous network environments, and multiple conflicting objectives. Method: This paper proposes a unified modeling framework based on Open-ended Reinforcement Learning (OEL). It jointly encodes attack patterns, network topologies, and defense objectives via task embeddings to establish a consistent state-action-reward space; integrates multi-objective reward shaping with a scalable simulation environment to enable continual learning and cross-task policy transfer. Contribution/Results: Experiments demonstrate significant improvements in agent robustness and adaptability to unseen attack types and network topologies. The framework provides a reusable training paradigm and principled guidelines for benchmark design in AI-driven autonomous network defense.