Enforcing LLM Safety through DMD-based Classification of Prompt-Response Embedding Dynamics

📅 2026-08-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过动态系统框架和Koopman算子对提示-响应嵌入动态进行分类,以有效检测大型语言模型生成的有害内容,提高模型安全性。
📝 Abstract
Large Language Models (LLMs) are increasingly deployed in high-stakes applications, yet their tendency to generate toxic, harmful, or policy-violating content poses significant risks. Detecting these unsafe outputs efficiently in a black-box manner remains an open challenge. In this paper, we extend a recently proposed dynamical systems framework designed for hallucination detection to LLM safety classification. By projecting both prompts and responses into high-dimensional embedding spaces and fitting separate Koopman-based predictive models for safe and unsafe regimes, we classify new outputs using a new differential residual score that compares prediction errors of the safe and unsafe regimes. A key contribution is the incorporation of the prompt and response embedding dynamics, yielding fitted Koopman operators that capture crucial interaction patterns. We evaluate our black-box method across three safety benchmarks using three embedding models. Our results show that incorporating prompt embeddings yields consistent improvements, particularly for interaction-dependent violations when paired with causal decoders (e.g., in Llama-3), while response-only violations benefit more from dense semantic embedding representations. These findings opens the door for using dynamical systems to analyze AI systems rather than the dominant paradigm of using AI to model dynamical systems.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
unsafe outputs
black-box detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Koopman-based predictive models
differential residual score
embedding dynamics
interaction patterns
causal decoders
Mohamed Akrout
Mohamed Akrout
Assistant Professor of Electrical Engineering and Computer Science, University of Tennessee
signal processingAI for dermatologywireless communicationcircuit theorydigital health
O
Olivera Kotevska
Computer Science and Mathematics Division, Oak Ridge National Laboratory, Oak Ridge, TN 37830, USA
D
Dan Wilson
Department of Electrical Engineering and Computer Science, University of Tennessee, Knoxville, TN 37996, USA