Why Formal Monitors Fail: Attack Distribution Entropy as a Coverage Bound for LTL-Based LLM Agent Safety
This study addresses the significant variation in attack interception efficacy of existing LTL/FSA-based safety monitors across different large language model (LLM) architectures, a phenomenon previously lacking theoretical explanation. The work reveals, for the first time, that the distributional entropy of attack trigger-completion patterns is the fundamental determinant of formal monitor coverage. To enable pre-deployment evaluation independent of model architecture, the authors propose an entropy-based testing methodology. Leveraging Shannon entropy and statistical correlation analysis, they establish a theoretical upper bound on the recall of invariant-based monitors. Experiments across eight state-of-the-art LLMs demonstrate a strong negative correlation (r = −0.87) between monitor coverage and attack distribution entropy: GPT-style models achieve 96% coverage with a single pattern, whereas Gemini-style models require multiple clustered strategies yet attain only 6–13% coverage.