PolicyMem: Geometric Policy Memory for LLM Governance

📅 2026-09-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决LLM治理问题,提出PolicyMem方法,将自然语言策略外部化为几何记忆对象,实现检测、干预和验证的一致性复用。
📝 Abstract
As large language models (LLMs) are increasingly deployed in real-world high-stakes applications, effective governance has become essential. Existing safeguards largely follow two paradigms: learning-based guards provide strong semantic discrimination but couple policy behavior to trained models and taxonomies, while programmable frameworks offer flexible control but require substantial manual prompt and workflow engineering. Neither externalizes policies as reusable operational states, making it difficult to consistently reuse policy evidence across detection, intervention, and verification. In this paper, we introduce PolicyMem, a geometric policy memory that externalizes natural-language policies as reusable geometric memory objects represented by low-rank subspaces in a shared representation space. A memory writer compiles natural-language policies into policy memory slots, and query-response pairs read the policy memory through projection energy. The resulting policy-evidence profile directly mediates the safety verdict and is reused for policy attribution and post-intervention verification. Coupled with a response rewriter, PolicyMem enables a detect-rewrite-verify loop for LLM governance. Across five widely used benchmarks, PolicyMem achieves state-of-the-art unsafe behavior detection while enabling effective policy attribution, rewriting, and post-intervention verification through the shared policy memory.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Governance
Policy Externalization
Reusable Operational States
Policy Evidence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Geometric Policy Memory
Natural-Language Policies
Reusable Geometric Memory Objects
Shared Representation Space
Detect-Rewrite-Verify Loop
🔎 Similar Papers
No similar papers found.