Policy Compliance of User Requests in Natural Language for AI Systems

📅 2026-02-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of assessing compliance with organizational policies in natural language user requests by introducing the first annotated dataset tailored to industrial settings, along with a corresponding evaluation benchmark that encompasses diverse security and regulatory constraints prevalent in the technology sector. Leveraging large language models (LLMs) integrated with various prompting and reasoning strategies, the study systematically evaluates the performance of current approaches on this task. The findings reveal significant limitations of existing LLMs in accurately interpreting complex policy provisions and making reliable compliance judgments, thereby underscoring the inherent difficulty of the problem. This research establishes a reproducible benchmark and provides a foundational dataset to support future advancements in policy-aware language understanding.

Technology Category

Application Category

📝 Abstract
Consider an organization whose users send requests in natural language to an AI system that fulfills them by carrying out specific tasks. In this paper, we consider the problem of ensuring such user requests comply with a list of diverse policies determined by the organization with the purpose of guaranteeing the safe and reliable use of the AI system. We propose, to the best of our knowledge, the first benchmark consisting of annotated user requests of diverse compliance with respect to a list of policies. Our benchmark is related to industrial applications in the technology sector. We then use our benchmark to evaluate the performance of various LLM models on policy compliance assessment under different solution methods. We analyze the differences on performance metrics across the models and solution methods, showcasing the challenging nature of our problem.
Problem

Research questions and friction points this paper is trying to address.

Policy Compliance
Natural Language Requests
AI Systems
Safety
Reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

policy compliance
natural language requests
benchmark dataset
large language models
AI safety
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Pedro Cisneros-Velarde
VMware Research