Benchmarking Identity-Sensitive LLM Outputs for Surveillance and Security Robots

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses identity bias and accessibility issues in large language model-generated descriptions for security robots. Leveraging 236 identity labels, we employ identity-conditioned prompting and multidimensional text analysis to evaluate generated content. Notably, this work establishes readability as a baseline metric for assessing identity-sensitive outputs, constructing a comprehensive benchmark framework encompassing lexical, semantic, and fairness dimensions. Our findings reveal significant readability disparities across varying prompt conditions and demographic attributes. By providing interpretable evaluation criteria for early-stage robotic development, this research effectively enhances the inclusivity and standardization of design descriptions, offering a systematic approach to mitigating bias in AI-generated technical documentation for security applications.
📝 Abstract
Large language models (LLMs) are increasingly used to generate textual robot design specifications, interaction policies, and risk assessments during early-stage robot development. Such outputs may influence how surveillance and security robots are conceptualized, documented, and ultimately implemented. This paper evaluates whether identity-conditioned prompts produce systematic differences in LLM-generated surveillance and security robot design descriptions. Using 236 demographic identity labels across single-label and model-augmented prompt conditions, we analyze readability as an initial benchmark for evaluating accessibility and identity-conditioned variation in generated robot design descriptions. The results show significant differences in readability across prompt conditions, design dimensions, and demographic identities. Although readability cannot determine whether an output is fair or socially appropriate, it provides an interpretable baseline within a broader benchmarking framework that also includes lexical, semantic, sentiment, syntactic, and fairness-focused analyses.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Surveillance Robots
Identity-Conditioned Prompts
Readability Benchmarking
Design Specifications
Innovation

Methods, ideas, or system contributions that make the work stand out.

Identity-Sensitive Benchmarking
Surveillance and Security Robots
Readability Analysis
Demographic Identity Labels
LLM Output Evaluation
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
N
Nneka Hyman
Hunter College, City University of New York
J
Jasmine Khan
Hunter College, City University of New York
Raj Korpan
Raj Korpan
Hunter College, City University of New York
Artificial IntelligenceHuman-Robot InteractionCognitive Science