🤖 AI Summary
This study addresses identity bias and accessibility issues in large language model-generated descriptions for security robots. Leveraging 236 identity labels, we employ identity-conditioned prompting and multidimensional text analysis to evaluate generated content. Notably, this work establishes readability as a baseline metric for assessing identity-sensitive outputs, constructing a comprehensive benchmark framework encompassing lexical, semantic, and fairness dimensions. Our findings reveal significant readability disparities across varying prompt conditions and demographic attributes. By providing interpretable evaluation criteria for early-stage robotic development, this research effectively enhances the inclusivity and standardization of design descriptions, offering a systematic approach to mitigating bias in AI-generated technical documentation for security applications.
📝 Abstract
Large language models (LLMs) are increasingly used to generate textual robot design specifications, interaction policies, and risk assessments during early-stage robot development. Such outputs may influence how surveillance and security robots are conceptualized, documented, and ultimately implemented. This paper evaluates whether identity-conditioned prompts produce systematic differences in LLM-generated surveillance and security robot design descriptions. Using 236 demographic identity labels across single-label and model-augmented prompt conditions, we analyze readability as an initial benchmark for evaluating accessibility and identity-conditioned variation in generated robot design descriptions. The results show significant differences in readability across prompt conditions, design dimensions, and demographic identities. Although readability cannot determine whether an output is fair or socially appropriate, it provides an interpretable baseline within a broader benchmarking framework that also includes lexical, semantic, sentiment, syntactic, and fairness-focused analyses.