Conformity and Social Impact on AI Agents
This study investigates conformity behavior in large language models (LLMs) acting as AI agents within multi-agent environments and the associated safety risks arising from social influence. By replicating a classic social psychology visual experiment and integrating multimodal LLMs with key group influence variables—such as group size, unanimity, and task difficulty—the work provides the first systematic validation that AI agents exhibit conformity tendencies consistent with established social influence theory. The findings reveal that even high-performing models, which demonstrate strong accuracy in isolation, are significantly swayed by group opinions, with their susceptibility intensifying as task complexity increases. This highlights critical vulnerabilities in multi-agent systems, particularly the potential for social manipulation and the propagation of biases through collective dynamics.