Role Steering of Language Models for Social Simulations
This study addresses the challenge of maintaining behavioral consistency in language model agents during social simulations and the absence of effective pre-deployment screening mechanisms. The authors propose a role-screening pipeline based on activation steering, which enables fine-grained control by defining role profiles, extracting role-specific activation directions, scanning steering coefficients, and evaluating behavioral consistency. For the first time, experiments across 275 roles demonstrate that steering intensity should be individually optimized per role rather than uniformly applied. Using the OLMo-3-7B-Instruct model with GPT-4.1-mini for auxiliary evaluation, the proposed method significantly outperforms baselines in role consistency (63.2 vs. 41.1) while preserving high lexical diversity. Notably, performance declines with increased steering strength for 38 roles, highlighting the necessity of role-adaptive calibration.