Rules or Character? Scaling Laws for AI Safety Design
This study investigates how AI systems should dynamically balance rule-based safety mechanisms against behavior-shaping approaches as deployment scale increases, aiming to minimize both expected harm and tail risk. We formalize this trade-off for the first time, introducing a parameter α to represent the allocation of resources between the two strategies. Incorporating factors such as filter degradation, common-mode failures, and baseline behavioral vulnerability, we evaluate the trade-off using comparative statics, a multiplicative Pareto damage model, Monte Carlo simulations, and Conditional Value-at-Risk (CVaR). Our results show that the baseline vulnerability of behavior shaping is the dominant determinant of the optimal strategy. As deployment scale grows, the optimal α* shifts modestly to substantially toward behavior shaping (Δα* = +0.01 to +0.21), reaching offsets up to 0.50 under high vulnerability; moreover, at large scales, the CVaR-optimal and expected-harm-optimal solutions converge.