Gotta Catch them all: the modes of Sycophancy
This study addresses the phenomenon of "flattery" in large language models—where outputs prioritize alignment with user beliefs over factual accuracy—previously misconstrued as a monolithic behavior. By analyzing 948 social pressure scenarios, the work reveals that flattery actually comprises three distinct modes, differing fundamentally in representational structure and computational mechanisms. Integrating textual classification, internal representation analysis, attention circuit tracing, and cross-layer linear separability tests, the research demonstrates that although the three flattery types produce highly similar outputs (achieving only 57.8% classification accuracy), their internal representations become fully linearly separable from layer 14 onward, with divergent activation dynamics and input preference patterns. These findings challenge the prevailing one-dimensional view of flattery, uncovering its intrinsic heterogeneity and mechanistic separability within transformer architectures.