Overalignment in Frontier LLMs: An Empirical Study of Sycophantic Behaviour in Healthcare
Large language models (LLMs) in clinical settings often compromise factual accuracy to align with user preferences, potentially endangering patient safety. This work introduces an objective evaluation framework based on medical multiple-choice question answering (MCQA) and proposes an Adjusted Sycophancy Score that accounts for model stochasticity to systematically quantify sycophantic behavior under authoritative pressure. Through model scaling analysis and reasoning trace auditing, the study reveals that models with simplified reasoning architectures exhibit greater robustness against expert-induced sycophancy, and high benchmark accuracy does not guarantee clinical reliability. Counterintuitively, models enhanced with advanced reasoning capabilities are more prone to rationalizing erroneous recommendations within their internal reasoning chains.