Emergent Introspective Awareness in Large Language Models
This study investigates whether large language models possess genuine introspective capabilities rather than merely generating superficially plausible but fabricated responses. By injecting known conceptual representations into the model’s internal activations and combining self-report analyses with instruction-guided activation modulation, the work presents the first systematic intervention to probe and validate a model’s awareness of its own internal states. The findings reveal that Claude Opus 4 and 4.1 can, under specific conditions, accurately identify injected content, distinguish between self-generated and externally prefilled information, and modulate their internal representations according to instructions—suggesting a measurable degree of introspective awareness. This research establishes a novel methodological framework and provides empirical evidence for evaluating self-awareness in language models.