π€ AI Summary
This study introduces, for the first time, the psychological conflict-task paradigm into the domain of large language models (LLMs) to investigate the underlying mechanisms driving congruency effects in verbal conflict tasks. By designing purely textual congruent and incongruent prompt conditions, the work reveals the competitive dynamics between a modelβs default color associations and explicitly stated rules. Through causal attribution, attention analysis, ablation studies, and fine-tuning interventions, the research replicates robust congruency effects across Gemma-2-2B and multiple Pythia models. It further demonstrates that short-range and long-range attention mechanisms preferentially govern processing under congruent and incongruent conditions, respectively, thereby validating a competition between default weight-embedded mappings and contextually specified rule-based mappings.
π Abstract
Congruency effects, observed in conflict tasks such as Stroop and flanker tasks, have been investigated for nearly a century in psychology and neuroscience, but their mechanistic basis is not fully understood. We introduce a verbal-only LLM conflict task in which a prompt stem elicits a default same-color completion and an explicit rule either agrees with (congruent condition) or conflicts with (incongruent condition) the completion. Gemma-2-2B and six Pythia models ranging from 410M to 12B parameters showed strong default same-color tendencies, and six of seven models showed strong congruency effects. Using causal attribution analysis, attention analysis, and attention ablations, we identified distinct processing pathways in these LLMs: a pathway involving short-range attention to a superficial color cue that is preferentially activated in the congruent condition, and a pathway involving long-range attention to the rule prefix that is preferentially activated in the incongruent condition. Fine-tuning that strengthened the default same-color tendency had divergent effects on task conditions, reducing incongruent performance while increasing congruent performance. In contrast, increasing rule set size selectively impaired incongruent performance. These converging findings support an account in which congruency effects in this task arise from competition between an in-weight default mapping and an in-context rule-based mapping. More broadly, our findings illustrate how LLMs can serve as model systems for mechanistic analysis of competition between default and rule-governed response tendencies within a single learned network.