Temporal Preference Concepts and their Functions in a Large Language Model
How large language models internally represent and balance short-term rewards against long-term consequences remains unclear. This study addresses this gap by applying mechanistic interpretability techniques to identify the causal neural subgraph underlying time preference in Qwen3-4B-Instruct-2507, revealing its geometric encoding structure within the residual stream. Through an integrated approach combining gradient attribution, activation patching, and steering vector interventions, we demonstrate that the model exhibits a significantly lower discount rate for future rewards than humans and displays context-dependent instability in temporal preferences. Furthermore, we show that time preference can be effectively modulated via targeted interventions at mid-to-upper-layer network nodes. Our findings provide both empirical grounding and a novel pathway for controllably shaping the planning capabilities of large language models.