🤖 AI Summary
This work addresses the challenge that existing memory localization methods cannot reliably determine whether an intervention direction operates through a selective mechanism, thereby limiting risk-controlled model editing. It reframes memory localization as a falsifiable problem of predicting selective interventions and introduces a predictive framework under low-dose causal responses. By modeling grid-based intervention paths, decoupling magnitude and direction, matching residual norms, and integrating static localization with supervised geometric constraints, the approach disentangles target movement, semantic interference, and capability degradation. This enables risk-aware intervention decisions without exhaustive scanning. Evaluated across nine datasets and fourteen domains, the method achieves a 13.1% hit rate for RFM/AGOP directions at layer 7—significantly outperforming random baselines—and attains macro AUROC scores of 0.801–0.828 on held-out records, effectively enhancing editing efficacy while mitigating damage to semantically proximate knowledge.
📝 Abstract
Activation steering turns localized representations into control directions, but localization alone does not reveal whether a direction has a selective operating regime. We introduce Predictive Memory Localization (PML), which treats the measured-grid intervention path as the predictive object of memory localization. PML separates random-calibrated target movement from semantic-neighbor and capability damage, and compares static localization and supervised geometry with a strength-disjoint low-dose causal response. Our frozen study covers 3,000 records from nine datasets and fourteen domains, yielding 30,000 distinct record-direction-layer paths and 210,000 distinct path-strength evaluations. At layer 7, the geometry-derived RFM/AGOP direction reaches 13.1% target-any and 12.3% clean-any, exceeding random by 3.6 and 3.4 percentage points under a record-paired bootstrap. Across record-, dataset-, and domain-grouped splits, responses at $|α|=0.1$ are the strongest signal for outcomes at disjoint strengths $|α|\in\{0.25,0.5\}$. On held-out records, a predictor-driven selector chooses a coefficient or abstains, improves utility and reduces semantic-neighbor damage relative to a train-tuned fixed-strength policy, and avoids most evaluations in a dense scan. Across three residual-norm-matched base models, learned directions retain selective-path gains and low-dose responses yield 0.801-0.828 record-held-out macro AUROC. PML therefore turns memory localization into a falsifiable forecast of margin-level selective outcomes and a risk-aware intervention decision.