🤖 AI Summary
This study addresses the challenge where lexical diversity confounds intrinsic dimensionality estimation in language models, obscuring the distinction between essential properties and data artifacts. We reveal a scale-dependent dual-regime effect governing how intrinsic dimensionality varies with lexical diversity. Through geometric manifold analysis and non-parametric theoretical modeling, we derive an exact inversion formula to eliminate estimation bias. Experimental validation demonstrates that this formula aligns perfectly with empirical observations across scales, effectively decoupling data construction factors from the geometric structure of linguistic representations. Consequently, this work establishes a novel paradigm for understanding the manifold organization principles and representational complexity of large language models by providing a rigorous method to isolate true geometric dimensionality from vocabulary-induced distortions.
📝 Abstract
Intrinsic dimensionality (ID) is widely used to probe the representational complexity of language models, but it remains unclear whether ID differences reflect properties of language itself or artefacts of how the underlying dataset was constructed. In this paper, we focus specifically on how lexical diversity, the number of unique last-token items present in a dataset, affects ID estimates of that dataset. We find a scale-dependent transition between two regimes: at low lexical diversity, conditions with fewer unique final words produce higher ID, while at high lexical diversity, this ordering reverses, and conditions with more unique words produce higher ID. We derive an exact, parameter-free formula for the point at which this reversal occurs, which matches the observed transition point at every scale tested. On the one hand, our results highlight how care must be taken when interpreting the intrinsic dimensionality of a set of representations as a straightforward cue of their complexity. On the other hand, our discovery of the two ID regimes reveals a general principle of organisation of linguistic data in LLMs that sheds new light on their inner manifold structures.