🤖 AI Summary
This work addresses a critical gap in representation learning: while much research focuses on refining existing representations, little attention has been paid to the mechanisms that trigger the emergence of representations at new levels of abstraction. The paper proposes that when current representations fail to account for observed organizational structures or dynamic patterns, the system actively initiates the generation of novel representations, thereby enabling a recursive, self-bootstrapping process of representational evolution. For the first time, the authors formalize “explanatory insufficiency” as a positive driving force behind this evolution and introduce a five-stage framework grounded in cognitive and systems theory to model representational emergence. This framework applies broadly to systems such as world models and foundation models, offering a new paradigm for AI design—one that endows systems with the capacity to recognize the limits of their own representations, thereby enabling more autonomous representation learning and scientific discovery.
📝 Abstract
Representation learning is central to modern machine learning, enabling transitions from handcrafted features to learned embeddings, latent spaces, foundation models, world models, and digital twins. Yet most research examines how representations are optimized after a representational framework has been selected, while less attention is given to when a new level of representation becomes necessary. We introduce the Bootstrap Theory of Representational Emergence (TBER), a framework describing how new representations arise when existing ones become explanatorily insufficient. In this view, representational innovation is not only driven by more data, larger models, or greater computational power, but also by persistent explanatory gaps: situations in which a representation can still describe observations but can no longer make their organization or transformations intelligible. TBER identifies explanatory insufficiency as a positive signal for representational transition. A representation becomes insufficient not because it is necessarily false, but because its explanatory domain has been exceeded. The bootstrap dynamic follows a recursive sequence: observations reveal anomalies; anomalies expose insufficiencies; insufficiencies motivate new representations; and these new representations generate further observations and possible new insufficiencies.We formalize this process through five stages: stabilized observation, anomaly detection, recognition of explanatory insufficiency, representational emergence, and provisional stabilization. We discuss applications to representation learning, latent spaces, foundation models, world models, digital twins, adaptive biological systems, and scientific discovery. TBER suggests that future AI systems may benefit from mechanisms for detecting the explanatory limits of their own internal representations.