Language Generation: Complexity Barriers and Implications for Learning
This paper investigates the practical feasibility of language generation in terms of sample efficiency. It establishes that, for classical language classes—including regular and context-free languages—successful generation may require a number of positive examples exceeding the bound of any computable function, rendering sample complexity uncomputable—even though these classes are theoretically learnable. Method: Integrating formal language theory, the PAC learning framework, and the Kleinberg–Mullainathan generative model, the paper rigorously derives and proves strong information-theoretic lower bounds on sample complexity. Contribution/Results: The work provides the first systematic characterization, from a computational complexity perspective, of fundamental sample barriers inherent to language generation. Crucially, it demonstrates that the empirical success of modern large language models cannot be fully explained by classical learnability theory alone; instead, it must rely on structural constraints unique to natural language. This insight offers a novel conceptual bridge between theoretical guarantees and practical performance.