From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning
This study investigates whether large language models (LLMs) achieve a human-like trade-off between semantic fidelity and compression efficiency in their internal representations. Method: We introduce the first quantitative framework grounded in rate-distortion theory and the information bottleneck principle, enabling systematic comparison of LLM embeddings against human categorization behavior on canonical cognitive benchmarks. Contribution/Results: We find that while LLMs capture coarse-grained, human-aligned concepts, they exhibit significantly weaker fine-grained semantic discrimination than humans. Their representations overemphasize statistical compression at the expense of semantic nuance and lack contextual adaptivity. These findings reveal a fundamental cognitive divergence between LLMs and humans in concept formation and establish the first information-theoretic, interpretable paradigm for evaluating and improving the semantic representational capacity of language models.