Large Model Empowered Metaverse: State-of-the-Art, Challenges and Opportunities
Metaverse applications face critical bottlenecks including high real-time rendering latency, poor adaptability to dynamic scenes, and limited scalability. To address these challenges, this paper proposes a large language model (LLM)-empowered cloud-edge-device collaborative generative AI rendering framework. It introduces two key innovations: (1) a mobility-aware pre-rendering mechanism that anticipates user movement for proactive resource allocation, and (2) a diffusion model–driven adaptive rendering strategy that dynamically optimizes visual fidelity and computational load based on scene complexity and device capabilities. The framework tightly integrates LLMs, video foundation models (e.g., Sora), and hierarchical distributed computing across cloud, edge, and end devices. Experimental evaluation demonstrates a 37% reduction in end-to-end rendering latency and significantly enhanced real-time immersion under high-concurrency, highly dynamic conditions. This work establishes a scalable, generative-AI-native technical pathway for next-generation metaverse systems.