UAVs Meet Embodied Intelligence: Bridging Human Intents and Flying Dynamics Via Harnessing Physical-Digital AI Agents
论文探讨了通过融合物理-数字AI代理来解决无人机理解人类意图和飞行动态的问题,提出了5+5框架以实现适应性和持续进化的无人机自主性。
论文探讨了通过融合物理-数字AI代理来解决无人机理解人类意图和飞行动态的问题,提出了5+5框架以实现适应性和持续进化的无人机自主性。
Existing world models struggle to balance inference efficiency and deployment cost, making them ill-suited for high-frame-rate real-time interaction. This work proposes MoWorld, an end-to-end efficient world model that integrates 3D-native data generation, curriculum-based cross-frame pretraining, diffusion denoising step distillation, and a mixed-precision parallel inference framework. Notably, it achieves real-time interaction at 50 FPS on neural processing units (NPUs) for the first time. Departing from conventional paradigms reliant on large-scale video corpora, MoWorld maintains cinematic visual quality while reducing inference costs to 30%–50% of current models, substantially enhancing practical deployability and environmental adaptability.
论文探讨了通过融合物理-数字AI代理来解决无人机理解人类意图和飞行动态的问题,提出了5+5框架以实现适应性和持续进化的无人机自主性。
Existing world models struggle to balance inference efficiency and deployment cost, making them ill-suited for high-frame-rate real-time interaction. This work proposes MoWorld, an end-to-end efficient world model that integrates 3D-native data generation, curriculum-based cross-frame pretraining, diffusion denoising step distillation, and a mixed-precision parallel inference framework. Notably, it achieves real-time interaction at 50 FPS on neural processing units (NPUs) for the first time. Departing from conventional paradigms reliant on large-scale video corpora, MoWorld maintains cinematic visual quality while reducing inference costs to 30%–50% of current models, substantially enhancing practical deployability and environmental adaptability.