Design of Seamless Multi-modal Interaction Framework for Intelligent Virtual Agents in Wearable Mixed Reality Environment
To address high interaction latency, modality fragmentation, and weak immersion in wearable mixed reality (MR) environments, this paper proposes a lightweight multimodal intelligent agent framework. Methodologically, it integrates spatial mapping, automatic speech recognition (ASR), gaze estimation, object detection, and knowledge-graph-driven dialogue, underpinned by a cloud-edge collaborative computing architecture for efficient computational offloading; it further introduces novel mechanisms for automatic speech–animation synchronization and human-like gaze modeling. The key contributions are: (1) the first realization of low-latency (2–4 seconds), high-naturalness virtual–physical interaction on resource-constrained edge devices; (2) a modular, cross-device-compatible framework supporting all SLAM-capable MR headsets. Evaluation in real-world museum and botanical garden deployments demonstrates significant improvements in user engagement and content retention rates.