🤖 AI Summary
To address high interaction latency, modality fragmentation, and weak immersion in wearable mixed reality (MR) environments, this paper proposes a lightweight multimodal intelligent agent framework. Methodologically, it integrates spatial mapping, automatic speech recognition (ASR), gaze estimation, object detection, and knowledge-graph-driven dialogue, underpinned by a cloud-edge collaborative computing architecture for efficient computational offloading; it further introduces novel mechanisms for automatic speech–animation synchronization and human-like gaze modeling. The key contributions are: (1) the first realization of low-latency (2–4 seconds), high-naturalness virtual–physical interaction on resource-constrained edge devices; (2) a modular, cross-device-compatible framework supporting all SLAM-capable MR headsets. Evaluation in real-world museum and botanical garden deployments demonstrates significant improvements in user engagement and content retention rates.
📝 Abstract
In this paper, we present the design of a multimodal interaction framework for intelligent virtual agents in wearable mixed reality environments, especially for interactive applications at museums, botanical gardens, and similar places. These places need engaging and no-repetitive digital content delivery to maximize user involvement. An intelligent virtual agent is a promising mode for both purposes. Premises of framework is wearable mixed reality provided by MR devices supporting spatial mapping. We envisioned a seamless interaction framework by integrating potential features of spatial mapping, virtual character animations, speech recognition, gazing, domain-specific chatbot and object recognition to enhance virtual experiences and communication between users and virtual agents. By applying a modular approach and deploying computationally intensive modules on cloud-platform, we achieved a seamless virtual experience in a device with limited resources. Human-like gaze and speech interaction with a virtual agent made it more interactive. Automated mapping of body animations with the content of a speech made it more engaging. In our tests, the virtual agents responded within 2-4 seconds after the user query. The strength of the framework is flexibility and adaptability. It can be adapted to any wearable MR device supporting spatial mapping.