Acquiring Grounded Representations of Words with Situated Interactive Instruction
This study addresses the challenge of jointly acquiring perceptual, semantic, and procedural knowledge in embodied lexical learning. We propose a human-robot hybrid active interaction framework enabling a robot to dynamically identify unknown concepts in real-world desktop environments, autonomously initiate contextualized learning requests to human teachers, and unify multimodal knowledge acquisition within a “perception–comprehension–planning–execution” semantic closed loop. The system is built upon the Soar cognitive architecture and integrates visual perception, natural language understanding, task planning, and robotic arm control. Experiments demonstrate significant improvements in rapid generalization to novel lexical items and instruction execution accuracy, along with zero-shot concept transfer capability. Our core contributions are (1) the first hybrid active interaction mechanism for embodied lexical learning, and (2) a unified representational paradigm for heterogeneous knowledge types—perceptual, semantic, and procedural—within a single cognitive architecture.