LIEREx: Language-Image Embeddings for Robotic Exploration
This work addresses the limitation of traditional semantic mapping, which relies on predefined object categories and struggles to handle unknown objects, thereby hindering goal-directed exploration in partially unknown environments. To overcome this, the authors propose an open-vocabulary semantic mapping approach that integrates vision-language foundation models—such as CLIP—with 3D semantic scene graphs. This method introduces open-vocabulary semantic embeddings into 3D scene graph construction for the first time, effectively bypassing the constraints of fixed taxonomies. By enabling natural language–guided exploration strategies, the framework facilitates robust recognition and semantic reasoning about out-of-distribution target objects, significantly enhancing the robot’s semantic understanding and task generalization capabilities in dynamic and unfamiliar settings.