๐ค AI Summary
Computer vision has long been constrained by the closed-world assumption, leading to highly fragmented tasks. Method: This paper introduces โOpen-World Detectionโ (OWD) as the first unified paradigm for general-purpose visual detection in open environments, systematically integrating saliency detection, out-of-distribution detection, zero-shot detection, open-world object detection, and vision-language models. It analyzes task evolution, constructs a coherent knowledge framework, and synthesizes key datasets and methodologies. Contribution/Results: The work reveals an emerging trend toward class-agnostic, scene-adaptive detection fusion. By transcending traditional closed-world constraints, the OWD framework establishes a theoretical foundation and technical roadmap for general visual perception, enabling detection models to robustly generalize beyond narrow, predefined task boundaries into dynamic, open-ended environments.
๐ Abstract
For decades, Computer Vision has aimed at enabling machines to perceive the external world. Initial limitations led to the development of highly specialized niches. As success in each task accrued and research progressed, increasingly complex perception tasks emerged. This survey charts the convergence of these tasks and, in doing so, introduces Open World Detection (OWD), an umbrella term we propose to unify class-agnostic and generally applicable detection models in the vision domain. We start from the history of foundational vision subdomains and cover key concepts, methodologies and datasets making up today's state-of-the-art landscape. This traverses topics starting from early saliency detection, foreground/background separation, out of distribution detection and leading up to open world object detection, zero-shot detection and Vision Large Language Models (VLLMs). We explore the overlap between these subdomains, their increasing convergence, and their potential to unify into a singular domain in the future, perception.