An Iterative Feedback Mechanism for Improving Natural Language Class Descriptions in Open-Vocabulary Object Detection

📅 2025-03-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge faced by non-technical users in formulating accurate natural language category descriptions for open-vocabulary object detection. We propose the first iterative human-in-the-loop feedback mechanism specifically designed for refining such textual class descriptions. Our method integrates text embedding analysis with contrastive example embedding synthesis, enabling users to dynamically define novel categories and iteratively improve description quality *in situ*, without retraining the detector. Evaluated across multiple state-of-the-art open-vocabulary detectors—including GLIP and GroundingDINO—the approach consistently improves detection accuracy (mAP gains of +3.2–5.7), while ensuring interpretability of outputs. Our key contributions are: (1) the first integration of human-AI iterative feedback into textual prompt engineering for open-vocabulary detection; (2) a novel description optimization paradigm grounded in contrastive embedding synthesis; and (3) empirical validation of the mechanism’s cross-model generalizability and robustness.

Technology Category

Application Category

📝 Abstract
Recent advances in open-vocabulary object detection models will enable Automatic Target Recognition systems to be sustainable and repurposed by non-technical end-users for a variety of applications or missions. New, and potentially nuanced, classes can be defined with natural language text descriptions in the field, immediately before runtime, without needing to retrain the model. We present an approach for improving non-technical users' natural language text descriptions of their desired targets of interest, using a combination of analysis techniques on the text embeddings, and proper combinations of embeddings for contrastive examples. We quantify the improvement that our feedback mechanism provides by demonstrating performance with multiple publicly-available open-vocabulary object detection models.
Problem

Research questions and friction points this paper is trying to address.

Improving natural language class descriptions for object detection
Enhancing non-technical users' target descriptions with feedback
Optimizing text embeddings for open-vocabulary detection models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Iterative feedback mechanism for class descriptions
Text embedding analysis for target refinement
Contrastive embedding combinations for performance improvement
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Louis Y. Kim
The Charles Stark Draper Laboratory, Inc., 555 Technology Square, Cambridge, MA 02139
M
Michelle Karker
The Charles Stark Draper Laboratory, Inc., 555 Technology Square, Cambridge, MA 02139
V
Victoria Valledor
The Charles Stark Draper Laboratory, Inc., 555 Technology Square, Cambridge, MA 02139
S
Seiyoung C. Lee
The Charles Stark Draper Laboratory, Inc., 555 Technology Square, Cambridge, MA 02139
K
Karl F. Brzoska
The Charles Stark Draper Laboratory, Inc., 555 Technology Square, Cambridge, MA 02139
Margaret Duff
Margaret Duff
The Charles Stark Draper Laboratory, Inc., 555 Technology Square, Cambridge, MA 02139
A
Anthony Palladino
The Charles Stark Draper Laboratory, Inc., 555 Technology Square, Cambridge, MA 02139