PointRL: Learning Point-Level Vision-Language Grounding from Verifiable Annotation Evidence

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出PointRL框架,通过强化学习和异构标注证据学习点级视觉-语言对齐,解决多实例指令中的目标覆盖、计数一致性和重复抑制问题。
📝 Abstract
Vision-language models (VLMs) increasingly rely on point coordinates as a compact and executable interface for visual grounding in GUI interaction, robotic manipulation, and interactive visual systems. However, learning reliable pointing behavior remains difficult because the supervision space is inherently non-unique: many coordinates may be valid within the same target region, while multi-instance instructions require target coverage, count consistency, and duplicate suppression. This work presents PointRL, a verifiable reinforcement learning framework that learns point-level grounding from existing heterogeneous annotation evidence. PointRL converts bounding boxes, masks, and instance labels into pointing instructions, while retaining their target supports, instance membership, and set constraints as hidden verifier evidence, i.e., annotations kept outside the prompt and used by a deterministic checker to score predictions. The proposed reward evaluates parseability, point validity, instance coverage, cardinality consistency, and redundant or missing predictions. On PointArena, PointRL improves the overall accuracy of Qwen3.5-4B from 56.11% to 65.58%. Further evaluations on RoboSpatial, BLINK, and Ref-Adv show same-backbone gains on the evaluated external benchmarks, suggesting that verifiable point-level feedback may benefit spatial grounding in these settings.
Problem

Research questions and friction points this paper is trying to address.

point-level grounding
non-unique supervision space
multi-instance instructions
target coverage
count consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

verifiable reinforcement learning
point-level grounding
annotation evidence
deterministic checker
reward mechanism
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
J
Jingyang Su
School of Intelligent Engineering and Automation, Beijing University of Posts and Telecommunications, Beijing, China
Pu Cao
Pu Cao
Beijing University of Posts and Telecommunications
Computer Vision
X
Xiuze Jin
School of Intelligent Engineering and Automation, Beijing University of Posts and Telecommunications, Beijing, China
L
Longyue Zhang
School of Intelligent Engineering and Automation, Beijing University of Posts and Telecommunications, Beijing, China
Qing Song
Qing Song
Apple Inc
Image processingvideo processingvideo compression
Lu Yang
Lu Yang
Beijing University of Posts and Telecommunications
computer visionmachine learning