Detect Anything in Graphic Design: Element-Level Rewards for Autoregressive Detection

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决图形设计中元素组成顺序未被利用的问题,提出DAD模型进行有序解构检测,并通过EleRPO方法优化训练,提高了检测性能。
📝 Abstract
Graphic designs, such as posters, advertisements, and infographics, are an important medium for communicating information and shaping understanding. Unlike natural images, they consist of layered elements with explicit compositional order. However, existing object detection models treat these elements as an unordered set, leaving compositional order unexploited. To address this limitation, we present Detect Anything in Graphic Design (DAD), a model that formulates graphic design detection as compositional deconstruction. It decodes elements in compositional order, using lower-layer elements to better detect higher-layer ones. The key feature of DAD is amodal detection, which predicts the full bounding box of each element, including regions occluded by elements placed above it. Building on this formulation, we propose Element Relative Policy Optimization (EleRPO), which extends GRPO from sequence-level supervision to element-level optimization. EleRPO provides fine-grained training signals that capture how each detected element contributes to overall detection quality, and works synergistically with compositional order to improve detection performance. To support training and evaluation, we build a dataset of 10 million graphic designs. Experiments show that DAD outperforms all baselines and achieves human-level performance in amodal detection, supporting effective image-to-layer decomposition. EleRPO consistently improves over GRPO across nine detection benchmarks.
Problem

Research questions and friction points this paper is trying to address.

Graphic Design
Compositional Order
Amodal Detection
Element-Level Rewards
Innovation

Methods, ideas, or system contributions that make the work stand out.

Amodal Detection
Compositional Order
Element Relative Policy Optimization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jiangning Zhu
BNRist, Tsinghua University
B
Bowen Li
BNRist, Tsinghua University
S
Shenyu Qiao
BNRist, Tsinghua University
Y
Yima Gu
BNRist, Tsinghua University
Z
Zhao Zhang
Canva Research
Yuhui Yuan
Yuhui Yuan
Canva CORE, ex-Microsoft Research Asia
Generative AI + DesignComputer Vision
Shixia Liu
Shixia Liu
Professor, Tsinghua University, IEEE Fellow
interactive machine learningData-Centric AIvisual analytics