Structured-Prior-Guided Diffusion Inpainting with Physical Consistency for Traffic Sign Augmentation

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
"This study addresses the challenges of long-tail distribution and the scarcity of rare traffic signs by proposing a structured prior-guided diffusion framework for generating physically consistent images. The method injects semantic, appearance, and geometric priors through three orthogonal paths and enforces physical consistency via color and edge structure losses. By leveraging JSON text prompts, front-view vector templates, and affine-aligned vector templates, with Stable Diffusion 1.5 as the backbone, this approach outperforms seven state-of-the-art methods in terms of reconstruction fidelity, physical consistency, and semantic controllability. Notably, it achieves a 91.1% OCR accurate match rate, significantly enhancing the detection performance of rare categories."
📝 Abstract
Traffic sign detection faces a long-tailed data distribution. Many rare signs matter as much as common ones from a regulatory standpoint, yet they have very few samples. Generative data augmentation is one way out. General-purpose inpainting models, however, distort digits, deform geometry and perspective, and shift colours when applied directly to sign regions. We trace this to a single gap: the conditioning signal is too abstract for the physical composition of a sign. We propose a structured-prior-guided diffusion inpainting framework with physical consistency. It injects the semantic, appearance and geometric priors of a sign through three orthogonal pathways: a JSON-formatted text prompt, a front-view vector template rendered with measured dominant colours (via IP-Adapter), and an affine-aligned vector template (via ControlNet). Two physical consistency losses constrain colour with a CIELAB chromaticity $L_1$ term and edge structure with a Sobel gradient term. We train by self-supervised reconstruction on a large set of images collected in-house at AMAP, then evaluate zero-shot on the public TT100K-2021 dataset, a different source. Our method uses a Stable Diffusion 1.5 backbone of about 1.4B parameters. It beats seven representative competitors on every metric of reconstruction fidelity, physical consistency and semantic controllability. Its OCR exact-match rate reaches 91.1\%, against 44.2\% for the 12B industrial model FLUX.1 Fill [dev], and it needs only $1/14$ of that model's inference time. Leave-one-out ablations confirm that each of the three prior pathways and both loss terms contribute on their own. In downstream detection, the synthetic data raises the group-pooled AP50 of rare classes by $1.23\times$ to $7.40\times$ over a real-data-only baseline. Code and pre-trained models are available at https://github.com/52hz-whale/TrafficSignInpaint.
Problem

Research questions and friction points this paper is trying to address.

traffic sign detection
long-tailed data distribution
data augmentation
inpainting
physical consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

structured-prior-guided
diffusion inpainting
physical consistency
traffic sign augmentation
semantic and geometric priors
💼 Related Jobs
No related jobs found.
L
Luo Li
AMAP, Alibaba Group, Beijing, China
C
Chongchong Huang
AMAP, Alibaba Group, Beijing, China
J
Jun Jia
AMAP, Alibaba Group, Beijing, China
Qiang Gao
Qiang Gao
Wuhan University
MoERAGNatural Language Processing
X
Xinlong Liu
AMAP, Alibaba Group, Beijing, China
G
Gui Yang
AMAP, Alibaba Group, Beijing, China
Liang Cao
Liang Cao
Massachusetts Institute of Technology; PhD-University of British Columbia
machine learningfault diagnosisprocess controlprocess monitoringrenewable energy