Evaluating Constrained Iterative Refinement for Scalable Vector Graphics Generation with Off-the-Shelf VLMs

πŸ“… 2026-08-25
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
η ”η©Άι€šθΏ‡ηΊ¦ζŸθΏ­δ»£δΌ˜εŒ–ζ–Ήζ³•οΌŒη»“εˆθ§†θ§‰ει¦ˆγ€η»“ζž„εŒ–ηΌ–θΎ‘ε’ŒηΊ¦ζŸθ§£η ζŠ€ζœ―οΌŒζŽ’η΄’ηŽ°ζœ‰θ§†θ§‰-θ―­θ¨€ζ¨‘εž‹η”Ÿζˆε―ηΌ©ζ”ΎηŸ’ι‡ε›Ύε½’ηš„θƒ½εŠ›δΈŽε±€ι™γ€‚
πŸ“ Abstract
Scalable Vector Graphics (SVGs) power much of the modern visual ecosystem, yet state-of-the-art generative models focus almost entirely on rasterized images. We explore whether inference-time methods can unlock SVG generation capabilities in off-the-shelf vision-language models (VLMs). We systematically evaluate a constrained iterative refinement harness that combines visual feedback, structured editing, and constrained decoding to characterize the capabilities and limitations of current VLMs for SVG generation. Across multiple VLMs and generation settings, we find that constrained decoding improves compilation success rates, while iterative refinement reveals a deficit in visual reasoning and self-correction. Our results highlight both the promise and current limitations of using inference-time methods to adapt general-purpose VLMs for SVG generation.
Problem

Research questions and friction points this paper is trying to address.

Scalable Vector Graphics
vision-language models
inference-time methods
iterative refinement
Innovation

Methods, ideas, or system contributions that make the work stand out.

constrained decoding
iterative refinement
visual feedback
structured editing