VectorGym: A Multitask Benchmark for SVG Code Generation, Sketching, and Editing

📅 2026-02-22
🏛️ arXiv.org
📈 Citations: 3
Influential: 0
📄 PDF
🤖 AI Summary
为解决SVG生成、编辑等缺乏专业级基准问题,提出VectorGym,采用多任务强化学习方法优化,并提供人类标注数据和评估指标。
📝 Abstract
We introduce VectorGym, a comprehensive benchmark suite for Scalable Vector Graphics (SVG) that spans generation from text and sketches, complex editing, and visual understanding. VectorGym addresses the lack of realistic, challenging benchmarks aligned with professional design workflows. Our benchmark comprises four tasks with expert human-authored annotations: the novel Sketch2SVG task (VG-Sketch); a new SVG editing dataset (VG-Edit) featuring complex, multi-step edits with higher-order primitives; Text2SVG generation (VG-Text); and SVG captioning (VG-Cap). Unlike prior benchmarks that rely on synthetic edits, VectorGym provides gold-standard human annotations that require semantic understanding and design intent. We also propose a multi-task reinforcement learning approach that jointly optimizes across all four tasks using rendering-based rewards. Our method, built on GRPO with curriculum learning, trains a Qwen3-VL 8B model that achieves state-of-the-art performance among open-source models, surpassing much larger models including Qwen3-VL 235B and matching GPT-4o. We also introduce a VLM-as-a-Judge metric for SVG generation, validated through human correlation studies. Our evaluation of frontier VLMs reveals significant performance gaps, positioning VectorGym as a rigorous framework for advancing visual code generation. VectorGym is publicly available on huggingface.co/datasets/ServiceNow/VectorGym.
Problem

Research questions and friction points this paper is trying to address.

VectorGym
SVG
benchmark
design workflow
generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

VectorGym
multi-task benchmark
human annotations
reinforcement learning
VLM-as-a-Judge
💼 Related Jobs
No related jobs found.
J
Joan Rodriguez
Mila, Quebec AI Institute; ÉTS Montréal; QuiverAI
Haotian Zhang
Haotian Zhang
Columbia University
Computer VisionNatural Language ProcessingMachine Learning
Abhay Puri
Abhay Puri
Applied Research Scientist, ServiceNow Research
Agent SecurityLarge Language ModelsComputer VisionMultiModal Foundational Models
H
Haoran Dai
Illinois Institute of Technology; QuiverAI
T
Tianyang Zhang
University of Massachusetts Amherst; QuiverAI
Meng Lin
Meng Lin
Southern University of Science and Technology
Solar energy conversion and utilization
Rishav Pramanik
Rishav Pramanik
PhD Student, Stony Brook University, New York, USA
Artificial IntelligenceMachine LearningComputer VisionMultimodal Learning
X
Xiaoqing Xie
Minzu University of China; QuiverAI
M
Marco Terral Rodriguez
Computer Vision Center; QuiverAI
D
Darsh Kaushik
Mila, Quebec AI Institute; QuiverAI
A
Aly Shariff
University of Waterloo
P
Perouz Taslakian
ServiceNow Research
Spandana Gella
Spandana Gella
ServiceNow AI Research
Multimodal Foundational ModelsGUI AgentsSafety & Security
Sai Rajeswar
Sai Rajeswar
Staff Research Scientist, Adjunct Professor, Mila, ServiceNow
machine learninggenerative modelsreinforcement learning
D
David Vazquez
ServiceNow Research
C
Christopher Pal
ServiceNow Research; Mila, Quebec AI Institute; Polytechnique Montréal; Canada CIFAR AI Chair
M
Marco Pedersoli
Mila, Quebec AI Institute; ÉTS Montréal