VAnim: Rendering-Aware Sparse State Modeling for Structure-Preserving Vector Animation
Existing methods for text-to-SVG animation generation struggle to simultaneously preserve topological consistency and model non-rigid deformations, while also lacking support for open-domain instructions. This work proposes VAnim—the first large language model framework tailored for open-domain, text-driven SVG animation generation—which formulates animation as sparse state updates over a persistent SVG DOM tree, drastically reducing sequence length by more than 9.8×. By integrating an identity-aware motion planning mechanism and rendering-aware reinforcement learning (GRPO) with a video-perceptual hybrid reward, VAnim achieves high-fidelity dynamic generation while maintaining structural validity and identity consistency. We further introduce SVGAnim-134k, the first vector animation benchmark, and demonstrate through experiments that VAnim significantly outperforms existing approaches in semantic alignment, motion quality, and structural preservation.