🤖 AI Summary
This work addresses the susceptibility of visual autoregressive models (VARs) to cross-scale error propagation in complex multi-object scenes, which degrades generation quality. To mitigate this issue, the authors propose SynVAR, the first training-free enhancement framework that improves VAR performance through a spatial-semantic co-control mechanism. SynVAR innovatively integrates a three-stage strategy: global spatial structure guidance, receptive field constraints to suppress early semantic confusion, and high-frequency detail compensation. This synergistic approach effectively alleviates error accumulation during autoregressive generation. Experimental results demonstrate that SynVAR substantially enhances the modeling capacity and image generation fidelity of VARs in complex scenes, without requiring any additional training.
📝 Abstract
VAR has gained widespread popularity due to its next-scale prediction paradigm. However, it faces substantial performance bottlenecks when handling complex scenes with multiple objects and attributes. Existing diffusion-based enhancement methods fail to adequately address the unique challenge of cross-scale error propagation and accumulation in VAR. To this end, we propose SynVAR, the first training-free enhancement framework specifically tailored for the VAR paradigm, which introduces a spatial-semantic collaborative control strategy to effectively suppress propagation error and improve generation quality. SynVAR comprises three key components: (1) Global guidance to ensure reasonable spatial structure in the early stages, (2) Receptive field constraints to mitigate early-stage semantic confusion, (3) High-frequency compensation to recover fine-grained details. Extensive quantitative and qualitative experiments demonstrate the significant improvements in the ability of SynVAR to enhance the VAR's capability for complex scene modeling.