OmniGen-AR: AutoRegressive Any-to-Image Generation
Existing autoregressive vision generation models typically support only a single modality of conditioning, limiting their ability to meet real-world demands for image synthesis driven by multiple types of control signals. This work proposes OmniGen-AR, a unified autoregressive framework that discretizes diverse multimodal conditions—including text, spatial layouts, and visual context—into a shared token space via a common visual-text tokenizer. To prevent information leakage during training, the model introduces Decoupled Causal Attention (DCA), which separates causal dependencies between conditioning and content tokens while preserving the standard autoregressive prediction pipeline during inference. Evaluated on benchmarks such as GenEval (0.63) and VBench (80.02), OmniGen-AR achieves state-of-the-art or competitive performance, demonstrating its effectiveness in flexible, high-fidelity multimodal image generation.