Detect Anything in Graphic Design: Element-Level Rewards for Autoregressive Detection
为解决图形设计中元素组成顺序未被利用的问题,提出DAD模型进行有序解构检测,并通过EleRPO方法优化训练,提高了检测性能。
为解决图形设计中元素组成顺序未被利用的问题,提出DAD模型进行有序解构检测,并通过EleRPO方法优化训练,提高了检测性能。
本文提出了一种新架构Giraffe,通过单个[IMG]标记将隐藏文本表示映射到视觉嵌入空间,以解决图形设计生成中多模态大语言模型生成媒体能力有限的问题。
该研究通过在预训练的图像编辑扩散变换器中隐式生成布局,解决了图形设计合成中布局僵硬和资产失真的问题。
This work addresses the inefficiency in traditional neural network optimization, where additive weight updates induce imbalanced relative perturbations across weights of differing magnitudes. To mitigate this, the authors propose a hybrid exponential-linear reparameterization of weights that integrates a sign-aware symmetric exponential pathway with an identity linear pathway. This construction, augmented with learnable scale, curvature, and offset parameters, induces a curved weight geometry wherein optimization step sizes scale proportionally with weight magnitudes. Coupled with a mismatched initialization strategy to encourage early symmetry breaking, the method achieves equivalent validation loss on OpenWebText using 1.32–1.49× fewer training steps across various Transformer architectures, with particularly pronounced gains for wide models.
This work addresses the limitation of existing DiT research, which overly relies on ImageNet class-conditional generation as a sole evaluation setting and fails to reflect model performance in broader applications such as text-to-image (T2I) synthesis. To this end, we propose NanoGen, a unified training and evaluation framework that supports diverse diffusion approaches—including RAE, VAE, pixel-space, and MeanFlow—across both ImageNet and T2I tasks with only a 12-line configuration switch. We further introduce DiffusionBench, a comprehensive benchmark encompassing both task types. Experiments across 21 latent diffusion models reveal a significant negative correlation between ImageNet and T2I performance, with Pearson correlation coefficients ranging from −0.377 to −0.580, underscoring the misleading nature of single-task evaluation and validating the necessity and effectiveness of DiffusionBench.
为解决图形设计中元素组成顺序未被利用的问题,提出DAD模型进行有序解构检测,并通过EleRPO方法优化训练,提高了检测性能。
本文提出了一种新架构Giraffe,通过单个[IMG]标记将隐藏文本表示映射到视觉嵌入空间,以解决图形设计生成中多模态大语言模型生成媒体能力有限的问题。
该研究通过在预训练的图像编辑扩散变换器中隐式生成布局,解决了图形设计合成中布局僵硬和资产失真的问题。
This work addresses the inefficiency in traditional neural network optimization, where additive weight updates induce imbalanced relative perturbations across weights of differing magnitudes. To mitigate this, the authors propose a hybrid exponential-linear reparameterization of weights that integrates a sign-aware symmetric exponential pathway with an identity linear pathway. This construction, augmented with learnable scale, curvature, and offset parameters, induces a curved weight geometry wherein optimization step sizes scale proportionally with weight magnitudes. Coupled with a mismatched initialization strategy to encourage early symmetry breaking, the method achieves equivalent validation loss on OpenWebText using 1.32–1.49× fewer training steps across various Transformer architectures, with particularly pronounced gains for wide models.
This work addresses the limitation of existing DiT research, which overly relies on ImageNet class-conditional generation as a sole evaluation setting and fails to reflect model performance in broader applications such as text-to-image (T2I) synthesis. To this end, we propose NanoGen, a unified training and evaluation framework that supports diverse diffusion approaches—including RAE, VAE, pixel-space, and MeanFlow—across both ImageNet and T2I tasks with only a 12-line configuration switch. We further introduce DiffusionBench, a comprehensive benchmark encompassing both task types. Experiments across 21 latent diffusion models reveal a significant negative correlation between ImageNet and T2I performance, with Pearson correlation coefficients ranging from −0.377 to −0.580, underscoring the misleading nature of single-task evaluation and validating the necessity and effectiveness of DiffusionBench.