GDB-Reward: From Evaluation Metrics to Training Rewards for Graphic Design
本文提出GDB-Reward框架,将图形设计评估指标转化为统一的强化学习奖励,有效提高文本到图像模型生成图形设计时的质量和准确性。
本文提出GDB-Reward框架,将图形设计评估指标转化为统一的强化学习奖励,有效提高文本到图像模型生成图形设计时的质量和准确性。
This study addresses the limitation of existing text-to-image models, which rely on holistic preference scores and fail to capture designers’ fine-grained judgments across multiple dimensions such as layout, color, and visual hierarchy. To bridge this gap, the authors introduce TASTE, a dataset comprising evaluations by ten professional designers who rate outputs from four state-of-the-art text-to-image models across nine design dimensions and annotate hallucinated content. They propose a criterion-agnostic, multidimensional preference evaluation framework that employs Kendall’s tau, majority probability, and Condorcet cycles for statistical significance testing. Experimental results demonstrate that all evaluated dimensions significantly deviate from random scoring, yet the highest agreement between current models and designer consensus reaches only 0.55. In contrast, a lightweight prediction head trained on TASTE achieves a correlation of 0.611—approaching the human upper bound of 0.741—highlighting the limited capacity of contemporary vision-language models to understand nuanced design preferences.
This work addresses the lack of standardized, automated evaluation methods for design animation video generation, which hinders objective assessment of generation quality under structured constraints. To bridge this gap, we propose the first multidimensional automatic evaluation framework tailored specifically for design animations. Leveraging computer vision and video analysis techniques, the framework quantifies key generative attributes across four dimensions: layout fidelity, motion correctness, temporal consistency, and content fidelity. Operating without human intervention, it delivers an objective and reproducible benchmark that enables fair comparison among diverse generative models and supports sustained progress in the field.
本文提出GDB-Reward框架,将图形设计评估指标转化为统一的强化学习奖励,有效提高文本到图像模型生成图形设计时的质量和准确性。
This study addresses the limitation of existing text-to-image models, which rely on holistic preference scores and fail to capture designers’ fine-grained judgments across multiple dimensions such as layout, color, and visual hierarchy. To bridge this gap, the authors introduce TASTE, a dataset comprising evaluations by ten professional designers who rate outputs from four state-of-the-art text-to-image models across nine design dimensions and annotate hallucinated content. They propose a criterion-agnostic, multidimensional preference evaluation framework that employs Kendall’s tau, majority probability, and Condorcet cycles for statistical significance testing. Experimental results demonstrate that all evaluated dimensions significantly deviate from random scoring, yet the highest agreement between current models and designer consensus reaches only 0.55. In contrast, a lightweight prediction head trained on TASTE achieves a correlation of 0.611—approaching the human upper bound of 0.741—highlighting the limited capacity of contemporary vision-language models to understand nuanced design preferences.
This work addresses the lack of standardized, automated evaluation methods for design animation video generation, which hinders objective assessment of generation quality under structured constraints. To bridge this gap, we propose the first multidimensional automatic evaluation framework tailored specifically for design animations. Leveraging computer vision and video analysis techniques, the framework quantifies key generative attributes across four dimensions: layout fidelity, motion correctness, temporal consistency, and content fidelity. Operating without human intervention, it delivers an objective and reproducible benchmark that enables fair comparison among diverse generative models and supports sustained progress in the field.