🤖 AI Summary
This study addresses the bottleneck of subject-personalized generation, which typically relies on expensive synthetic data and complex preprocessing. We propose CRAFT, a framework grounded in the "where to look" principle that fine-tunes a reference-aware multimodal diffusion transformer via attention alignment, mask gating, and multi-dimensional reward mechanisms. This approach achieves high-fidelity identity preservation without requiring synthesized target supervision. Evaluated on XVerseBench, CRAFT attains state-of-the-art performance using only 10,000 reference images, significantly outperforming existing methods that depend on millions of synthetic samples. Consequently, this work substantially reduces both data costs and training barriers for personalized generative modeling.
📝 Abstract
Subject-driven image personalization---generating new images that preserve the identity of one or several reference subjects in novel scenes---is a foundational capability for modern visual content creation. It is currently dominated by generalized methods that fine-tune a pretrained multimodal diffusion transformer (MMDiT) on hundreds of thousands to millions of paired \emph{(reference, composed-target)} examples, where each composed target is a synthesized image of the subject in a novel scene. Producing such targets demands a costly multi-stage curation pipeline---LLM-based prompt generation, T2I-based composed-target synthesis, reference-subject extraction, VLM-based quality filtering, and correspondence labeling---and tightly couples each method to a particular target synthesizer and curation choice. We introduce \emph{CRAFT} (Constrained Reward via Attention Fine-Tuning), a single-step ReFL framework that fine-tunes a pre-trained \emph{reference-aware} MMDiT via LoRA adapters using a compact reference-only data construction---$10$K reference images and subject masks, with no composed-target supervision. CRAFT realizes a \emph{Where to look} principle: attention-level rewards align noise- and phrase-token attention with the correct reference subject, and the resulting per-subject attention masks gate a pixel-level identity reward to keep image-space supervision consistent with the learned attention routing. Applied to FLUX.2-klein-9B, CRAFT achieves state-of-the-art performance on XVerseBench \rev{while using no composed-target supervision---only $10$K reference-only samples, whereas prior generalized methods require $150$K to over $2$M composed-target pairs}. The same recipe transfers to other reference-aware backbones, consistently improving performance. Project page: https://jihun999.github.io/projects/CRAFT/.