PALATE: Personalized Aesthetic Learning through Adaptive Taste Evolution for Multi-User Portrait Retouching
本文提出PALATE框架,通过固定图像编辑器并个性化选择重绘候选图像来解决多用户肖像修饰中的个人审美差异问题。
本文提出PALATE框架,通过固定图像编辑器并个性化选择重绘候选图像来解决多用户肖像修饰中的个人审美差异问题。
This work addresses the challenge of achieving pixel-level precision in sketch-based image editing, which is hindered by the absence of high-quality datasets that jointly encode geometric constraints and semantic instructions. To overcome this limitation, we propose SI-Edit, a novel framework accompanied by SI-Data—the first high-quality dataset specifically designed for instruction-guided sketch editing. SI-Data comprises quadruplets of images, sketches, natural language instructions, and corresponding edited results, automatically generated using multimodal large language models to enable spatial-semantic co-learning. By jointly modeling geometric sketches and semantic instructions, our method achieves high-fidelity local deformations and significantly outperforms existing approaches in both structural preservation and alignment with user intent.
This work addresses the challenges of native 2K-resolution virtual try-on with multiple garments, where high resolution leads to prohibitive memory consumption and diffusion models tend to oversmooth fine fabric details. To overcome these limitations, the authors propose an end-to-end, mask-free generative framework featuring three key innovations: an Adaptive Token Packing (ATP) mechanism that dynamically compresses sequence length in 2D feature space, a Multi-dimensional Try-on Reward (MTR) system that jointly optimizes texture fidelity and physical plausibility, and WearWow-2K—the first native 2K-resolution triplet dataset for multi-garment try-on. The proposed method substantially reduces memory overhead while preserving high-frequency textile details, achieving state-of-the-art performance in multi-clothing synthesis and outperforming existing commercial baselines.
This work addresses the challenge of maintaining identity and geometric consistency under extreme illumination changes in image relighting. To this end, it formulates relighting as an illumination feature transfer problem and introduces a Consistent Feature Transfer (CFT) training framework. CFT jointly models noise-to-image generation and source-to-target illumination transfer via rectified flows and incorporates trajectory-level supervision to explicitly disentangle illumination and content features. The authors also construct the first large-scale portrait dataset featuring complex lighting conditions. Experimental results demonstrate that the proposed method significantly outperforms existing approaches in both relighting quality and content fidelity, and further generalizes effectively to other editing tasks such as style transfer.
This work addresses semantic misalignment and detail inconsistency in multi-reference image generation caused by redundant and ambiguous natural language instructions. To mitigate these issues, we propose a structured dictionary-based intent representation that explicitly encodes the semantic intents of multiple reference images, replacing conventional textual prompts to reduce ambiguity. We introduce the first high-quality structured dataset tailored for complex multi-reference scenarios, along with an aligned training framework and a dedicated evaluation benchmark. Experimental results demonstrate that our approach significantly outperforms existing models on both public and newly curated benchmarks, achieving superior semantic alignment and detail fidelity—particularly under complex instruction settings.
本文提出PALATE框架,通过固定图像编辑器并个性化选择重绘候选图像来解决多用户肖像修饰中的个人审美差异问题。
This work addresses the challenge of achieving pixel-level precision in sketch-based image editing, which is hindered by the absence of high-quality datasets that jointly encode geometric constraints and semantic instructions. To overcome this limitation, we propose SI-Edit, a novel framework accompanied by SI-Data—the first high-quality dataset specifically designed for instruction-guided sketch editing. SI-Data comprises quadruplets of images, sketches, natural language instructions, and corresponding edited results, automatically generated using multimodal large language models to enable spatial-semantic co-learning. By jointly modeling geometric sketches and semantic instructions, our method achieves high-fidelity local deformations and significantly outperforms existing approaches in both structural preservation and alignment with user intent.
This work addresses the challenges of native 2K-resolution virtual try-on with multiple garments, where high resolution leads to prohibitive memory consumption and diffusion models tend to oversmooth fine fabric details. To overcome these limitations, the authors propose an end-to-end, mask-free generative framework featuring three key innovations: an Adaptive Token Packing (ATP) mechanism that dynamically compresses sequence length in 2D feature space, a Multi-dimensional Try-on Reward (MTR) system that jointly optimizes texture fidelity and physical plausibility, and WearWow-2K—the first native 2K-resolution triplet dataset for multi-garment try-on. The proposed method substantially reduces memory overhead while preserving high-frequency textile details, achieving state-of-the-art performance in multi-clothing synthesis and outperforming existing commercial baselines.
This work addresses the challenge of maintaining identity and geometric consistency under extreme illumination changes in image relighting. To this end, it formulates relighting as an illumination feature transfer problem and introduces a Consistent Feature Transfer (CFT) training framework. CFT jointly models noise-to-image generation and source-to-target illumination transfer via rectified flows and incorporates trajectory-level supervision to explicitly disentangle illumination and content features. The authors also construct the first large-scale portrait dataset featuring complex lighting conditions. Experimental results demonstrate that the proposed method significantly outperforms existing approaches in both relighting quality and content fidelity, and further generalizes effectively to other editing tasks such as style transfer.
This work addresses semantic misalignment and detail inconsistency in multi-reference image generation caused by redundant and ambiguous natural language instructions. To mitigate these issues, we propose a structured dictionary-based intent representation that explicitly encodes the semantic intents of multiple reference images, replacing conventional textual prompts to reduce ambiguity. We introduce the first high-quality structured dataset tailored for complex multi-reference scenarios, along with an aligned training framework and a dedicated evaluation benchmark. Experimental results demonstrate that our approach significantly outperforms existing models on both public and newly curated benchmarks, achieving superior semantic alignment and detail fidelity—particularly under complex instruction settings.