EditCrafter: Tuning-free High-Resolution Image Editing via Pretrained Diffusion Model
Existing diffusion models struggle to edit high-resolution images with arbitrary aspect ratios or resolutions significantly exceeding their training scale (e.g., 512×512), as naive tiling often introduces structural distortions and content duplication. This work proposes a fine-tuning-free editing framework that integrates tiled latent-space inversion with an enhanced noise-damped classifier-free guidance strategy (NDCFG++). By effectively leveraging the generative priors of pre-trained text-to-image diffusion models, the method achieves coherent and photorealistic high-resolution edits while preserving image identity. It supports inputs of arbitrary dimensions and consistently produces structurally consistent and detail-rich results across diverse resolutions, without requiring model fine-tuning or additional optimization.