Analyzing Coarse-to-fine Generation of Diffusion Models through Image Editing Perspectives

Citations

SCOPUS

0

초록

Diffusion models involve a forward noising process and a reverse denoising process. Their diffusion-like sampling process have an intrinsic nature of coarse-to-fine generation (i.e., the initial stages build coarse-grained outlines and the latter stages focus on fine details), which has been implicitly demonstrated in prior works. The goal of this work is to present an analysis method that explicitly reveals the diffusion models' coarse-to-fine behavior, through conducting image editing that mixes the latent features of a pair of samples (source and reference). Different from the prior analysis that defined a specific stage based on the step number T (e.g., fine: [0, 333], medium: [333, 666], and coarse: [666, 999]), our analysis more carefully partitions the generative sequence based on variation of noise and U-Net latent features during the sequence. We then show that, with a few modifications, our analysis method can perform the image editing task working on top of existing pre-trained diffusion models while allowing controllability, without any additional training or fine-tuning. Finally, we propose an extended metric based on LPIPS (Learned Perceptual Image Patch Similarity), which is more suitable for evaluating the controllable image editing task. Experiments on several widely-used benchmarks show the effectiveness of our method. Our code is available at: https://github.com/gyeomo/Coarse-to-fine_Generation_of_Diffusion.

키워드

coarse-to-fine generationdiffusion modelsimage editingDiffusion
제목
Analyzing Coarse-to-fine Generation of Diffusion Models through Image Editing Perspectives
저자
Kim, SeonggyeomKim, MinjuBang, MinjuChae, Dong-Kyu
DOI
10.1145/3748522.3779967
발행일
2026-06
유형
Conference Paper
저널명
Proceedings of the ACM Symposium on Applied Computing
페이지
1020 ~ 1028