TAECE : T2I-Adapter with Enhanced Color Expression for Improving Conditional Text-to-Image Generation Capabilities

Citations

WEB OF SCIENCE

1
Citations

SCOPUS

1

초록

The text-to-image diffusion model has advanced, enabling the generation of complex images from text as well as sketches, key poses, and segmentation maps. However, these models face challenges in accurately representing detailed scenes or real-world elements. This study addresses these challenges by proposing a method to enhance image generation ability based on both text and sketch. Our approach introduces an adapter incorporating a deformable convolution network (DCN) to process sketch inputs, allowing structural information to be retained in generated images. Additionally, we integrate large language models (LLMs) to enrich textual descriptions with nuanced color expressions. By combining structural input and enriched text, our model produces images that are not only realistic but visually appealing. This method significantly enhances the model's capacity to capture intricate details. Experimental results demonstrate that our model outperforms existing conditional text-to-image models in visual quality. Overall, this study contributes to image generation technology by advancing color representation via LLMs, fostering the creation of more visually consistent and detailed images. The proposed approach presents broad applicability, offering a notable contribution to text-to-image synthesis and advancing image generation techniques for greater realism.

키워드

computer visionimage generationtext-to-image synthesisComplex imageDiffusion modelImage diffusionImage generationsImages synthesisKey poseLanguage modelReal-worldSegmentation mapText-to-image synthesis
제목
TAECE : T2I-Adapter with Enhanced Color Expression for Improving Conditional Text-to-Image Generation Capabilities
저자
Seo, HyeinJeong, YunaChoi, Yong Suk
DOI
10.1145/3672608.3707847
발행일
2025-05
유형
Proceedings Paper
저널명
40TH ANNUAL ACM SYMPOSIUM ON APPLIED COMPUTING
페이지
1180 ~ 1187