Creative Painting with Latent Diffusion Models

Xianchao Wu


Abstract
Artistic painting has achieved significant progress during recent years. Using a variational autoencoder to connect the original images with compressed latent spaces and a cross attention enhanced U-Net as the backbone of diffusion, latent diffusion models (LDMs) have achieved stable and high fertility image generation. In this paper, we focus on enhancing the creative painting ability of current LDMs in two directions, textual condition extension and model retraining with Wikiart dataset. Through textual condition extension, users’ input prompts are expanded with rich contextual knowledge for deeper understanding and explaining the prompts. Wikiart dataset contains 80K famous artworks drawn during recent 400 years by more than 1,000 famous artists in rich styles and genres. Through the retraining, we are able to ask these artists to draw artistic and creative paintings on modern topics. Direct comparisons with the original model show that the creativity and artistry are enriched.
Anthology ID:
2022.cai-1.8
Volume:
Proceedings of the Second Workshop on When Creative AI Meets Conversational AI
Month:
October
Year:
2022
Address:
Gyeongju, Republic of Korea
Venue:
CAI
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
59–80
Language:
URL:
https://aclanthology.org/2022.cai-1.8
DOI:
Bibkey:
Cite (ACL):
Xianchao Wu. 2022. Creative Painting with Latent Diffusion Models. In Proceedings of the Second Workshop on When Creative AI Meets Conversational AI, pages 59–80, Gyeongju, Republic of Korea. Association for Computational Linguistics.
Cite (Informal):
Creative Painting with Latent Diffusion Models (Wu, CAI 2022)
Copy Citation:
PDF:
https://preview.aclanthology.org/ingestion-script-update/2022.cai-1.8.pdf