Recurrent Affine Transformation for Text-to-image Synthesis

Senmao Ye, Fei Liu, Minkui Tan

2022-04-22Text-to-Image Generation

Abstract

Text-to-image synthesis aims to generate natural images conditioned on text descriptions. The main difficulty of this task lies in effectively fusing text information into the image synthesis process. Existing methods usually adaptively fuse suitable text information into the synthesis process with multiple isolated fusion blocks (e.g., Conditional Batch Normalization and Instance Normalization). However, isolated fusion blocks not only conflict with each other but also increase the difficulty of training (see first page of the supplementary). To address these issues, we propose a Recurrent Affine Transformation (RAT) for Generative Adversarial Networks that connects all the fusion blocks with a recurrent neural network to model their long-term dependency. Besides, to improve semantic consistency between texts and synthesized images, we incorporate a spatial attention model in the discriminator. Being aware of matching image regions, text descriptions supervise the generator to synthesize more relevant image contents. Extensive experiments on the CUB, Oxford-102 and COCO datasets demonstrate the superiority of the proposed model in comparison to state-of-the-art models \footnote{https://github.com/senmaoy/Recurrent-Affine-Transformation-for-Text-to-image-Synthesis.git}

Results

Task	Dataset	Metric	Value	Model
Image Generation	COCO (Common Objects in Context)	FID	14.6	RAT-GAN
Image Generation	Oxford 102 Flowers	FID	16.04	RAT-GAN
Image Generation	Oxford 102 Flowers	Inception score	4.09	RAT-GAN
Image Generation	CUB	FID	10.21	RAT-GAN
Image Generation	CUB	Inception score	5.36	RAT-GAN
Text-to-Image Generation	COCO (Common Objects in Context)	FID	14.6	RAT-GAN
Text-to-Image Generation	Oxford 102 Flowers	FID	16.04	RAT-GAN
Text-to-Image Generation	Oxford 102 Flowers	Inception score	4.09	RAT-GAN
Text-to-Image Generation	CUB	FID	10.21	RAT-GAN
Text-to-Image Generation	CUB	Inception score	5.36	RAT-GAN
10-shot image generation	COCO (Common Objects in Context)	FID	14.6	RAT-GAN
10-shot image generation	Oxford 102 Flowers	FID	16.04	RAT-GAN
10-shot image generation	Oxford 102 Flowers	Inception score	4.09	RAT-GAN
10-shot image generation	CUB	FID	10.21	RAT-GAN
10-shot image generation	CUB	Inception score	5.36	RAT-GAN
1 Image, 2*2 Stitchi	COCO (Common Objects in Context)	FID	14.6	RAT-GAN
1 Image, 2*2 Stitchi	Oxford 102 Flowers	FID	16.04	RAT-GAN
1 Image, 2*2 Stitchi	Oxford 102 Flowers	Inception score	4.09	RAT-GAN
1 Image, 2*2 Stitchi	CUB	FID	10.21	RAT-GAN
1 Image, 2*2 Stitchi	CUB	Inception score	5.36	RAT-GAN

Recurrent Affine Transformation for Text-to-image Synthesis

Abstract

Results

Related Papers

Recurrent Affine Transformation for Text-to-image Synthesis

Abstract

Results

Related Papers