TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Papers/Semantic Object Accuracy for Generative Text-to-Image Synt...

Semantic Object Accuracy for Generative Text-to-Image Synthesis

Tobias Hinz, Stefan Heinrich, Stefan Wermter

2019-10-29Text-to-Image GenerationImage CaptioningImage Generation
PaperPDFCodeCode(official)

Abstract

Generative adversarial networks conditioned on textual image descriptions are capable of generating realistic-looking images. However, current methods still struggle to generate images based on complex image captions from a heterogeneous domain. Furthermore, quantitatively evaluating these text-to-image models is challenging, as most evaluation metrics only judge image quality but not the conformity between the image and its caption. To address these challenges we introduce a new model that explicitly models individual objects within an image and a new evaluation metric called Semantic Object Accuracy (SOA) that specifically evaluates images given an image caption. The SOA uses a pre-trained object detector to evaluate if a generated image contains objects that are mentioned in the image caption, e.g. whether an image generated from "a car driving down the street" contains a car. We perform a user study comparing several text-to-image models and show that our SOA metric ranks the models the same way as humans, whereas other metrics such as the Inception Score do not. Our evaluation also shows that models which explicitly model objects outperform models which only model global image characteristics.

Results

TaskDatasetMetricValueModel
Image GenerationCOCO (Common Objects in Context)FID24.7OP-GAN
Image GenerationCOCO (Common Objects in Context)Inception score27.88OP-GAN
Image GenerationCOCO (Common Objects in Context)SOA-C35.85OP-GAN
Text-to-Image GenerationCOCO (Common Objects in Context)FID24.7OP-GAN
Text-to-Image GenerationCOCO (Common Objects in Context)Inception score27.88OP-GAN
Text-to-Image GenerationCOCO (Common Objects in Context)SOA-C35.85OP-GAN
10-shot image generationCOCO (Common Objects in Context)FID24.7OP-GAN
10-shot image generationCOCO (Common Objects in Context)Inception score27.88OP-GAN
10-shot image generationCOCO (Common Objects in Context)SOA-C35.85OP-GAN
1 Image, 2*2 StitchiCOCO (Common Objects in Context)FID24.7OP-GAN
1 Image, 2*2 StitchiCOCO (Common Objects in Context)Inception score27.88OP-GAN
1 Image, 2*2 StitchiCOCO (Common Objects in Context)SOA-C35.85OP-GAN

Related Papers

fastWDM3D: Fast and Accurate 3D Healthy Tissue Inpainting2025-07-17Synthesizing Reality: Leveraging the Generative AI-Powered Platform Midjourney for Construction Worker Detection2025-07-17FashionPose: Text to Pose to Relight Image Generation for Personalized Fashion Visualization2025-07-17A Distributed Generative AI Approach for Heterogeneous Multi-Domain Environments under Data Sharing constraints2025-07-17Pixel Perfect MegaMed: A Megapixel-Scale Vision-Language Foundation Model for Generating High Resolution Medical Images2025-07-17Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos2025-07-16FADE: Adversarial Concept Erasure in Flow Models2025-07-16CharaConsist: Fine-Grained Consistent Character Generation2025-07-15