CLEVR-X
ImagesBSD-3-Clause LicenseIntroduced 2022-04-05
CLEVR-X is a dataset that extends the CLEVR dataset with natural language explanations in the context of VQA. It consists of 3.6 million natural language explanations for 850k question-image pairs.
For each image-question pair in the CLEVR dataset, CLEVR-X contains multiple structured textual explanations which are derived from the original scene graphs. By construction, the CLEVR-X explanations are correct and describe the reasoning and visual information that is necessary to answer a given question.
The CLEVR-X dataset consists of:
- A training set of 2,401,275 natural language explanations for 70,000 images.
- A validation set of 599,711 natural language explanations for 14,000 images.
- A test set of 644,151 natural language explanations for 15,000 images.