FineCops-Ref

Introduced 2024-09-23

FineCops-Ref is a dataset for Compositional Referring Expression Comprehension (REC) that rigorously evaluates Vision-Language Models (VLMs) on compositional reasoning and their ability to identify inconsistencies between images and text. Beyond standard REC tasks, it challenges models with fine-grained correspondences involving objects, attributes, and relationships. The dataset comprises both training and testing sets, designed to thoroughly assess model performance across various difficulty level.