19,997 machine learning datasets
19,997 dataset results
The Wikidata Reference Logo Dataset (WiRLD), a comprehensive collection of reference logos specifically designed to address the challenges of large-scale logo identification. Recognizing the limitations of existing logo datasets, which often have a restricted number of logo classes or lack public availability, the authors curated WiRLD to facilitate research on more realistic, large-scale logo identification tasks. WiRLD contains 100,000 reference logo images sourced from Wikidata, representing 100,000 distinct logo classes. Each entity in the dataset has one corresponding logo image. The dataset's focus on providing a vast and readily accessible collection of reference logos makes it particularly valuable for evaluating one-shot logo identification methods, especially in large-scale scenarios
The Wikidata Reference Logo Dataset (WiRLD), a comprehensive collection of reference logos specifically designed to address the challenges of large-scale logo identification. Recognizing the limitations of existing logo datasets, which often have a restricted number of logo classes or lack public availability, the authors curated WiRLD to facilitate research on more realistic, large-scale logo identification tasks. WiRLD contains 100,000 reference logo images sourced from Wikidata, representing 100,000 distinct logo classes. Each entity in the dataset has one corresponding logo image. The dataset's focus on providing a vast and readily accessible collection of reference logos makes it particularly valuable for evaluating one-shot logo identification methods, especially in large-scale scenarios.
A dataset of images obtained from DALL-E 3 for 67 countries and 10 concept classes, similar to DollarStreet images.
We introduce a new style- and category-agnostic floor plan image parsing benchmark developed in collaboration with professional architectural designers. This benchmark includes 25 categories of space and adjacency labels (19 space elements and 6 adjacency elements), offering a more diverse and comprehensive representation of common design elements across various graphical styles and design categories. It sets a new standard for the level of diversity and complexity of floor plan image parsing tasks oriented towards real-world applications, far exceeding the scope of existing datasets. This benchmark is available at https://doi.org/10.7910/DVN/MDIRHE.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts. To evaluate their understanding of such knowledge, we introduce WorldCuisines, a massive-scale benchmark for multilingual and multicultural, visually grounded language understanding. This benchmark includes a visual question answering (VQA) dataset with text-image pairs across 30 languages and dialects, spanning 9 language families and featuring over 1 million data points, making it the largest multicultural VQA benchmark to date. It includes tasks for identifying dish names and their origins. We provide evaluation datasets in two sizes (12k and 60k instances) alongside a training dataset (1 million instances). Our findings show that while VLMs perform better with correct location context, they struggle with adversarial contexts and predicting specific regional cuisines and languages. To support future research, we release
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Two versions of the dataset are offered: one is the full dataset used to train the models in DeformPAM, and the other is a mini dataset for easier examination. Both datasets include data for the supervised and finetuning stages of granular pile shaping, rope shaping, and T-shirt unfolding.
Omni-Image is built as a challenging but tractable dataset for continual learning and few-shot learning.
A dataset of 2D robot recordings with 21 different symbols.
Existing arithmetic benchmarks have a limited number of multiple-choice questions. To address this gap, MathMC is created including 1,000 Chinese mathematical multiple-choice questions with detailed explanations and focusing on math problems typically encountered in grades 4 to 6. It features a wide range of question types, including arithmetic, algebra, geometry, statistics, reasoning, and more, enhancing the diversity of current Chinese arithmetic datasets.
We provide synthetic reflectance, direct shading (shading due to surface geometry and illumination conditions), ambient light and shadow cast ground-truth images. The dataset contains garden/park like natural (out-door) scenes including trees, plants, bushes, fences, etc. Furthermore, scenes are rendered with different types of terrains, landscapes, and lighting conditions. Addition-ally, real HDR sky images with a parallel light source are used to provide realistic ambient light. Moreover, light source properties are designed to model daytime lighting conditions to enrich the photometric effects.
The synthetic ShapeNet intrinsic image decomposition dataset of 90,000 images. 50,000 of them were used for training the deep CNN models of CVIU'2021 - see Section 4 of the paper. This is the extension of the first release of the synthetic ShapeNet intrinsic image decomposition dataset of 20,000 images used for training the deep CNN models IntrinsicNet and RetiNet of CVPR'2018. See Section 4.1 of the CVPR paper for the details of the data rendering.
Data URL: https://data.mendeley.com/datasets/9k892pzkfx/1
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Release decompile-ghidra-100k, a subset of 100k training samples (25k per optimization level). We provide a training script that runs in ~3.5 hours on a single A100 40G GPU. It achieves a 0.26 re-executability rate, with a total cost of under $20 for quick replication of LLM4Decompile.
Characterising multimedia content with relevant, reliable and discriminating tags is vital for multimedia information retrieval. With the rapid expansion of digital multimedia content, alternative methods to the existing explicit tagging are needed to enrich the pool of tagged content. Currently, social media websites encourage users to tag their content. However, the users’ intent when tagging multimedia content does not always match the information retrieval goals. A large portion of user defined tags are either motivated by increasing the popularity and reputation of a user in an online com-munity or based on individual and egoistic judgments. Moreover, users do not evaluate media content on the same criteria. Some might tag multimedia content with words to express their emotion while others might use tags to describe the content. For example, a picture receive different tags based on the objects in the image, the camera by which the picture was taken or the emotion a user felt look