19,997 machine learning datasets
19,997 dataset results
Image dataset with about 600K Flickr photos.
This dataset consists of 2,192 high-quality traditional Chinese landscape paintings (中国山水画). All paintings are sized 512x512, from the following sources: * Princeton University Art Museum, 362 paintings * Harvard University Art Museum, 101 paintings * Metropolitan Museum of Art, 428 paintings * Smithsonian's Freer Gallery of Art, 1,301 paintings
The dataset is composed of 100 video sequences densely annotated with 60K bounding boxes, 17 sequence attributes, 13 action verb attributes and 29 target object attributes.
A new high accuracy Turkish morphology dataset.
Twitter Cyberthreat Detection Dataset is a dataset that contains tweets from two sets of accounts related to cybersecurity. The tweets are annotated with different information such as whether they contain security-related information and named entities.
This dataset contains two subsets of flood images from Twitter: The Harz17 dataset comprises images from tweets containing flood-related keywords during the occurrence of a flood in the Harz region in Germany in July 2017. Similarly, the Rhine18 dataset comprises images related to a flood of the river Rhine in January 2018.
The TWT16 dataset contains ~30k conversations in Twitter, collected from January to June 2016.
Contains three difficult real-world scenarios: uncontrolled videos taken by UAVs and manned gliders, as well as controlled videos taken on the ground. Over 160,000 annotated frames forhundreds of ImageNet classes are available, which are used for baseline experiments that assess the impact of known and unknown image artifacts and other conditions on common deep learning-based object classification approaches.
This dataset comprises over 26,000 full names annotated with genders.
UMC005 English-Urdu is a parallel corpus of texts in English and Urdu language with sentence alignments. The corpus can be used for experiments with statistical machine translation.
The MultiUN parallel corpus is extracted from the United Nations Website , and then cleaned and converted to XML at Language Technology Lab in DFKI GmbH (LT-DFKI), Germany. The documents were published by UN from 2000 to 2009.
Urban Dict spelling variant is a variant spelling dataset for use of NLP research in the informal domain. It consists of around 25k variant spelling pairs form UrbanDictionary.
Comprises twenty individuals picking up and placing objects of varying weights to and from cabinet and table locations at various heights.
The Visual Discriminative Question Generation (VDQG) dataset contains 11202 ambiguous image pairs collected from Visual Genome. Each image pair is annotated with 4.6 discriminative questions and 5.9 non-discriminative questions on average.
Enable visual relation detection and serves as an extension to Visual Genome (VG).
The VIA dataset is a dataset for aiding the visually impaired. The proposed datase1 consists of 342 images divided into two classes: 175 of them are “clear-path” and 167 are “nonclear” path. They were taken using a smartphone camera and resized to 750 × 1000 pixels. The smartphone was placed in the user chest height and inclined approximately 30 to 60 from the ground, so it could capture a few meters of the path ahead, and beyond the reach of a regular white cane
A large video dataset with dynamic content.
A novel dataset that consists of content-based and video-specific features extracted from publicly available scientific video lectures and several metrics related to user engagement.
Collects 60 reference sequences and 540 impaired sequences.
The WIDER Attribute dataset is a human attribute recognition dataset with human attribute and image event annotations. Images are selected from the WIDER dataset. There are a total of 13,789 images. A bounding box is annotated for each person in these images, with no more than 20 people (with top resolutions) in a crowd image, resulting in 57,524 boxes in total and 4+ boxes per image on average. For each bounding box, 14 distinct human attributes are labelled. There are 805,336 labels in total.