19,997 machine learning datasets
19,997 dataset results
This dataset is used to train and evaluate models for the detection of machine-paraphrased text.
The entire dataset consists of 61 different subjects, for each of which 12 radial OCT B-scans are collected at the Ophthalmology Department of Shanghai General Hospital by using DRI OCT-1 Atlantis (Topcon Corporation, Tokyo, Japan). The image size is 1024 × 992 pixels, corresponding to a field of view of 20.48 mm × 7.94 mm. For each subject, 2 radial OCT B-scans were randomly selected to ensure mutual exclusion. Two graders annotated these images manually through ITK-SNAP software into the optic disc and nine retinal layers under the supervision of a glaucoma specialist. For more details, please refer to our paper.
The SBCoseg dataset includes 889 groups of images and each group consists of 18 images with a common object, leading to 16002 images in total. The whole dataset is divided into five subsets: with ECFB, with TR, with MH, with SD, and Normal (normal data). The five subsets contain 193, 251, 82, 83, and 280 image groups, respectively. Each original image is in JPG format with a pixel size of 360 ×360, and each ground-truth image is in PNG format.
This dataset for abusive content detection in Twitter consists of two sets of annotations for the same set of tweets, one where the human annotators had access to the tweet's content and one where they didn't know the context.
10x Genomics dataset of sequenced TCRs barcoded by a panel of pMHCs (arranged on a dextramer)
Adaptive Biotechnologies' dataset of sequenced T cell repertoires labelled by patient age, HLA type, and CMV serostatus
ArtDL is a novel painting data set for iconography classification composed of images collected from online sources. Most of the paintings are from the Renaissance period and depict scenes or characters of Christian art. The data set is annotated with classes representing specific characters belonging to the Iconclass classification system.
fGn series used in the article to develop the simulations.
Seven different types of dry beans were used in this research, taking into account the features such as form, shape, type, and structure by the market situation. A computer vision system was developed to distinguish seven different registered varieties of dry beans with similar features in order to obtain uniform seed classification. For the classification model, images of 13,611 grains of 7 different registered dry beans were taken with a high-resolution camera. Bean images obtained by computer vision system were subjected to segmentation and feature extraction stages, and a total of 16 features; 12 dimensions and 4 shape forms, were obtained from the grains.
RUSS (Rapid Universal Support Service) is a dataset that consists of a collection of 741 real-world step-by-step natural language instructions (raw and annotated) from the open web, and for each: its corresponding webpage DOM, ground-truth ThingTalk, and ground-truth actions.
MSRB is a benchmarking dataset for marine snow removal of underwater images. Marine snow is one of the main degradation sources of underwater images that are caused by small particles, e.g., organic matter and sand, between the underwater scene and photosensors. The dataset consists of large-scale pairs of ground-truth and degraded images to calculate objective qualities for marine snow removal and to train a deep neural network. We propose two marine snow removal tasks using the dataset and show the first benchmarking results of marine snow removal.
The U.S. Broadband Coverage data set is a publicly available dataset that reports broadband coverage percentages at a zip code-level. The authors have used differential privacy to guarantee that the privacy of individual households is preserved. The data set also contains error ranges estimates, providing information on the expected error introduced by differential privacy per zip code.
First of its kind paired win-fail action understanding dataset with samples from the following domains: “General Stunts,” “Internet Wins-Fails,” “Trick Shots,” & “Party Games.” The task is to identify successful and failed attempts at various activities. Unlike existing action recognition datasets, intra-class variation is high making the task challenging, yet feasible.
Dataset for multimodal skills assessment focusing on assessing piano player’s skill level. Annotations include player's skills level, and song difficulty level. Bounding box annotations around pianists' hands are also provided.
AMT Objects is a large dataset of object centric videos suitable for training and benchmarking models for generating 3D models of objects from a small number of photos of the objects. The dataset consists of multiple views of a large collection of object instances.
RepLab 2013 dataset uses Twitter data in English and Spanish (more than 142,000 tweets). The balance between both languages depends on the availability of data for each of the entities included in the dataset. The corpus consists of a collection of tweets referring to a selected set of 61 entities from four domains: automotive, banking, universities and music/artists. The domain selection was done to offer a variety of scenarios for reputation studies.
Omiverse Object is a large-scale synthetic dataset of 60,000 images including both transparent and opaque objects in different scenes. It is used for depth completion of transparent objects from a single RGB-D view.
Auto-KWS is a dataset for customized keyword spotting, the task of detecting spoken keywords. The dataset closely resembles real world scenarios, as each recorder is assigned with an unique wake-up word and can choose their recording environment and familiar dialect freely.
Content of this dataset This dataset includes following files: