19,997 machine learning datasets
19,997 dataset results
CPM-Real is a dataset consisting of 3895 images representing real - makeup styles.
The COPA-HR dataset (Choice of plausible alternatives in Croatian) is a translation of the English COPA dataset by following the XCOPA dataset translation methodology. The dataset consists of 1000 premises (My body cast a shadow over the grass), each given a question (What is the cause?), and two choices (The sun was rising; The grass was cut), with a label encoding which of the choices is more plausible given the annotator or translator (The sun was rising).
CASP13 MQA is a dataset that contains predicted models for CASP13 targets and their scores.
GE852 is a dataset of 852 game engine repositories mined from GitHub in two languages, namely Java and C++. The dataset contains metadata of all the mined repositories including commits, pull requests, issues and so on. This dataset can lays the foundation for empirical investigation in the area of game engines.
Cry Wolf is a dataset for cyber security analysis tasks. It is an open-access dataset of 73 true and false Intrusion Detection System (IDS) alarms derived from real-world examples of "impossible travel" scenarios.
IITM-Bandersnatch is a dataset to evaluate traffic analysis techniques. The dataset comprises of data points of the form {encrypted traces, ground truth choices}. To collect each data point, we asked the viewer to watch Bandersnatch from the beginning and note down the choices made by them. At the same time, we collected the encrypted network traffic. As of now, our dataset contains information corresponding to 100 viewers who volunteered for this study.
This dataset is composed of the URLs of the top 1 million websites. The domains are ranked using the Alexa traffic ranking which is determined using a combination of the browsing behavior of users on the website, the number of unique visitors, and the number of pageviews. In more detail, unique visitors are the number of unique users who visit a website on a given day, and pageviews are the total number of user URL requests for the website. However, multiple requests for the same website on the same day are counted as a single pageview. The website with the highest combination of unique visitors and pageviews is ranked the highest
MAI is a dataset for multi-scene recognition in single aerial images. It consists of 3,923 labelled large-scale images from Google Earth imagery that covers the United States, Germany, and France. The size of each image is 512 ×512, and spatial resolutions vary from 0.3 m/pixel to 0.6 m/pixel. After capturing aerial images, multiple scene-level labels were manually assigned to each image from in total 24 scene categories, including apron, baseball, beach, commercial, farmland, woodland, parking lot, port, residential, river, storage tanks, sea, bridge, lake, park, roundabout, soccer field, stadium, train station, works, golf course, runway, sparse shrub, and tennis court
MLDS is a collection of thousands of trained neural networks labelled with the data used to train them. MLDS allows meta weight-space analysis across thousands of networks trained with identical or similar training data.
Continuous Defect Prediction (CDP) is a dataset of more than 11 million data rows, representing files involved in Continuous Integration (CI) builds, that synthesize the results of CI builds with data mined from software repositories. The dataset embraces 1,265 software projects, 30,022 distinct commit authors and several software process metrics that in earlier research appeared to be useful in software defect prediction. In this particular dataset the authors used TravisTorrent as the source of CI data. TravisTorrent synthesizes commit level information from the Travis CI server and GitHub open-source projects repositories.
This dataset of approximately 178,000 unique Dockerfiles collected from GitHub to facilitate sophisticated semantics-aware static analysis of Dockerfiles. To enhance the usability of this data, the authors use five representations for working with, mining from, and analyzing these Dockerfiles. Each Dockerfile representation builds upon the previous ones, and the final representation, created by three levels of nested parsing and abstraction, makes tasks such as mining and static checking tractable.
This dataset focuses on 50 articles about climate science, which were annotated completely by 49 students, 26 Upwork workers, 3 science and 3 journalism experts.
Acticipate is a publicly available dataset with recordings of human body-motion and eye-gaze, acquired in an experimental scenario with an actor interacting with three subjects. It contains synchronised and labelled video+gaze and body motion in a dyadic scenario of interaction.
APND (Arm Point Nav Dataset) is a dataset for the generalizable object manipulation task called ARMPOINTNAV, which consists on moving an object in the scene from a source location to a target location.
This dataset contains annotations for 5000 music files on the following music properties:
JoCAD is a dataset for anomaly detection in citation networks.
This is a dataset for benchmarking in-hand manipulation on different robot platforms.
SumeCzech-NER contains named entity annotations of SumeCzech 1.0, a Czech news-based summarization dataset.
AskUbuntu question dataset is a preprocessed collection of questions taken from the AskUbuntu.com 2014 corpus dump. It also comes with 400*20 manual annotations, marking pairs of questions as "similar" or "non-similar".