19,997 machine learning datasets
19,997 dataset results
Replication datasets (200 million rows) used in experiments by Yancey & Settles (2020). (2019-06-11)
SubSumE Dataset This repository contains the SubSumE dataset for subjective document summarization. See the paper and the talk for details on dataset creation. Also check out our work SuDocu on example-based document summarization.
AnswerSumm is a dataset of 4,631 CQA threads for answer summarization, curated by professional linguists.
It comprises synthetic mesh sequences from Deformation Transfer for Triangle Meshes.
Dataset contains cumulative reported cases, hospital admission and discharge, and mortality data as parsed from the publicly available press releases by the Ministry of Health and National Emergency Operations Centre (NEOC) of the Government of Samoa. The data spans the initial press release at the end of September 2019 through to the final press release at the end of January 2020.
Digital Edition: Essays from Hannah Arendt We have created a NER dataset from the digital edition "Sechs Essays" by Hannah Arendt. It consists of 23 documents from the period 1932-1976, which are available as TEI files online (see https://hannah-arendt-edition.net/3p.html?lang=de).
Digital Edition: Sturm Edition Source: Schrade, Torsten: „Startseite“, in: DER STURM. Digitale Quellenedition zur Geschichte der internationalen Avantgarde, erarbeitet und herausgegeben von Marjam Trautmann und Torsten Schrade. Mainz, Akademie der Wissenschaften und der Literatur, Version 1 vom 16. Jul. 2018.
The dataset describes 150 patients with the following demographic characteristics : sex, age, HOMA-IR , systolic and diastolic blood pressure , and LDL-Cholesterol . These patients were followed for 28 years . The characteristics are the mean of the following measurement each year. In each year a liver biopsy was taken to record the stage of fibrosis and then the count of transition among stages is recorded in the columns called ( lambda ij) where the i,j represents the stages where the patients move between them . In other word , for each patient , there is a column for the transition count this patient had made from stage 0 to stage 1 , then there is another column for the transition counts he made from stage 1 to stage 2 in this follow up 28 years , and so on , that is to mean there are 9 columns for the count of each transition for each patient as there are 9 transitions that could be made : from 0 to 1 , from 1 to 2 , from 2 to 3 , from 3 to 4 , from 1 to 0 , from 2 to 1 , from 3
DataCLUE is the first Data-Centric benchmark applied in NLP field.
We introduce a challenging new dataset for simultaneous object category and viewpoint classification—the Biased-Cars dataset. Our dataset features photo-realistic outdoor scene data with fine control over scene clutter (trees, street furniture, and pedestrians), car colors, object occlusions, diverse backgrounds (building/road textures) and lighting conditions (sky maps). Biased-Cars consists of 15K images of five different car models seen from viewpoints varying between 0-90 degrees of azimuth, and 0-50 degrees of zenith across multiple scales. Our dataset offers complete control over the joint distribution of categories, viewpoints, and other scene parameters, and the use of physically based rendering ensures photo-realism.
A high-resolution version of VGGFace2 for academic face editing purposes. This project uses GFPGAN for image restoration and insightface for data preprocessing (crop and align).
Hands Guns and Phones (HGP) dataset contains 2199 images (1989 for training an 210 for testing) of people using guns or phones in real-world scenarios (people making phones reviews, shooting drills, or making calls). Every image of this dataset is labeled with the bounding boxes of Hands, Phones and Guns. All the aforementioned images were collected from Youtube videos and have different sizes.
Temporal Hands Guns and Phones (THGP) dataset, is a collection of 5960 video frames (5000 for training and 960 for testing). The training part is composed with 50 videos of 100 frames (720 × 720 pixels). This dataset contains 20 videos of shooting drills, 20 videos of armed robberies, and 10 videos of people making calls. The testing part contains 48 videos of 20 frames (720 × 720). Videos contained in the testing dataset includes phone calls, gun reviews, shooting drills, people making calls, and armed robberies at convenience stores. This dataset is labeled with the bounding boxes of hands, phones, and guns.
Freely licensed dataset with warrants for 2k authentic arguments from news comments. On this basis, we present a new challenging task, the argument reasoning comprehension task. Given an argument with a claim and a premise, the goal is to choose the correct implicit warrant from two options. Both warrants are plausible and lexically close, but lead to contradicting claims.
WikiContradiction is a novel wiki dataset for self-contradiction Wikipedia article detection.
Product Page is a large-scale and realistic dataset of webpages. The dataset contains 51,701 manually labeled product pages from 8,175 real e-commerce websites. The pages can be rendered entirely in a web browser and are suitable for computer vision applications. This makes it substantially richer and more diverse than other datasets proposed for element representation learning, classification and prediction on the web.
Because there is no publicly available free dataset for speech dereverberation, we prepared a dataset based on the clean speech from VoiceBank-DEMAND [26] (discard the noisy speech) and convolved them with the room impulse response (RIR) from OpenSLR.
The ComMA Dataset v0.2 is a multilingual dataset annotated with a hierarchical, fine-grained tagset marking different types of aggression and the "context" in which they occur. The context, here, is defined by the conversational thread in which a specific comment occurs and also the "type" of discursive role that the comment is performing with respect to the previous comment. The initial dataset, being discussed here (and made available as part of the ComMA@ICON shared task), consists of a total 15,000 annotated comments in four languages - Meitei, Bangla, Hindi, and Indian English - collected from various social media platforms such as YouTube, Facebook, Twitter and Telegram. As is usual on social media websites, a large number of these comments are multilingual, mostly code-mixed with English.
Original dataset for "HIGH PRECISION MEDICINE BOTTLES VISION ONLINE INSPECTION SYSTEM AND CLASSIFICATION BASED ON MULTI-FEATURES AND ENSEMBLE LEARNING VIA INDEPENDENCE TEST"
533 parallel examples sampled from TACRED, translated into Russian and Korean (and 3 additional examples in Russian), accompanied with tranlsation of a list of trigger words collected for the different relations.