19,997 machine learning datasets
19,997 dataset results
CoRAL is a language and culturally aware Croatian Abusive dataset covering phenomena of implicitness and reliance on local and global context.
A maintained database tracks ICLR submissions and reviews, augmented with author profiles and higher-level textual features.
The dataset contains 179 photographs taken by a UAV flying at 120 meters altitude. The photographs are high resolution georeferenced (altitude and lontitude) orthoimages. It was originally used for developing a visual-based localization algorithm for UAVs. The dataset could be used for training machine learning models for localization purposes or building maps.
#chinahate dataset contains a total of 2,172,333 tweets hashtagged #china posted during the time it was collected. It is designed for the task of hate speech detection.
The BWB corpus consists of Chinese novels translated by experts into English, and the annotated test set is designed to probe the ability of machine translation systems to model various discourse phenomena.
DTBM is a benchmark dataset for Digital Twins that reflects these characteristics and look into the scaling challenges of different knowledge graph technologies.
Thorsten-Voice (Thorsten-21.02-neutral) is a neutrally spoken voice dataset recorded by Thorsten Müller, audio optimized by Dominik Kreutz and licenced under CC0 to provide it for anybody without any financial or licence struggle. It is intended to be used for speech synthesis in German as a single speaker dataset. It contains about 23 hours of high quality audio
VoxForge is an open speech dataset that was set up to collect transcribed speech for use with Free and Open Source Speech Recognition Engines (on Linux, Windows and Mac).
The M-AILABS Speech Dataset is the first large dataset that we are providing free-of-charge, freely usable as training data for speech recognition and speech synthesis. Most of the data is based on LibriVox and Project Gutenberg. The training data consist of nearly thousand hours of audio and the text-files in prepared format. A transcription is provided for each clip. Clips vary in length from 1 to 20 seconds and have a total length of approximately shown in the list (and in the respective info.txt-files) below. The texts were published between 1884 and 1964, and are in the public domain. The audio was recorded by the LibriVox project and is also in the public domain
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Unique radiogenomic dataset from a Non-Small Cell Lung Cancer (NSCLC) cohort of 211 subjects. The dataset comprises Computed Tomography (CT), Positron Emission Tomography (PET)/CT images, semantic annotations of the tumors as observed on the medical images using a controlled vocabulary, segmentation maps of tumors in the CT scans, and quantitative values obtained from the PET/CT scans. Imaging data are also paired with gene mutation, RNA sequencing data from samples of surgically excised tumor tissue, and clinical data, including survival outcomes.
Harry Potter Dialogue is the first dialogue dataset that integrates with scene, attributes and relations which are dynamically changed as the storyline goes on. Our work can facilitate research to construct more human-like conversational systems in practice. For example, virtual assistant, NPC in games, etc. Moreover, HPD can both support dialogue generation and retrieval tasks.
"Identifying Sensitive URLs at Web-Scale" dataset at IMC20
The Ambiguous VQA dataset is a dataset of ambiguous questions about images. It consists of a set of ambiguous images and their answers. It is used to train and evaluate question generation models in English.
Marine Microalgae Detection in Microscopy Images dataset contains a total number of images in the dataset is 937 and all the objects in these images were annotated. The total number of annotated objects is 4201. The training set contains 537 images and the testing set contains 430 images.
Hinglish-TOP is a human annotated code-switched semantic parsing dataset containing 10k human annotations for Hindi-English (HINGLISH) code switched utterances, and over 170K CST5 generated code-switched utterances from the TOPv2 dataset.
Discovering Interacted Objects (DIO) is a benchmark containing 51 interactions and 1,000+ objects designed for Spatio-temporal Human-Object Interaction (ST-HOI) detection.
CVE stands for Common Vulnerabilities and Exposures. CVE is a glossary that classifies vulnerabilities. The glossary analyzes vulnerabilities and then uses the Common Vulnerability Scoring System (CVSS) to evaluate the threat level of a vulnerability. A CVE score is often used for prioritizing the security of vulnerabilities.
The dataset contains standard contexts of the lattices of all atomic lattices in the Concept Explorer format.
This dataset contains stereo images, depth data, camera position/orientation, and camera information for 100 sorghum panicles (the seed-bearing head of the sorghum stalk), as well as semantic segmentation labels for a subset of the data. The 100 sampled sorghum stalks are drawn from 10 different species, in groups of 10. For image capture an illumination-invariant flash camera developed at CMU was swept around the panicle using a UR5 robotic arm, and approximately 150 image pairs were captured for each panicle, giving a full 3D view of each stalk.