TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

CoRAL dataset (CoRAL: a Context-aware Croatian Abusive Language Dataset)

CoRAL is a language and culturally aware Croatian Abusive dataset covering phenomena of implicitness and reliance on local and global context.

1 papers0 benchmarksTexts

ICLR Database (ICLR Database (with Textual Covariates))

A maintained database tracks ICLR submissions and reviews, augmented with author profiles and higher-level textual features.

1 papers0 benchmarksTabular, Texts

High Altitude Georeferenced UAV Images

The dataset contains 179 photographs taken by a UAV flying at 120 meters altitude. The photographs are high resolution georeferenced (altitude and lontitude) orthoimages. It was originally used for developing a visual-based localization algorithm for UAVs. The dataset could be used for training machine learning models for localization purposes or building maps.

1 papers0 benchmarks

#chinahate

#chinahate dataset contains a total of 2,172,333 tweets hashtagged #china posted during the time it was collected. It is designed for the task of hate speech detection.

1 papers0 benchmarksTexts

BWB

The BWB corpus consists of Chinese novels translated by experts into English, and the annotated test set is designed to probe the ability of machine translation systems to model various discourse phenomena.

1 papers0 benchmarksTexts

DTBM (Digital Twin Benchmark Model)

DTBM is a benchmark dataset for Digital Twins that reflects these characteristics and look into the scaling challenges of different knowledge graph technologies.

1 papers0 benchmarks

Thorsten voice 21.02 neutral

Thorsten-Voice (Thorsten-21.02-neutral) is a neutrally spoken voice dataset recorded by Thorsten Müller, audio optimized by Dominik Kreutz and licenced under CC0 to provide it for anybody without any financial or licence struggle. It is intended to be used for speech synthesis in German as a single speaker dataset. It contains about 23 hours of high quality audio

1 papers1 benchmarksAudio

Voxforge German

VoxForge is an open speech dataset that was set up to collect transcribed speech for use with Free and Open Source Speech Recognition Engines (on Linux, Windows and Mac).

1 papers2 benchmarksAudio

M-AILabs speech dataset

The M-AILABS Speech Dataset is the first large dataset that we are providing free-of-charge, freely usable as training data for speech recognition and speech synthesis. Most of the data is based on LibriVox and Project Gutenberg. The training data consist of nearly thousand hours of audio and the text-files in prepared format. A transcription is provided for each clip. Clips vary in length from 1 to 20 seconds and have a total length of approximately shown in the list (and in the respective info.txt-files) below. The texts were published between 1884 and 1964, and are in the public domain. The audio was recorded by the LibriVox project and is also in the public domain

1 papers2 benchmarksAudio

KGRED (Knowledge-graph-enhanced relation extraction datasets--)

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksTexts

NSCLC-Radiogenomics

Unique radiogenomic dataset from a Non-Small Cell Lung Cancer (NSCLC) cohort of 211 subjects. The dataset comprises Computed Tomography (CT), Positron Emission Tomography (PET)/CT images, semantic annotations of the tumors as observed on the medical images using a controlled vocabulary, segmentation maps of tumors in the CT scans, and quantitative values obtained from the PET/CT scans. Imaging data are also paired with gene mutation, RNA sequencing data from samples of surgically excised tumor tissue, and clinical data, including survival outcomes.

1 papers0 benchmarksImages, Medical

Harry Potter Dialogue Dataset

Harry Potter Dialogue is the first dialogue dataset that integrates with scene, attributes and relations which are dynamically changed as the storyline goes on. Our work can facilitate research to construct more human-like conversational systems in practice. For example, virtual assistant, NPC in games, etc. Moreover, HPD can both support dialogue generation and retrieval tasks.

1 papers5 benchmarks

SURL (IMC 2020 Curlie URL Dataset)

"Identifying Sensitive URLs at Web-Scale" dataset at IMC20

1 papers0 benchmarks

Ambiguous VQA

The Ambiguous VQA dataset is a dataset of ambiguous questions about images. It consists of a set of ambiguous images and their answers. It is used to train and evaluate question generation models in English.

1 papers0 benchmarksImages, Texts

Marine Microalgae Detection in Microscopy Images

Marine Microalgae Detection in Microscopy Images dataset contains a total number of images in the dataset is 937 and all the objects in these images were annotated. The total number of annotated objects is 4201. The training set contains 537 images and the testing set contains 430 images.

1 papers0 benchmarksBiology, Images

Hinglish-TOP

Hinglish-TOP is a human annotated code-switched semantic parsing dataset containing 10k human annotations for Hindi-English (HINGLISH) code switched utterances, and over 170K CST5 generated code-switched utterances from the TOPv2 dataset.

1 papers0 benchmarksTexts

DIO (Discovering Interacted Objects)

Discovering Interacted Objects (DIO) is a benchmark containing 51 interactions and 1,000+ objects designed for Spatio-temporal Human-Object Interaction (ST-HOI) detection.

1 papers0 benchmarksVideos

CVE (Common Vulnerabilities and Exposures)

CVE stands for Common Vulnerabilities and Exposures. CVE is a glossary that classifies vulnerabilities. The glossary analyzes vulnerabilities and then uses the Common Vulnerability Scoring System (CVSS) to evaluate the threat level of a vulnerability. A CVE score is often used for prioritizing the security of vulnerabilities.

1 papers0 benchmarksTexts

standard atomic contexts (standard contexts for the lattices of atomic lattices)

The dataset contains standard contexts of the lattices of all atomic lattices in the Concept Explorer format.

1 papers0 benchmarksTabular

High-Resolution Stereo Scans of 100 Sorghum Panicles

This dataset contains stereo images, depth data, camera position/orientation, and camera information for 100 sorghum panicles (the seed-bearing head of the sorghum stalk), as well as semantic segmentation labels for a subset of the data. The 100 sampled sorghum stalks are drawn from 10 different species, in groups of 10. For image capture an illumination-invariant flash camera developed at CMU was swept around the panicle using a UR5 robotic arm, and approximately 150 image pairs were captured for each panicle, giving a full 3D view of each stalk.

1 papers0 benchmarks
PreviousPage 444 of 1000Next