TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Flight Scheduling Data

Dataset was introduced by Jones Granatyr in his book https://iaexpert.academy/2016/10/25/review-de-livro-programando-a-inteligencia-coletiva where he scraped flight schedules.

1 papers0 benchmarksTabular

GeBiD (Geometric shapes Bimodal Dataset)

We provide a custom synthetic bimodal dataset, called GeBiD, designed specifically for the comparison of the joint- and cross-generative capabilities of Multimodal Variational Autoencoders. It comprises RGB images of geometric primitives and textual descriptions. The dataset offers 5 levels of difficulty (based on the number of attributes) to find the minimal functioning scenario for each model. Moreover, its rigid structure enables automatic qualitative evaluation of the generated samples.

1 papers0 benchmarksImages, Texts

Eurovision 2018 votes

Eurovision 2018 official votes dataset. more details can be found here: https://towardsdatascience.com/social-network-analysis-from-theory-to-applications-with-python-d12e9a34c2c7

1 papers0 benchmarks

SoccerTrack Dataset

The SoccerTrack dataset comprises top-view and wide-view video footage annotated with bounding boxes. GNSS coordinates of each player are also provided. We hope that the SoccerTrack dataset will help advance the state of the art in multi-object tracking, especially in team sports.

1 papers0 benchmarksRGB Video, Tracking, Videos

DistNLI

This dataset is named as the DistNLI dataset, which is a synthesized benchmark aiming to probe neural network models from the aspect of conjunctions on distributivity in NLI task in American English. DistNLI consists of sentence minimal pairs (premise and hypothesis) differentiated with conjunction structure within the pair and distributivity-related linguistic phenomenon. DistNLI is compiled with 328 sentences so far (164 for distributive and 164 for ambiguous predicates), annotated by 4 proficient English speakers with a background in NLP and Linguistics. Due to the specificity of the linguistic phenomenon involved and its size, this DistNLI dataset should only be used as an adversarial dataset in the investigation of distributivity of verb predication.

1 papers0 benchmarksTexts

SELTO Dataset

A Benchmark Dataset for Deep Learning-based Methods for 3D Topology Optimization.

1 papers0 benchmarks3D, 3d meshes

Eth-ICO (ICO-wallets in Ethereum)

The sampled 2-hop subgraphs centered on ICO-wallet accounts on the Ethereum Interaction graph.

1 papers0 benchmarks

Eth-Mining (Mining in Ethereum)

The sampled 2-hop subgraphs centered on Mining accounts on the Ethereum Interaction graph.

1 papers0 benchmarks

EOSIO-Robot (Robot account in EOSIO)

The sampled 2-hop subgraphs centered on Robot accounts on the EOSIO Interaction graph.

1 papers0 benchmarks

Censored_Planet_Quack (Censored Planet HyperQuack Echo)

Hyperquack v.2 response data which contains structured data records in JSON.

1 papers0 benchmarksTime series

DMAD (Deepfake Massively Annotated Databases)

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

Regressors-Regressions Dataset

This dataset is a collection of 5348 links from bug-introducing and bug-fixing commit sets extracted from Mozilla's Bugzilla with the use of bugbug. In this repository, you will find two shapes of it:

1 papers0 benchmarksTexts

JundeCOFG

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

DCASE 2021 Task3

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

nuScenes (Cross-City UDA)

A cross-city UDA benchmark built upon nuScenes.

1 papers0 benchmarks3D, LiDAR, Point cloud

Labeled data for citation field extraction

Citations are an important part of scientific papers, and the proper handling of them is indispensable for the science of science. Citation field extraction is the task of parsing citations: given a citation string, extract authors, title, venue, doi etc. Since the number of citations is counted by hundreds millions, efficient computer based methods for this task are very important.

1 papers0 benchmarks

doges-dogaresse (Doges and dogaresse of the Venetian Republic)

This is the list of all doges of the Venetian Republic, as well as their wives, if there's a record that they existed. They include name, surname if known, and date of their office, as well as the date of their weddings. Data has been extracted from the Wikipedia, with some errors fixed checking against other sources.

1 papers0 benchmarksGraphs

SKINL2 (Light Field Image Dataset of Skin Lesions)

The SKINL2 dataset comprises a total of 376 light fields acquired under similar conditions. The images were classified using eight categories, according to the type of skin lesion/ICD code:

1 papers0 benchmarks

7-point criteria evaluation Database

"We provide a database for evaluating computerized image-based prediction of the 7-point skin lesion malignancy checklist. The dataset includes over 2000 clinical and dermoscopy color images, along with corresponding structured metadata tailored for training and evaluating computer aided diagnosis (CAD) systems. "

1 papers0 benchmarks

CZ Software Mentions dataset, a new dataset of software mentions in biomedical papers.

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks
PreviousPage 437 of 1000Next