TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Pins Face Recognition

This images has been collected from Pinterest and cropped. There are 105 celebrities and 17534 faces.

1 papers0 benchmarks

NIH 3T3 microtubule cell dataset

The data consists of 21 images of microtubules in PFA-fixed NIH 3T3 mouse embryonic fibroblasts (DSMZ: ACC59) labeled with a mouse anti-alpha-tubulin monoclonal IgG1 antibody (Thermofisher A11126, primary antibody) and visualized by a blue-fluorescent Alexa Fluor 405 goat anti-mouse IgG antibody (Thermofisher A-31553, secondary antibody). Acquisition of the images was performed using a confocal microscope (Olympus IX81).

1 papers0 benchmarksImages

sim-combi

Data simulator for polypharmacies / drug combinations TL;DR python create_dataset.py [--config path/to/config.json --seed your_seed]

1 papers0 benchmarks

Accompnaying Dataset for: Chemical Heredity as Group Selection at the Molecular Level

Accompnaying Dataset for: Chemical Heredity as Group Selection at the Molecular Level. File descriptions are provided in the Appendix of [Markovitch, Witkowski and Virgo; Chemical Heredity as Group Selection at the Molecular Level, arXiv (2018)] (https://arxiv.org/abs/1802.08024).

1 papers0 benchmarks

AutoPoster dataset

Dataset proposed by ACM MM 2023 paper "AutoPoster: A Highly Automatic and Content-aware Design System for Advertising Poster Generation"

1 papers0 benchmarks

Voxceleb-3D

A dataset for voice and 3D face structure study. It contains about 1.4K identities with their 3D face models and voice data. 3D face models are fitted from VGGFace using BFM 3D models, and voice data are processed from Voxceleb

1 papers10 benchmarks3d meshes, Speech

USPTO-30K

We introduce USPTO-30K, a large-scale benchmark dataset of annotated molecule images, which overcomes these limitations. It is created using the pairs of images and MolFiles by the United States Patent and Trademark Office. Each molecule was independently selected among all the available documents from 2001 to 2020. The set consists of three subsets to decouple the study of clean molecules, molecules with abbreviations and large molecules.

1 papers0 benchmarksGraphs, Images

MolGrapher-Synthetic-300K

The set is created using molecule SMILES retrieved from the database PubChem. Images are then generated from SMILES using the molecule drawing library RDKit. The synthetic set is augmented at multiple levels:

1 papers0 benchmarksGraphs, Images

NeRF-MVL (object-centric multi-view LiDAR dataset)

We establish an object-centric multi-view LiDAR dataset, which we dub the NeRF-MVL dataset, containing carefully calibrated sensor poses, acquired from multi-LiDAR sensor data from real autonomous vehicles. It contains more than 76k frames covering two types of collecting vehicles, three LiDAR settings, two collecting paths, and nine object categories.

1 papers0 benchmarks

EchoNet LVH

Echocardiography, or cardiac ultrasound, is the most widely used and readily available imaging modality to assess cardiac function and structure. Combining portable instrumentation, rapid image acquisition, high temporal resolution, and without the risks of ionizing radiation, echocardiography is one of the most frequently utilized imaging studies in the United States and serves as the backbone of cardiovascular imaging. For diseases ranging from heart failure to valvular heart diseases, echocardiography is both necessary and sufficient to diagnose many cardiovascular diseases. In addition to our deep learning model, we introduce a new large video dataset of echocardiograms (parasternal long axis view) for computer vision research. The EchoNet-LVH dataset includes 12,000 labeled echocardiogram videos and human expert annotations (measurements, tracings, and calculations) to provide a baseline to study cardiac chamber size and wall thickness.

1 papers0 benchmarks

DEEP-VOICE: DeepFake Voice Recognition (Jordan Bird)

DEEP-VOICE: Real-time Detection of AI-Generated Speech for DeepFake Voice Conversion This dataset contains examples of real human speech, and DeepFake versions of those speeches by using Retrieval-based Voice Conversion.

1 papers2 benchmarks

TYC Dataset (The TYC Dataset for Understanding Instance-Level Semantics and Motions of Cells in Microstructures)

We introduce the trapped yeast cell (TYC) dataset, a novel dataset for understanding instance-level semantics and motions of cells in microstructures. We release $105$ dense annotated high-resolution brightfield microscopy images, including about $19$k instance masks. We also release $261$ curated video clips composed of $1293$ high-resolution microscopy images to facilitate unsupervised understanding of cell motions and morphology.

1 papers0 benchmarksImages, Videos

Pylon Benchmark (Pylon Table Union Search Benchmark)

We create a new dataset from GitTables, a data lake of 1.7M tables extracted from CSV files on GitHub. The benchmark comprises 1,746 tables including union-able table subsets under topics selected from Schema.org: scholarly article, job posting, and music playlist. We end up with these three topics since we can find a fair number of union-able tables of them from diverse sources in the corpus (we can easily find union-able tables from a single source but they are less interesting for table union search as simple syntactic methods can identify all of them because of the same schema and consistent value representations).

1 papers0 benchmarksTabular

Quechua-SER

Quechua Collao corpus for automatic emotion recognition in speech. Audios are provided, alongside csv files with labels from 4 annotators for valence, arousal, and dominance values, using a 1 to 5 scale.

1 papers4 benchmarksAudio, Speech

HardZiPA Dataset

The HardZiPA folder contains illuminance and RGB data as well as CO2 and TVOC data for five sensing devices.

1 papers0 benchmarks

SD7K (Shadow Document 7K)

SD7K is the only large-scale high-resolution dataset that satisfies all important data features about document shadow currently, which covers a large number of document shadow images. Mean resolution is $2462 \times 3699$

1 papers0 benchmarksImages

WikiFANE_Gold (Fine-grained Arabic Named Entity Corpora)

The gold-standard and automatically-developed fine-grained Arabic named entity corpora are resources created by annotating Named Entities into 50 fine-grained classes.

1 papers0 benchmarks

MIMIC-GAZE-JPG

1083 cases from the MIMIC-CXR dataset. For each case, a gray-scaled X-ray image with the size of around 3000x3000, eye-gaze data, and ground-truth classification labels are provided. These cases are classified into 3 categories: Normal, Congestive Heart Failure (CHF), and Pneumonia.

1 papers0 benchmarks

Taobao (TGN Style)

Taobao dataset which is pre-processed in TGN Style.

1 papers0 benchmarks

ML25m (TGN Style)

ML25m dataset which is pre-processed in TGN Style.

1 papers0 benchmarks
PreviousPage 470 of 1000Next