TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Accidental Turntables

Accidental Turntables contains a challenging set of 41,212 images of cars in cluttered backgrounds, motion blur and illumination changes that serves as a benchmark for 3D pose estimation.

1 papers0 benchmarksImages, Videos

ScanEnts3D

Scan Entities in 3D (ScanEnts3D) is a large-scale dataset which provides explicit correspondences between 369k objects across 84k natural referentural sentences, covering 705 real-world scenes.

1 papers0 benchmarks3D, Texts

Selection from FFHQ & StyleGAN2:FFHQ (used in "Testing Human Ability To Detect Deepfake Images of Human Faces" study)

This dataset is the image stimulus pool of 50 deepfake and 50 real images, used for the experiment in the study titled "Testing Human Ability To Detect Deepfake Images of Human Faces".

1 papers0 benchmarksImages

Text2shape

A large dataset of natural language descriptions for physical 3D objects in the ShapeNet dataset.

1 papers0 benchmarks

THU-FVFDT (Tsinghua University Finger Vein and Finger Dorsal Texture Database)

THU-FVFDT is a dataset containing raw finger vein and finger dorsal texture images of 220 different subjects. Images are captured in two different sessions with interval of about dozens of seconds. One session is for training and the other for testing. Four finger vein images and four finger dorsal texture images are captured simultaneously in each session. We only offer one of four images in that there is approximately no difference between them. The size of raw images is 720×576 pixels.

1 papers0 benchmarks

2D_NACA_RANS

Dataset of low fidelity resolutions of the RANS equations over airfoils.

1 papers0 benchmarksGraphs, Physics, Point cloud

MiST

MiST (Modals In Scientific Text) is a dataset containing 3737 modal instances in five scientific domains annotated for their semantic, pragmatic, or rhetorical function.

1 papers0 benchmarksTexts

OIVIO (A Benchmark for Visual-Inertial Odometry Systems Employing Onboard Illumination)

It consists of 36 sequences, recorded in mines, tunnels, and other dark environments, totaling more than 145 minutes of stereo camera video and IMU data. In each sequence, the scene is illuminated by an onboard light of approximately 1350, 4500, or 9000 lumens. We accommodate both direct and indirect VIO methods by providing the geometric and photometric camera calibrations. The full dataset, including sensor data, calibration sequences, and evaluation scripts can be downloaded here.

1 papers0 benchmarks

UMA-VI Dataset (The UMA-VI dataset: Visual--inertial odometry in low-textured and dynamic illumination environments)

The dataset contains 32 sequences for the evaluation of VI motion estimation methods, totalling ∼80 min of data. The dataset covers challenging conditions (mainly illumination changes and low textured environments) in different degrees and a wide rage of scenarios (including corridors, parking, classrooms, halls, etc.) from two different buildings at the University of Malaga. In general, we provide at least two different sequences within the same scenario, with different illumination conditions or following different trajectories. All sequences were recorded with our VI sensor handheld, except a few that were recorded while mounted in a car.

1 papers0 benchmarks

AIROGS (Rotterdam EyePACS AIROGS)

The Rotterdam EyePACS AIROGS dataset (in full, so including train and test) contains 113,893 color fundus images from 60,357 subjects and approximately 500 different sites with a heterogeneous ethnicity.

1 papers0 benchmarksImages, Medical

CRCDX (TCGA-CRC-DX)

Histological images of colorectal cancer, derived from the TCGA database

1 papers0 benchmarksImages, Medical

FreCDo (French cross-domain)

FreCDo is a corpus for French dialect identification comprising 413,522 French text samples collected from public news websites in Belgium, Canada, France and Switzerland.

1 papers0 benchmarksTexts

multiRAW

To encourage reproducible research, a labeled MultiRAW dataset containing>7k RAW images acquired using multiple camera sensors is made publicly accessible for RAW-domain processing.

1 papers0 benchmarksImages

Robust Summarization Evaluation Benchmark

Robust Summarization Evaluation Benchmark is a large human evaluation dataset consisting of over 22k summary-level annotations over state-of-the-art systems on three datasets.

1 papers0 benchmarksTexts

TBBR Raw (Hyperspectral (RGB + Thermal) drone images of Karlsruhe, Germany)

This dataset contains the raw images for the dataset of Thermal Bridges on Building Rooftops (TBBR) dataset.

1 papers0 benchmarksHyperspectral images

MuReD Dataset (Multi-Label Retinal Diseases Dataset)

Early detection of retinal diseases is one of the most important means of preventing partial or permanent blindness in patients. One of the major stumbling blocks for manual retinal examination is the lack of a sufficient number of qualified medical personnel per capita to diagnose diseases. Computer-aided diagnosis systems (CAD) have proven to be very effective in helping physicians reduce the time taken to make a diagnosis and minimize variability in image interpretation. Still, they are not flexible enough to accommodate the simultaneous presence of multiple retinal diseases, which is a common situation in real-world applications. In the past years, few datasets that focus on the classification of numerous retinal pathologies present at the same time, i.e., multi-label classification have been proposed, but there are some shared problems with all of them, such as a narrow range of pathologies to classify, high level of class imbalance, low amount of samples for the underrepresented

1 papers3 benchmarks

FETA Car-Manuals (FETA Car-Manuals dataset, image-text retrieval for foundation models' expert data performance.)

FETA benchmark focuses on text-to-image and image-to-text retrieval in public car manuals and sales catalogue brochures. The FETA Car-Manuals dataset consists of a total of 349 PDF documents from 5 car manufacturers, namely Nissan, Toyota, Mazda, Renault, Chevrolet.

1 papers6 benchmarksImages, Texts

FETA IKEA

FETA benchmark focuses on text-to-image and image-to-text retrieval in public car manuals and sales catalogue brochures. The FETA IKEA dataset contains 26 documents with 7366 pages total, approximately 9574 images and 23927 texts automatically extracted from those pages.

1 papers0 benchmarksImages, Texts

Verifee

Verifee is a dataset of news articles with fine-grained trustworthiness annotations. It contains over 10, 000 unique articles from almost 60 Czech online news sources. These are categorized into one of the 4 classes across the credibility spectrum we propose, raging from entirely trustworthy articles all the way to the manipulative ones.

1 papers0 benchmarksTexts

Werewolf Among Us

Werewolf Among Us is a dataset multimodal dataset for modeling persuasion behaviors. It contains 199 dialogue transcriptions and videos captured in a multi-player social deduction game setting, 26,647 utterance level annotations of persuasion strategy, and game level annotations of deduction game outcomes.

1 papers0 benchmarksDialog, Videos
PreviousPage 450 of 1000Next