TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

MARS (Multimodal Analogical Reasoning dataSet)

Analogical reasoning is fundamental to human cognition and holds an important place in various fields. However, previous studies mainly focus on single-modal analogical reasoning and ignore taking advantage of structure knowledge. We introduce the new task of multimodal analogical reasoning over knowledge graphs, which requires multimodal reasoning ability with the help of background knowledge. Our dataset MARS contains 10,685 training, 1,228 validation and 1,415 test instances.

1 papers4 benchmarksImages

LiDAR-CS

LiDAR-CS is a dataset for 3D object detection in real traffic. It contains 84,000 point cloud frames under 6 groups of different sensors but with same corresponding scenarios, captured from hybrid realistic LivDAR simulator.

1 papers0 benchmarksLiDAR

Coffereview Dataset

The data set is based on roughly 6,000 coffee bean review spublished on the website Coffereviews going back to 1997. All of these reviews are scored with the q-grading scale (Coffee Review, 2021)

1 papers0 benchmarks

Processed CMIP5 EWS Data

Processed CMIP5 data used for testing the CNN-LSTM model. Details in Zenodo description

1 papers0 benchmarks

Fraunhofer EZRT XXL-CT Instance Segmentation Me163

The 'Me 163' was a Second World War fighter airplane and a result of the German air force secret developments. One of these airplanes is currently owned and displayed in the historic aircraft exhibition of the 'Deutsches Museum' in Munich, Germany. To gain insights with respect to its history, design and state of preservation, a complete CT scan was obtained using an industrial XXL-computer tomography scanner at Fraunhofer EZRT .

1 papers0 benchmarks3D, Images

Motion Capture Data for Hand Motion Embodiment

Motion Capture Data for Hand Motion Embodiment contains demonstrations of different hand motion recorded with the Qualisys MOCAP system.

1 papers0 benchmarks

DigiCall (DigiCall: Earning Calls Dataset)

We release 3.691 earning call transcripts and also annotated data set, labeled particularly for the digital strategy maturity by linguists. https://github.com/hpataci/DigiCall

1 papers0 benchmarks

Rice Grains BRRI (Rice Grain Disease Dataset for False Smut and Neck Blast)

dataset (balanced) of 200 images consists of three classes - False Smut, Neck Blast and healthy grain class. Some of these images contain both diseases together. Field data collected under the supervision of staff from the Bangladesh Rice Research Institute (BRRI).

1 papers0 benchmarks

Fontenay Dataset

This data set encompasses 104 images and transcriptions of digital images of original charters from the Cistercian abbey Fontenay in Burgundy (France), dating mainly from the 12th c. and until 1213. The original data set was created as part of the ANR ORIFLAMMS (ANR-12-CORP-0010) project. Texts were transcribed in the original TEI-XML format, rendering both abbreviated and expanded forms of the original text. The alignment data was produced by merging coordinates created through the Oriflamms software. A new version was prepared in March-June 2022 as part of the research for the following paper: Camps, Jean-Baptiste, Chahan Vidal-Gorène, Dominique Stutzmann, Marguerite Vernet, and Ariane Pinche. « Data Diversity in Handwritten Text Recognition: Challenge or Opportunity? » In Digital Humanities 2022. Conference Abstracts (The University of Tokyo, Japan, 25-29 July 2022), published by DH2022 Local Organizing Committee, 160‑65. Tokyo, 2022.

1 papers0 benchmarks

MaxwellBlobs

The dataset consists of random electromagnetic scatterers and their associated fields when illuminated by a 1000nm plane-wave illumination. The scatterers have a refractive index of $n=1.5$ and are surrounded by air ($n=1.0$), and the side-length of the simulated area is 5.12 microns.

1 papers0 benchmarks

MotionID: IMU specific motion (User verification)

Dataset for User Verification part of MotionID: Human Authentication Approach. Data type: bin (should be converted by attached notebook). ~50 hours of IMU (Inertial Measurement Units) data for one specific motion pattern, provided by 101 users.

1 papers0 benchmarksActions, Time series

MotionID: IMU all motions part1 (Motion Patterns Identification)

Dataset (part 1/3) for Motion Patterns Identification part of MotionID: Human Authentication Approach. Data type: bin (should be converted by attached notebook).

1 papers0 benchmarksActions, Time series

MotionID: IMU all motions part2 (Motion Patterns Identification)

Dataset (part 2/3) for Motion Patterns Identification part of MotionID: Human Authentication Approach. Data type: bin (should be converted by attached notebook).

1 papers0 benchmarksActions, Time series

MotionID: IMU all motions part3 (Motion Patterns Identification)

Dataset (part 3/3) for Motion Patterns Identification part of MotionID: Human Authentication Approach. Data type: bin (should be converted by attached notebook).

1 papers0 benchmarksActions, Time series

IISc VEED (Indian Institute of Science Virtual Environment Exploration Database of Static Scenes)

IISc VEED consists of 200 diverse indoor and outdoor scenes (see samples below). The videos are rendered with blender and the blend files are obtained for the scenes mainly from blendswap and turbosquid. 4 different camera trajectories are added to each scene and thus we have a total of 800 videos. The videos are rendered at full HD resolution (1920 x 1080) and at 30fps and contain 12 frames each.

1 papers0 benchmarks

WUSTL_EHMS_2020 (WUSTL EHMS 2020 Dataset for Internet of Medical Things (IoMT) Cybersecurity Research)

The WUSTL-EHMS-2020 dataset was created using a real-time Enhanced Healthcare Monitoring System (EHMS) testbed [1]. This testbed collects both the network flow metrics and patients' biometrics due to the scarcity of a dataset that combines these biometrics.

1 papers0 benchmarks

APRICOT-Mask

We present the APRICOT-Mask dataset, which augments the APRICOT dataset with pixel-level annotations of adversarial patches. We hope APRICOT-Mask along with the APRICOT dataset can facilitate the research in building defenses against physical patch attacks, especially patch detection and removal techniques.

1 papers0 benchmarks

DocOIE

We manually annotate 800 sentences from 80 documents in two domains (Healthcare and Transportation) to form a DocOIE dataset for evaluation.

1 papers0 benchmarksTexts

Dataset for MPLP

This dataset is used for MPLP considering time windows constraints of customers and parking space. To randomly generate the dataset, please visit the link: https://github.com/Yubin-Liu/Hybrid-Q-Learning-Network-Approach-for-MPLP.

1 papers0 benchmarks

OSASUD (OSASUD: A dataset of stroke unit recordings for the detection of Obstructive Sleep Apnea Syndrome)

Polysomnography (PSG) is a fundamental diagnostical method for the detection of Obstructive Sleep Apnea Syndrome (OSAS). Historically, trained physicians have been manually identifying OSAS episodes in individuals based on PSG recordings. Such a task is highly important for stroke patients, since in such cases OSAS is linked to higher mortality and worse neurological deficits. Unfortunately, the number of strokes per day vastly outnumbers the availability of polysomnographs and dedicated healthcare professionals. The data in this work pertains to 30 patients that were admitted to the stroke unit of the Udine University Hospital, Italy. Unlike previous studies, exclusion criteria are minimal. As a result, data are strongly affected by noise, and individuals may suffer from several comorbidities. Each patient instance is composed of overnight vital signs data deriving from multi-channel ECG, photoplethysmography and polysomnography, and related domain expert’s OSAS annotations. The datas

1 papers0 benchmarks
PreviousPage 454 of 1000Next