TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

VR Curve on Surface Drawing Dataset

The datasets includes curves drawn on 3D surfaces (triangle meshes) in Virtual Reality. A total of 2,880 curves were created using two different techniques by 20 users on 6 meshes. For each curve, a 3D curve executed by the user is provided, the projected curve created on the mesh, and the ground truth target curve on the mesh. For collecting the data, two different task types were employed, which are described in the paper.

1 papers0 benchmarks3D, Interactive

Machine Learning Quantum Reaction Rate Constants (Evan Komp and Stephanie Valleau)

Dataset of 1,517,419 quantum reaction rate constant products kQM(T)QR(T) computed from the transmission coefficient for model single and double barrier minimum energy paths. Here kQM(T) is the quantum reaction rate constant at temperature T and QR(T) is the reactant partition function computed with the rigid rotor and harmonic oscillator approximations.This dataset was created for Ref [1] where it was used to train and test a DNN to predict logkQM(T)QR(T).

1 papers0 benchmarks

Novel COVID-19 Chestxray Repository

Authors of the Dataset:

1 papers1 benchmarksImages, Medical

TVRecap

TVRecap a story generation dataset that requires generating detailed TV show episode recaps from a brief summary and a set of documents describing the characters involved. Unlike other story generation datasets, TVRecap contains stories that are authored by professional screenwriters and that feature complex interactions among multiple characters. Generating stories in TVRecap requires drawing relevant information from the lengthy provided documents about characters based on the brief summary. In addition, by swapping the input and output, TVRecap can serve as a challenging testbed for abstractive summarization.

1 papers0 benchmarksTexts

HYouTube

HYouTube is a video for Video harmonization, which aims to adjust the foreground of a composite video to make it compatible with the background. The dataset was created by adjusting the foreground of real videos to create synthetic composite videos. It is based on Youtube-VOS

1 papers0 benchmarksVideos

EUEN17037_Daylight_and_View_Standard_TestDataSet

EUEN17037 Daylight and View Standard Test Dataset.

1 papers0 benchmarks3D, Point cloud, Tabular

CI-ToD

CI-ToD is a dataset for Consistency Identification in Task-oriented Dialog system.

1 papers0 benchmarksTexts

METEOR

METEOR is a complex traffic dataset which captures traffic patterns in unstructured scenarios in India. METEOR consists of more than 1000 one-minute video clips, over 2 million annotated frames with ego-vehicle trajectories, and more than 13 million bounding boxes for surrounding vehicles or traffic agents. METEOR is a unique dataset in terms of capturing the heterogeneity of microscopic and macroscopic traffic characteristics.

1 papers0 benchmarksVideos

Diagnosis of COVID-19 and its clinical spectrum

This dataset contains anonymized data from patients seen at the Hospital Israelita Albert Einstein, at São Paulo, Brazil, and who had samples collected to perform the SARS-CoV-2 RT-PCR and additional laboratory tests during a visit to the hospital.

1 papers0 benchmarks

WHPA (Wellhead Protection Area prediction from breakthrough curves)

This dataset was created as part of the following study, which was published in the Journal of Hydrology: A new framework for experimental design using Bayesian Evidential Learning: the case of wellhead protection area https://doi.org/10.1016/j.jhydrol.2021.126903. The pre-print is available on arXiv: https://arxiv.org/pdf/2105.05539.pdf

1 papers0 benchmarks

IECSIL FIRE-2018 Shared Task

The dataset is taken from the First shared task on Information Extractor for Conversational Systems in Indian Languages (IECSIL) . It consists of 15,48,570 Hindi words in Devanagari script and corresponding NER labels. Each sentence end is marked by \newline" tag. Fig. 1 shows a snapshot of one sentence in the dataset. Our Dataset has nine classes, namely, Datenum, Event, Location, Name, Number, Occupation, Organization, Other, Things.

1 papers1 benchmarks

ChMusic

ChMusic is a traditional Chinese music dataset for training model and performance evaluation of musical instrument recognition. This dataset cover 11 musical instruments, consisting of Erhu, Pipa, Sanxian, Dizi, Suona, Zhuiqin, Zhongruan, Liuqin, Guzheng, Yangqin and Sheng.

1 papers0 benchmarksAudio, Music

TFRD (Temperature Field Reconstruction Dataset)

TFRD is a dataset to evaluate machine learning modelling methods for theTemperature field reconstruction of heat source systems (TFR-HSS).

1 papers0 benchmarks

CAT (Context Adjustment Training)

CAT is a specialized dataset for co-saliency detection - one of the core tasks in the field of computer vision. This dataset is intended for both helping to assess the performance of vision algorithms and supporting research that aims to exploit large volumes of annotated data, e.g., for training deep neural networks. CAT consists of 33,500 images

1 papers0 benchmarksImages

FewGLUE_64_labeled (A new version of FewGLUE with 64 training examples)

Introduction The FewGLUE_64_labeled dataset is a new version of FewGLUE dataset. It contains a 64-sample training set, a development set (the original SuperGLUE development set), a test set, and an unlabeled set. It is constructed to facilitate the research of few-shot learning for natural language understanding tasks.

1 papers0 benchmarksTexts

VQA-MHUG

VQA-MHUG is a 49-participant dataset of multimodal human gaze on both images and questions during visual question answering (VQA) collected using a high-speed eye tracker.

1 papers0 benchmarksImages, Texts

Lincolnbeet

The Lincolnbeet dataset is an object detection dataset designed to encourage research in the identification of items in environments with high levels of occlusion, and in the development of better approaches to evaluate object detection models in practical scenarios. This dataset was introduced in the paper: "Towards practical object detection for weed spraying in precision agriculture".

1 papers0 benchmarksImages

JDDC 2.0

JDDC 2.0 is a large-scale multimodal multi-turn dialogue dataset collected from a mainstream Chinese E-commerce platform JD.com, containing about 246 thousand dialogue sessions, 3 million utterances, and 507 thousand images, along with product knowledge bases and image category annotations. The dataset is divided into the training set, the validation set, and the test set according to the ratio of 80%, 10%, and 10%.

1 papers0 benchmarksTexts

StoryDB

StoryDB is a broad multi-language dataset of narratives. StoryDB is a corpus of texts that includes stories in 42 different languages. Every language includes 500+ stories. Some of the languages include more than 20 000 stories. Every story is indexed across languages and labeled with tags such as a genre or a topic. The corpus shows rich topical and language variation and can serve as a resource for the study of the role of narrative in natural language processing across various languages including low resource ones.

1 papers0 benchmarks

SCIMAT

SCIMAT is a large question-answer dataset for mathematics and science problems; such dataset can have impact on online education, intelligent tutoring and automated grading.

1 papers0 benchmarksTexts
PreviousPage 407 of 1000Next