TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

RSP Dataset

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksImages

WebLI (Web Language Image)

WebLI (Web Language Image) is a web-scale multilingual image-text dataset, designed to support Google’s vision-language research, such as the large-scale pre-training for image understanding, image captioning, visual question answering, object detection etc.

1 papers0 benchmarks

EHR Dataset for Patient Treatment Classification

The dataset is Electronic Health Record Predicting collected from a private Hospital in Indonesia. It contains the patients laboratory test results used to determine next patient treatment whether in care or out care patient. The task embedded to the dataset is classification prediction.

1 papers0 benchmarks

Extended MP-16 Dataset (EMP-16)

To overcome the need for a full installation of a reverse geocoder such as Nominatim, we provide the post-processed output of the reverse geocoding for the MP-16 dataset along with the validation set (YFCC-Val26k) which originally comprised photos and respective GPS coordinates. Both datasets are subsets of the YFCC100M dataset which are crawled from Flickr.

1 papers0 benchmarks

Events in Invasion Games Dataset - Handball (EIGD-H)

This dataset contains the broadcast video streams of handball matches along with synchronized official positional data and human event annotations for 125min raw data in summary.

1 papers0 benchmarks

emojiSpace embedding

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

BPAEC (bovine pulmonary artery endothelial cells)

This dataset contains the confocal fluorescence microscopy images of nucleus, actin and mitochondria, where each clear image corresponds to 6 out-of-focus images with different degree of blurring. Images acquired below and above the optimal focal plane are blurry and out-of-focus images. In detail, scans were acquired in z-stack of 15 layers spanning the depth (8.4 μm) with 0.6 μm between each slice, z = 7 is the optimal focal plane, z = 1–6 are below the focal plane, z = 8–15 are above the focal plane. We make layers from z = 4 to z = 10 publicly available. Their visual variations in 3-dimensional structure can be negligible. In the end, datasets of actin and nucleus contain 100 in-focus images and 600 out-of-focus images respectively, dataset of mitochondria contain 97 in-focus images and 582 out-of-focus images. Each image is composed of 1024 × 1024 pixels in 8-bit JPG format.

1 papers0 benchmarks

Leishmania parasite dataset

This dataset includes sharp-blur pairs of Leishmania image, which is a protozoan parasite microscopy image dataset of Leishmania, obtained from the preserved slides stained with Giemsa. The paired blur-sharp images are acquired by employing a bright-field microscope (Olympus IX53) with 100× magnification oil immersion objectives.We first capture the sharp images as ground truth, then acquire its corresponding out-of-focus images. The extent and nature of defocusing are random along the optical axis, where the degree of out-of-focus is inconsistent from image-to-image. This dataset includes 764 in-focus and 764 corresponding out-of-focus images, where each image is composed of 2304 × 1728 pixels in 24-bit JPG format.

1 papers0 benchmarksImages, Medical

VILT (Video Instructions Linking for Complex Tasks)

VILT is a new benchmark collection of tasks and multimodal video content. The video linking collection includes annotations from 10 (recipe) tasks, which the annotators chose from a random subset of the collection of 2,275 high-quality 'Wholefoods' recipes. There are linking annotations for 61 query steps across these tasks which contain cooking techniques, chosen from the 189 total recipe steps. As each method results in approximately 10 videos to annotate, the collection consists of 831 linking judgments.

1 papers0 benchmarksVideos

DrugComb

DrugComb is an open-access, community-driven data portal where the results of drug combination screening studies for a large variety of cancer cell lines are accumulated, standardized and harmonized. An actively expanding array of data visualization and computational tools is provided for the analysis of drug combination data. All the data and informatics tools are made freely available to a wider community of cancer researchers.

1 papers0 benchmarks

Wisture Dataset

https://ieee-dataport.org/documents/wi-fi-signal-strength-measurements-smartphone-various-hand-gestures

1 papers0 benchmarks

V-MIND

V-MIND enhanced the MIND dataset with news pictures.

1 papers0 benchmarksImages

FCGEC (FCGEC: Fine-Grained Corpus for Chinese Grammatical Error Correction)

a fine-grained corpus to detect, identify and correct the chinese grammatical errors. collected mainly from multi-choice questions in public school Chinese examinations with multiple references Online Evaluation Site for test set: https://codalab.lisn.upsaclay.fr/competitions/8020

1 papers2 benchmarksTexts

AesVQA

AesVQA is a dataset that contains 72168 high-quality images and 324756 pairs of aesthetic questions. This dataset addresses the task of aesthetic VQA and introduces subjectiveness into VQA tasks.

1 papers0 benchmarksImages, Texts

ChiQA (Chinese VQA)

ChiQA is a dataset designed for visual question answering tasks that not only measures the relatedness but also measures the answerability, which demands more fine-grained vision and language reasoning. It contains more than 40K questions and more than 200K question-images pairs. The questions are real-world image-independent queries that are more various and unbiased.

1 papers0 benchmarksImages, Texts

iDesigner

Fashion trends are constantly evolving, but a trained eye can estimate with some accuracy the signature elements of a particular designer's style.

1 papers3 benchmarks

Kannada Treebank

This dataset was build as a part of development of Treebanks for Indian Languages funded by MeitY, Govt. of India. The Kannada treebank consists of 13.1 K sentences from general, tourism, conversational domains.

1 papers0 benchmarks

BGG dataset (PUBG Gun Sound Dataset)

We recorded gun sounds by changing the type and position of guns to diversify distances and angles in the PUBG environment. The BGG dataset consists of 2,195 samples with 37 different types of guns and five directions, including a silence in which there is no gunfire, but noises exist. The distance from the firearms ranged from 0 meters to 600 meters. The audio was recorded in stereo (i.e., two-channel audio), and each sample contains various environmental noises (e.g., water splashing, walking, and bullet friction).

1 papers0 benchmarksAudio, Stereo

Citations to invalid DOI-identified entities obtained from processing DOI-to-DOI citations to add in COCI

This dataset contains a two-column CSV file, where the first column ("Valid_citing_DOI") contains the DOI of a citing entity retrieved in Crossref, while the second column ("Invalid_cited_DOI") contains the invalid DOI of a cited entity identified by looking at the field "reference" in the JSON document returned by querying the Crossref API with the citing DOI.

1 papers0 benchmarksTabular

time-agnostic-library: benchmarks on execution times and memory

This deposit contains benchmark code, data and results to assess the Python software time-agnostic-library v4.3.0.

1 papers0 benchmarks
PreviousPage 439 of 1000Next