TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Identity Access Management dataset

We release 280 synthetic IAM graphs generated using IAM graphs of commercial companies. Specifically, we vary the number of nodes, but keep graph density as is, i.e. in the range of 0.259 ± 0.198 (avg ± std). To generate a synthetic graph, we first sample the number of users and datastores from uniform distributions over the following intervals [10, 150] and [50, 300] respectively that cover variations of those parameters across real graphs. After fixing node counts we sample with replacement the actual nodes from a real world graph, which is chosen at random. Then we add Gaussian N(0, 0.01) noise to node embeddings and renormalize them. To match the graph density with the density of the underlying baseline we sample edges from a multinomial distribution, where each component is proportional to the cosine distance between a user and a datastore embeddings. Also we enforce the invariant that dynamic edges are always a subset of all permission edges. A synthetic graph generated in such

1 papers0 benchmarksGraphs

GMD-12

A dataset for medical consultation dialogues. See our related paper for more details: https://arxiv.org/pdf/2204.13953.pdf

1 papers0 benchmarks

r/transprogrammer survey results

Questions regarding computer science education for members of the r/transprogrammer Reddit. Used for the paper "Why The Trans Programmer?" by Skye Kychenthal.

1 papers0 benchmarks

NLU Evaluation Corpora

This project is a collection of three corpora which can be used for evaluating chatbots or other conversational interfaces. Two of the corpora were extracted from StackExchange, one from a Telegram chatbot.

1 papers0 benchmarks

OntoRock

OntoRock is a benchmark for evaluating the robustness of existing NER models via a systematic evaluation protocol.

1 papers0 benchmarks

UAGD (Uniform Age and Gender Dataset)

The source images of UAGD is manually selected from APPA-REAL, UTKFace and AgeDB datasets very carefully, which means only face images that are having large poses, containing noise pixels, bearing various expressions, and under different illuminations could be chosen. We also double clean and remove the images that have wrong or not pretty sure label by crowdsourcing platform. UAGD has almost the same number of female and male images in each age, about 75 female and 75 male, total 150 face.

1 papers0 benchmarks

WikiMulti (WikiMulti: a Corpus for Cross-Lingual Summarization)

wikimulti is a dataset for cross-lingual summarization based on Wikipedia articles in 15 languages.

1 papers0 benchmarks

COVMis-Stance

COVMis-Stance is a stance detection dataset for COVID-19 misinformation. It consists of fake news and claims related to COVID. Fake news was collected from articles fact-checking sites, and fake claims were from the WHO official Twitter. It contains 2631 tweets annotated for stance towards 111 COVID19 misinformation items.

1 papers0 benchmarksTexts

Custom Spatio-Temporal Action Video Dataset

This spatio-temporal actions dataset for video understanding consists of 4 parts: original videos, cropped videos, video frames, and annotation files. This dataset uses a proposed new multi-person annotation method of spatio-temporal actions. First, we use ffmpeg to crop the videos and frame the videos; then use yolov5 to detect human in the video frame, and then use deep sort to detect the ID of the human in the video frame. By processing the detection results of yolov5 and deep sort, we can get the annotation file of the spatio-temporal action dataset to complete the work of customizing the spatio-temporal action dataset.

1 papers0 benchmarksVideos

TASTEset

TASTEset Recipe Dataset and Food Entities Recognition is a dataset for Named Entity Recognition (NER) which consists of 700 recipes with more than 13,000 entities to extract.

1 papers0 benchmarksTexts

Pirá (Pirá: A Bilingual Portuguese-English Dataset for Question-Answering about the Ocean)

A large set of questions and answers about the ocean and the Brazilian coast both in Portuguese and English. Pirá is a crowdsourced question answering (QA) dataset on the ocean and the Brazilian coast designed for reading comprehension.

1 papers0 benchmarksTexts

Danish Airs and Grounds

Danish Airs and Grounds (DAG) is a large collection of street-level and aerial images targeting such cases. Its main challenge lies in the extreme viewing-angle difference between query and reference images with consequent changes in illumination and perspective. The dataset is larger and more diverse than current publicly available data, including more than 50 km of road in urban, suburban and rural areas. All images are associated with accurate 6-DoF metadata that allows the benchmarking of visual localization methods.

1 papers0 benchmarksImages

Kinetics-GEB+

Kinetics-GEB+ (Generic Event Boundary Captioning, Grounding and Retrieval) is a dataset that consists of over 170k boundaries associated with captions describing status changes in the generic events in 12K videos.

1 papers40 benchmarksVideos

Sen4AgriNet (A Sentinel-2 multi-year, multi-country benchmark dataset for crop classification and segmentation with deep learning)

A Sentinel-2 based time series multi country benchmark dataset, tailored for agricultural monitoring applications with Machine and Deep Learning. Sen4AgriNet dataset is annotated from farmer declarations collected via the Land Parcel Identification System (LPIS) for harmonizing country wide labels. Sen4AgriNet is the only multi-country, multi-year dataset that includes all spectral information. It is constructed to cover the period 2016-2020 for Catalonia and France, while it can be extended to include additional countries. Currently, it contains 42.5 million parcels, which makes it significantly larger than other available archives.

1 papers0 benchmarksImages

YouTube-GDD (YouTube-GDD: A challenging gun detection dataset with rich contextual information)

YouTubeGun Detection Dataset is collected from 343 high-definition YouTube videos and contains 5000 well-chosen images, in which 16064 instances of gun and 9046 instances of person are annotated. Compared to other datasets, YouTube-GDD is "dynamic", containing rich contextual information

1 papers0 benchmarksImages

OC-Cityscape (Out-of-Context Cityscapes)

Out-of-Context Cityscapes (OC-Cityscapes) is a new dataset build by replacing roads in the validation data of Cityscapes with various textures such as water, sand, grass, etc.

1 papers0 benchmarks

ANUBIS (Skeleton-Based Action Recognition Dataset)

ANUBIS is a large-scale human skeleton dataset containing 80 actions. Compared with previously collected datasets, ANUBIS is advantageous in the following four aspects: (1) employing more recently released sensors; (2) containing novel back view; (3) encouraging high enthusiasm of subjects; (4) including actions of the COVID pandemic era.

1 papers0 benchmarksImages

Cross-View Cross-Scene Multi-View Crowd Counting Dataset

A large synthetic multi-camera crowd counting dataset with a large number of scenes and camera views to capture many possible variations, which avoids the difficulty of collecting and annotating such a large real dataset.

1 papers0 benchmarksImages

TuGebic (A Turkish-German Bilingual Code-Switching Corpus)

TuGebic is a corpus of recordings of spontaneous speech samples from Turkish-German bilinguals, and the compilation of a corpus called TuGebic. Participants in the study were adult Turkish and German bilinguals living in Germany or Turkey at the time of recording in the first half of the 1990s. The data were manually tokenised and normalised, and all proper names (names of participants and places mentioned in the conversations) were replaced with pseudonyms. Token-level automatic language identification was performed, which made it possible to establish the proportions of words from each language.

1 papers0 benchmarksTexts

VIS-TIR

A visible-light and thermal-infrared images dataset for dual-spectrum depth estimation.

1 papers0 benchmarksImages
PreviousPage 427 of 1000Next