TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

French Dialect Samples

We collated subcorpora each between 50,000 and 70,000 words, containing samples of national dialects of French across different countries: Algeria, Democratic Republic of Congo, France, Ivory Coast, Morocco and Senegal.

1 papers0 benchmarks

LGI-PPGI (LGI-PPGI-Face-Video-Database)

LGI-PPGI is a dataset for heart Rate estimation from face videos in the wild.

1 papers0 benchmarks

Cryptocurrency User Attitudes Towards the Environmental Impact of Proof-of-Work in Nigeria: Online Survey Results

The "Cryptocurrency User Attitudes Towards the Environmental Impact of Proof-of-Work in Nigeria" online survey is a convenience sample survey of residents of Nigeria aged 16 and over that have participated in Bitcoin transactions in the past. Between November 2021 and March 2022, participants were asked their opinions on cryptocurrencies, the environmental effects of using them, and attitudes towards these effects.

1 papers0 benchmarks

Auditory Detection of Sound (ADS)

Test dataset for unsupervised anomaly detection in sound (ADS).

1 papers0 benchmarksAudio

Sequence Consistency Evaluation (SCE) tests

Sequence Consistency Evaluation (SCE) consists of a benchmark task for sequence consistency evaluation (SCE).

1 papers0 benchmarksImages, Time series

Handwritten Devanagari Character Recognition

This is an image database of Handwritten Devanagari characters. There are 46 classes of characters with 2000 examples each. The dataset is split into training set(85%) and testing set(15%).

1 papers0 benchmarks

FullTextPeerRead

FullTextPeerRead is a dataset created by Jeong et al. for context-aware citation recommendation. It contains context sentences to cited references and paper metadata, which makes it a well-organized dataset for a context-aware paper recommendation.

1 papers0 benchmarks

arXiv-200

A newly proposed dataset for local citation recommendation, consisting of 3.2 million local citation sentences along with the title and the abstract of both the citing and the cited papers. Around 1.66 million papers' titles and abstracts are available in the database.

1 papers0 benchmarks

MOS Dataset (Microblog Opinion Summarisation)

This dataset was used in the paper 'Template-based Abstractive Microblog Opinion Summarisation' (to be published at TACL, 2022). The data is structured as follows: each file represents a cluster of tweets which contains the tweet IDs and a summary of the tweets written by journalists. The gold standard summary follows a template structure and depending on its opinion content, it contains a main story, majority opinion (if any) and/or minority opinions (if any).

1 papers0 benchmarks

Box-Jenkins (Box-Jenkins Gas Furnace Problem)

Box-Jenkins gas furnace, a well-known time series forecasting problem

1 papers0 benchmarks

32vis

Dataset for the 32 years of IEEE VIS

1 papers0 benchmarks

Bus Stop Spacings for Transit Providers in the US

Transit agencies use the General Transit Feed Specification (GTFS) to publish transit data. More and more cities across the globe are adopting this GTFS format to represent their transit network. However, the GTFS format is not convenient as the information is spread across multiple files and is cumbersome with rules. This dataset is a collection of over 600 Transit agencies in the US with a concise and easy representation of GTFS data in the form of segments.

1 papers0 benchmarks

Bengali Ekman's Six Basic Emotions Corpus

The dataset contains 36000 Bangla data based on Ekman's six basic emotions. This data was first introduced in the paper Alternative non-BERT model choices for the textual classification in low-resource languages and environments. The whole dataset is balanced and evenly distributed among all the six classes.

1 papers1 benchmarksTexts

Capriccio (Sentiment Analysis + Data Drift)

Capriccio is a sentiment classification dataset on tweets that simulates data drift. It is created by slicing the Sentiment140 dataset (homepage, Huggingface datasets) with a sliding window of 500,000 tweets, resulting in 38 slices. Thus, each slice can be used to represent the training/validation dataset of a sentiment classification model that is re-trained every day. Each slice has 425,000 tweets for training (file named %d_train.json) and 75,000 tweets for validation (file named %d_val.json).

1 papers0 benchmarksTexts

ABCD Study (Adolescent Brain Cognitive Development)

The ABCD Study is a prospective longitudinal study starting at the ages of 9-10 and following participants for 10 years. The study includes a diverse sample of nearly 12,000 youth enrolled at 21 research sites across the country. It measures brain development (via structural, task functional, and resting state functional imaging), social, emotional, and cognitive development, mental health, substance use and attitudes, gender identity and sexual health, bio-specimens, as well as a variety of physical health, and environmental factors.

1 papers0 benchmarks3D, Medical

NCANDA (National Consortium on Alcohol and Neurodevelopment in Adolescence)

The NCANDA consortium is composed of an Administrative component at the University of California San Diego, a Data Analysis and Informatics component at SRI International, and five research sites (University of California San Diego, SRI International, Duke University, the University of Pittsburgh, and the Oregon Health & Science University). A sample of 831 individuals (ages 12-21) were recruited for the study across the five research sites. The enrolled participants are followed in an accelerated longitudinal design that involves structural and functional imaging of the brain along with extensive neuropsychological and clinical assessments.

1 papers0 benchmarks3D, Medical

CrossDomainTypes4Py

A Python Dataset for Cross-Domain Evaluation of Type Inference Systems

1 papers0 benchmarks

AquaTrash

This dataset contains 369 images of Trash used for deep learning. Each image is manually labelled by our team for accurate detections making a total of 470 bounding boxes. There are total 4 classes {(0: glass), (1:paper), (2:metal), (3:plastic)}

1 papers5 benchmarksImages

Survival Analysis of Heart Failure Patients

The dataset contains cardiovascular medical records taken from 299 patients. The patient cohort comprised of 105 women and 194 men between 40 and 95 years in age. All patients in the cohort were diagnosed with the systolic dysfunction of the left ventricle and had previous history of heart failures. As a result of their previous history every patient was classified into either class III or class IV of New York Heart Association (NYHA) classification for various stages of heart failure.

1 papers0 benchmarks

Anime Face Dataset by Character Name

This dataset is suitable for the image classification model. Train image classification model to classify anime characters by face image. This dataset includes 130 characters with 75 images each, scrapped from Danbooru.

1 papers0 benchmarks
PreviousPage 435 of 1000Next