TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Names pairs dataset

Includes co-referent name string pairs along with their similarities.

1 papers0 benchmarksTexts

Sample EEG dataset and looking times data for NEAR (Newborn EEG Artifact Removal)

The sample EEG dataset consists of the newborn EEG data recorded for the work published as:

1 papers0 benchmarks

Korea Composite Stock Price Index

The data contains the following attributes for Korea Stock Price Index (KOSPI) for January 2000–December 2016: 1. Date (YYYY.M(M).D(D)) 2. Opening Price for the date, PX_OPEN 3. Highest Price for the date, PX_HIGH 4. Lowest Price for the date, PX_LOW 5. Closing Price for the date, PX_LAST 6. Total volume traded on the date, PX_VOLUME

1 papers2 benchmarks

SC2ReSet: StarCraft II Esport Replaypack Set

Raw StarCraft II data is subject to processing under the Blizzard end user license agreement (EULA), and in special cases Blizzard AI and Machine Learning License may be applied. Please refer to the materials listed below.

1 papers0 benchmarksReplay data

SC2EGSet: StarCraft II Esport Game State Dataset

SC2EGSet: StarCraft II Esport Game State Dataset

1 papers0 benchmarksReplay data

TGRDB (Tour-Guide Robot Dataset and Benchmark)

Our TGRDB dataset was collected with a 180 fisheye RGB camera on-board of a moving tour-guide robot. A first dataset in tour-guide scenario. Statistical comparisons between TGRDB and existing datasets are refferred to https://arxiv.org/abs/2207.03726. We hope this dataset will drive the progress of research in service robotics, long-term multi-person tracking, and fine-grained or clothes-inconsistency person re-identification.

1 papers0 benchmarks

CareerCoach 2022

The CareerCoach 2022 gold standard is available for download in the NIF and JSON format, and draws upon documents from a corpus of over 99,000 education courses which have been retrieved from 488 different education providers.

1 papers0 benchmarksTexts

Multilingual Persuasion Detection

This dataset contains dialogue lines from the games Knights of the Old Republic 1 & 2 and Neverwinter Nights 1. Some of the dialogue lines are marked as persuasive (which is when the player character is attempting a Persuade skill check.)

1 papers0 benchmarksTexts

Taskography (PDDLGym Taskography)

PDDL dataset of Rearrangement tasks in large-scale 3D scene graphs.

1 papers0 benchmarksTexts

CoCaHis (Colon Cancer Histology Dataset)

Highlights

1 papers0 benchmarksBiomedical, Images, Medical

MatriVasha: (MatriVasha: Compound Character atasetD)

MatriVasha the largest dataset of handwritten Bangla compound characters for research on handwritten Bangla compound character recognition. The proposed dataset contains 120 different types of compound characters that consist of 306,464‬ images written where 152,950 male and 153,514 female handwritten Bangla compound characters. This dataset can be used for other issues such as gender, age, district base handwriting research because the sample was collected that included district authenticity, age group, and an equal number of men and women.

1 papers0 benchmarksImages, Texts

The Mafia Dataset

The Mafia Dataset was created to model the behavior of deceptive actors in the context of the Mafia game, as described in the paper “Putting the Con in Context: Identifying Deceptive Actors in the Game of Mafia”. We hope that this dataset will be of use to others studying the effects of deception on language use.

1 papers0 benchmarksDialog, Interactive, Texts

ANTILLES (ANTILLES: An Open French Linguistically Enriched Part-of-Speech Corpus)

ANTILLES is a part-of-speech tagging corpus based on UD_French-GSD which was originally created in 2015 and is based on the universal dependency treebank v2.0.

1 papers1 benchmarksTexts

WinoNB

This dataset consists of Winograd schemas that test coreference resolution systems' ability to differentiate singular vs plural they/them pronouns. It consists of 4077 templates, each with a group of people, a singular person (which can be filled with a name or a generic "someone") and a single they/them pronoun to resolve.

1 papers0 benchmarks

Wind speed and power potential for Switzerland

Summary:

1 papers0 benchmarks

Article Bias Prediction

Article-Bias-Prediction Dataset The articles crawled from www.allsides.com are available in the ./data folder, along with the different evaluation splits.

1 papers0 benchmarks

Crowd Activity Dataset

This dataset concentrates on the activities of the crowd for a fine-grained image classification task, named as Crowd Activity dataset, as automatically understanding crowd activity is meaningful for social security. This dataset is newly collected, where the images are mainly searched on the Internet and collected from streets by mobile phones. All images in this dataset contain at least one text instance. The categories come from activities of daily living and demonstrations stimulated by hot events in recent years. Specifically, this dataset consists of 21 categories and 8785 images in total. The 21 categories broadly fall into two types: activities of daily living(i.e., celebrating Christmas, holding sport meeting, holding concert, celebrating birthday party, celebrity speech, teaching, graduation ceremony, picnic, press briefing, shopping, celebrating Thanks giving day) and demonstrations (i.e., protecting animals, protecting environment, appealing for peace, Brexit, COVID-19, ele

1 papers0 benchmarksImages

HTDM (Hypertention Disease Medication)

Hypertention Disease Medication dataset.

1 papers0 benchmarksGraphs, Medical

Locount

Loucount is a retail object detection and and counting dataset with rich annotations in retail stores, which consists of 50, 394 images with more than 1.9 million object instances in 140 categories

1 papers0 benchmarksImages

Replication Data for: Do uHear? Validation of uHear App for Preliminary Screening of Hearing Ability in Soundscape Studies

Audiogram data based on a "gold standard" audiometer and the uHear iOS application of 163 participants

1 papers0 benchmarksBiomedical
PreviousPage 433 of 1000Next