TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Essays (Stream-of-consciousness Essays)

J. W. Pennebaker and L. A. King, “Linguistic styles: Language use as an individual difference,” J. Pers. Soc. Psychol., vol. 77, no. 6, pp. 1296–1312, Dec. 1999, doi: 10.1037/0022-3514.77.6.1296.

1 papers4 benchmarks

ECTF (Early COVID-19 Twitter Fake news)

ECTF is a dataset for Twitter fake news detection in the Covid-19 domain.

1 papers0 benchmarksTexts

Extended UCF Crime

The Extended UCF Crime extends the UCF Crime data set that consists of 13 anomaly classes. The extension adds two different anomaly classes to the data set, which are ”molotov bomb” and ”protest” classes. It also adds 33 videos to the fighting class. In total, the extension adds 216 videos to the training set, 17 videos to the test set.

1 papers0 benchmarksVideos

Android Common Libraries

This dataset was constructed from an analysis of about 1.5 million apps from Google Play to identify a set of common libraries, to facilitate Android app analysis. It contains 1,113 libraries supporting common functionalities and 240 libraries for advertisement.

1 papers0 benchmarks

Workflow Trace Archive

The Workflow Trace Archive (WTA) is an open-access archive of workflow traces from diverse computing infrastructures. The WTA includes >48 million workflows captured from >10 computing infrastructures, representing a broad diversity of trace domains and characteristics.

1 papers0 benchmarks

BLEBeacon

The BLEBeacon dataset is a collection of Bluetooth Low Energy (BLE) advertisement packets/traces generated from BLE beacons carried by people following their daily routine inside a university building for a whole month. A network of Raspberry Pi 3 (RPi)-based edge devices were deployed inside a multi-floor facility continuously gathering BLE advertisement packets and storing them in a cloud-based environment. The focus is on presenting a real-life realization of a location-aware sensing infrastructure, that can provide insights for smart sensing platforms, crowd-based applications, building management, and user-localization frameworks.

1 papers0 benchmarks

AdobeIndoorNav

AdobeIndoorNav is a dataset collected in real-world to facilitate the research in DRL based visual navigation. The dataset includes 3D reconstruction for real-world scenes as well as densely captured real 2D images from the scenes. It provides high-quality visual inputs with real-world scene complexity to the robot at dense grid locations.

1 papers0 benchmarks

nicolingua-0003-west-african-radio-corpus (West African Radio Corpus)

This dataset contains 17,090 audio clips of length 30 seconds sampled from archives collected from 6 Guinean radio stations. The broadcasts consist of news and various radio shows in languages including French, Guerze, Koniaka, Kissi, Kono, Maninka, Mano, Pular, Susu, and Toma. Some radio shows include phone calls, background and foreground music, and various noise types. We collected this dataset for the purpose of unsupervised speech representation learning. A validation set of 300 tagged audio clips is also included.

1 papers0 benchmarks

nicolingua-0004-west-african-va-asr-corpus (West African Virtual Assistant Speech Recognition Corpus)

This dataset contains 10,083 recorded utterances in French, Maninka, Pular and Susu from 49 speakers (16 female and 33 male) ranging from 5 to 76 years old on a variety of devices.

1 papers0 benchmarks

BuGL

BuGL is a large-scale cross-language dataset for bug localization in code. BuGL constitutes of more than 10,000 bug reports drawn from open-source projects written in four programming languages, namely C, C++, Java, and Python. The dataset consists of information which includes Bug Reports and Pull-Requests. BuGL aims to unfold new research opportunities in the area of bug localization.

1 papers0 benchmarks

BIRD (Big Impulse Response Dataset)

BIRD (Big Impulse Response Dataset) is an open dataset that consists of 100,000 multichannel room impulse responses (RIRs) generated from simulations using the Image Method, making it the largest multichannel open dataset currently available. These RIRs can be used to perform efficient online data augmentation for scenarios that involve two microphones and multiple sound sources.

1 papers0 benchmarksAudio

BCSD (Bank Check Segmentation Dataset)

The dataset consists of images of 158 filled out bank checks containing various complex backgrounds, and handwritten text and signatures in the respective fields, along with both pixel-level and patch-level segmentation masks for the signatures on the checks. Please visit the dataset homepage for more details.

1 papers0 benchmarksImages

AbuseAnalyzer Dataset

The dataset contains 7,601 Gab posts classified on three different aspects: abuse presence or not, abuse severity and abuse target.

1 papers0 benchmarksTexts

DX7 Timbre Dataset

This is a dataset of 22.5 hours of synthesized audio using the open-source learnfm clone of the DX7 FM synthesizer, based upon 31K presets from Bobby Blue. These represent "natural'' synthesis sounds---i.e.presets devised by humans.

1 papers0 benchmarksAudio

WhatsApp, Doc?

This is a large-scale dataset collected from WhatsApp public groups. It has been created from 178 public groups containing around 45K users and 454K messages. This dataset allows researchers to ask questions like (i) Are WhatsApp groups a broadcast, multicast or unicast medium? (ii) How interactive are users, and how do these interactions emerge over time? (iii) What geographical span do WhatsApp groups have, and how does geographical placement impact interaction dynamics? (iv) What role does multimedia content play in WhatsApp groups, and how do users form interaction around multimedia content? (v) What is the potential of WhatsApp data in answering further social science questions, particularly in relation to bias and representability?

1 papers0 benchmarks

THÖR

THÖR is a dataset with human motion trajectory and eye gaze data collected in an indoor environment with accurate ground truth for position, head orientation, gaze direction, social grouping, obstacles map and goal coordinates. THOR also contains sensor data collected by a 3D lidar and involves a mobile robot navigating the space.

1 papers0 benchmarksLiDAR

FacebookVideoLive18

FacebookVideosLive18 dataset includes 1,000,000 Facebook live videos with their metadata (title, source, length, creation time, description, etc.), broadcasters locations and viewers locations. We are using a set of synchronised scripts that allow to have a global view of the real time streaming system every 3 minutes. We believe that our dataset is the first that tracks the locations and behaviors of live viewers. We expect FacebookVideosLive18 to support various trending research areas such as cloud computing, multimedia data allocation, multi-cloud allocation, edge computing, edge caching and transcoding, data analytics, etc.

1 papers0 benchmarks

FTR-18

FTR-18 is a multilingual rumour dataset on football transfer news. Transfer rumours are continuously published by sports media. They can both harm the image of player or a club or increase the player's market value. The proposed dataset includes transfer articles written in English, Spanish and Portuguese. It also comprises Twitter reactions related to the transfer rumours. FTR-18 is suited for rumour classification tasks and allows the research on the linguistic patterns used in sports journalism.

1 papers0 benchmarksTexts

GitHub Repository Deduplication

This is a dataset of 10.6 million GitHub projects that are copies of others, and link each record with the project's ultimate parent. The ultimate parents were derived from a ranking along six metrics. The related projects were calculated as the connected components of an 18.2 million node and 12 million edge denoised graph created by directing edges to ultimate parents. The graph was created by filtering out more than 30 hand-picked and 2.3 million pattern-matched clumping projects. Projects that introduced unwanted clumping were identified by repeatedly visualizing shortest path distances between unrelated important projects.

1 papers0 benchmarks

PersianQA (Persian Question Answering Dataset)

PersianQA: a dataset for Persian Question Answering Persian Question Answering (PersianQA) Dataset is a reading comprehension dataset on Persian Wikipedia. The crowd-sourced the dataset consists of more than 9,000 entries. Each entry can be either an impossible-to-answer or a question with one or more answers spanning in the passage (the context) from which the questioner proposed the question. Much like the SQuAD2.0 dataset, the impossible or unanswerable questions can be utilized to create a system which "knows that it doesn't know the answer".

1 papers0 benchmarksTexts
PreviousPage 389 of 1000Next