TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

CASCONet

CASCONet is a a collection of data about the CAS Conference (CASCON) for the past 25 years including information about papers, technology showcase demos, workshops, and keynote presentations.

1 papers0 benchmarks

US-Accidents

This is a countrywide traffic accident dataset, which covers 49 states of the United States. The data is continuously being collected from February 2016, using several data providers, including two APIs which provide streaming traffic event data. These APIs broadcast traffic events captured by a variety of entities, such as the US and state departments of transportation, law enforcement agencies, traffic cameras, and traffic sensors within the road-networks. Currently, there are about 4.2 million accident records in this dataset.

1 papers0 benchmarks

JSESM Publications

1 papers0 benchmarks

HoaxItaly

HoaxItaly consists of over 1 million tweets shared during 2019 and containing links to thousands of news articles published on two classes of Italian outlets: (1) disinformation websites, i.e. outlets which have been repeatedly flagged by journalists and fact-checkers for producing low-credibility content such as false news, hoaxes, click-bait, misleading and hyper-partisan stories; (2) fact-checking websites which notably debunk and verify online news and claims. The dataset includes title and body for approximately 37k news articles.

1 papers0 benchmarksGraphs, Texts

GED (Gridded Establishment Dataset)

GED is a dataset on the economic activity of mainland China, which measures the volume of establishments at a 0.01 latitude by 0.01 longitude scale. Specifically, the dataset captures the geographically based opening and closing of approximately 25.5 million firms that registered in mainland China over the period 2005-2015. The characteristics of fine granularity and long-term observability give the GED a high application value.

1 papers0 benchmarks

Interactive Gibson Environment

Interactive Gibson is a comprehensive benchmark for training and evaluating Interactive Navigation: robot navigation strategies where physical interaction with objects is allowed and even encouraged to accomplish a task. The benchmark has two main components:

1 papers0 benchmarksEnvironment

Dataset of Video Game Development Problems

This is a grounded dataset describing software-engineering problems in video-game development extracted from postmortems. The dataset was created using an iterative method through which the authors manually coded more than 200 postmortems spanning 20 years (1998 to 2018) and extracted 1,035 problems related to software engineering while maintaining traceability links to the postmortems. The problems were grouped in 20 different types. This dataset is useful to understand the problems faced by developers during video-game development, providing researchers and practitioners a starting point to study video-game development in the context of software engineering.

1 papers0 benchmarks

A Dataset of State-Censored Tweets

This is a dataset of 583,437 tweets by 155,715 users that were censored between 2012-2020 July. It also contains 4,301 accounts that were censored in their entirety. Additionally, another set of tweets is related, consisting of 22,083,759 supplemental tweets made up of all tweets by users with at least one censored tweet as well as instances of other users retweeting the censored user.

1 papers0 benchmarksTexts

Enterprise-Driven Open Source Software

This is a dataset of open source software developed mainly by enterprises rather than volunteers. This can be used to address known generalizability concerns, and, also, to perform research on open source business software development. Based on the premise that an enterprise's employees are likely to contribute to a project developed by their organization using the email account provided by it, we mine domain names associated with enterprises from open data sources as well as through white- and blacklisting, and use them through three heuristics to identify 17,264 enterprise GitHub projects. We provide these as a dataset detailing their provenance and properties. A manual evaluation of a dataset sample shows an identification accuracy of 89%.

1 papers0 benchmarks

Test Scene Dataset for Physically Based Rendering

This is a comprehensive test database of scenes that treat different light setups in conjunction with diverse materials. It delivers a comprehensive foundation for evaluating existing and newly developed rendering techniques.

1 papers0 benchmarks

DSSN (DAIICT Spatio-Temporal Network)

DSSN is a spatiotemporal dataset of 0.7 million data points of continuous location data logged at an interval of every 2 minutes by mobile phones of 46 subjects. The total number of data points reported in this dataset are 6,59,268. The total number of subjects using the application to record data are 74, however with cleaning based on quality checks. The number was reduced to 46. The data recorded varies in accuracy with an average accuracy of 36.0 meters.

1 papers0 benchmarks

CoronaVis

CoronaVis is a dataset of tweets related to coronavirus.

1 papers0 benchmarksTexts

DroidBugs

DroidBugs is a benchmark for Automated Program Repair (APR) of Android applications.

1 papers0 benchmarks

Near-Collision

Near-Collision is a large-scale dataset of 13,658 egocentric video snippets of humans navigating in indoor hallways. In order to obtain ground truth annotations of human pose, the videos are provided with the corresponding 3D point cloud from LIDAR.

1 papers0 benchmarksLiDAR, Point cloud, Videos

Apiza Corpus

The Apiza Corpus is a WoZ-like (Wizard of Oz) set of dialogues between 30 programmers and a simulated virtual assistant. This corpus can be used to study or train a virtual assistant for software engineering.

1 papers0 benchmarksTexts

YoutubeGraph-Dyn

YoutubeGraph-Dyn is an evolving graph dataset generated from YouTube real-world interactions. It can be used to study temporal evolution on graphs. YoutubeGraph-Dyn provides intra-day time granularity (with 416 snapshots taken every 6 hours for a period of 104 days), multi-modal relationships that capture different aspects of the data, multiple attributes including timestamped, non-timestamped, word embeddings, and integers.

1 papers0 benchmarksGraphs

HoMG

HoMG is a holoscopic 3D micro-gesture dataset captured with a holoscopic 3D camera. HoMG database recorded the image sequence of 3 conventional gestures from 40 participants under different settings and conditions. For the purpose of H3D micro-gesture recognition, HoMG has a video subset of 960 videos and a still image subset with 30,635 images.

1 papers0 benchmarks

BA (Binary Alphabet)

1 papers1 benchmarks

Dataset of Rendered Chess Game State Images

This dataset contains 4,888 synthetic images of chess game states that occurred in games played by Magnus Carlsen. The images were rendered in Blender at different angles and lighting conditions.

1 papers0 benchmarks

Risk-Aware Planning Dataset

Risk-Aware Planning is a dataset that contains the overhead images and their semantic segmentation captured by a drone from the CityEnviron environment in AirSim simulator.

1 papers0 benchmarksImages
PreviousPage 390 of 1000Next