TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

r/AmITheAsshole Reddit threads

This dataset is made of 6366 threads collected from the r/AmITheAsshole community on Reddit. The dataset contains a total of 6,372,251 comments. The collected threads constitute the “top” submissions — those having the highest score, measured as the difference between upvotes and downvotes of a post. We downloaded them using PRAW, running 10 different queries across various temporal scopes, and then cleaning the obtained dataset by removing duplicated threads. Please refer to the paper, specifically to Table 3, for more details about the dataset.

1 papers0 benchmarksTexts, Time series

Regional AQ Datasets

The primary environmental health threat in the WHO European Region is air pollution, impacting the daily health and well-being of its citizens significantly. To effectively understand the impact, and dynamics of air quality a detailed investigation of different environmental, weather, and land cover indices is appropriate. To this end, this paper introduces three European cities’ spatiotemporal datasets, customized for air pollution monitoring at a regional level. The datasets are composed of major air quality, weather measurements and land use information. The duration is approximately from 2020 to 2023 with an hourly temporal resolution and a spatial resolution of 0.005°. The temporal and spatiotemporal datasets are publicly released aiming to provide a solid foundation for researchers, analysts, and practitioners to conduct in-depth analyses of air pollution dynamics.

1 papers0 benchmarksTime series

Hilti-Oxford Dataset

This benchmark is based on the HILTI-OXFORD Dataset, which has been collected on construction sites as well as on the famous Sheldonian Theatre in Oxford, providing a large range of difficult problems for SLAM.

1 papers0 benchmarks

Newer College Dataset Extension (Multi-Camera LiDAR Inertial)

In this expansion dataset, we have extended the ground truth pointcloud of New College to include the colllege’s cloister (featured in Harry Potter films) and the Monk’s Passage. We also introduce another outdoor large scale environment, Maths Institute at the University of Oxford, which is shown in the figure below (right). Collection 1 and 2 contain datasets collected in New college, where the Park sequence mimics the original long sequence of the 2011 New College Dataset. Collection 3 contains three difficulty levels around the Maths Institute environment. Note that in this dataset Lidar, Cameras and IMU are all hardware synchronised through PTP.

1 papers0 benchmarks

NL2GQL Dataset (NL2GQL developped for R3-NL2GQL)

A bilingual (English and Chinese natural language queries) dataset which has NL queries annotated with their corresponding GQL queries (i.e. nGQL). Each data sample in the train data contains 4 pieces of information: prompt represents a natural language query, content represents a standard nGQL, reason represents the inference part that needs to be output by the reranker, and schema represents the code structure schema corresponding to this sentence. Each data sample in the test data contains 6 pieces of information, prompt represents natural language query, content represents gold nGQL, text_schema is used for the vanilla experiment, schema represents the code structure schema corresponding to this sentence, class represents which graph database space this sentence corresponds to, and result represents the results obtained using gold nGQL.

1 papers0 benchmarksTexts

A Dataset for Mechanical Mechanisms

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

geometricShapes14k (Images made up of geometric shapes - ControlNet)

The dataset was curated from the 1% data sample file of the Wikipedia-based Image Text (WIT) Dataset. Captions for images were generated using the BLIP model. Control images were derived using the Primitive software, resulting in a dataset of 14,279 pairs of control and target images with captions.

1 papers0 benchmarks

SECURES-Met

Click to add a brisef description of the datdaset (Markdown and LaTeX enabled).

1 papers0 benchmarksTexts

https://doi.org/10.6084/m9.figshare.1512427.v5 (Figshare Brain Dataset)

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

ChangeVPR

Scene change detection (SCD) dataset tailored for generalizable SCD algorithm. It consists of change-labeld images from SF-XL, St Lucia, Nordland which are widely used in VPR research.

1 papers1 benchmarksImages

MEEG

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers2 benchmarks

WHU - Audio ENF (WHU - Audio Electric Network Frequency)

The Whu dataset is an audio only dataset thought for testing ENF detection. It is divided into two parts: one with recordings containing ENF traces (H1) and one without them (H0). The recordings with ENF traces are coupled with the corresponding ENF reference (H1_ref). The dataset consists of 60 real-world audio recordings captured around Wuhan University campus, featuring a diverse range of environments and conditions. The recordings were made at a sampling rate of 44.1 kHz with 16-bit quantization and mono channel. Among the 60 recordings, 50 were found to have captured and verified ENF signals, which were confirmed by comparing the recording times with a reference database. The remaining 10 recordings were made in open exterior environments with strong noise and interference, and were often affected by Doppler effects due to the user walking while recording. The final dataset has 130 audio recordings in H1 and 40 in H0, obtained by randomly cropping them to durations varying from 5

1 papers0 benchmarksAudio

ENF moving video (Electric Network Frequency Moving Video Dataset)

The ENF moving video dataset, which is a subset of the dataset used in Temporal Localization of Non-Static Digital Videos Using the Electrical Network Frequency , consists of video recording without the audio channel coupled with the corresponding power ENF signal reference in WAV format at a rate of 1 kHz. The dataset is made of 8 video clips recorded in Europe at 29.97 frames per second, with a duration of approximately 11-12 minutes, using a GoPro Hero 4 Black and an NK AC3061-4KN camera. In terms of content, videos 1-3 are entirely stationary, videos 4-5 are predominantly stationary with some movement, and videos 6-8 are non-stationary, meaning the camera is fixed, but there are moving objects in most frames. All videos depict natural, everyday indoor scenes (i.e., not plain backgrounds).

1 papers0 benchmarksAudio, Videos

Candombe (Candombe Recordings Dataset)

35 recordings of Candombe music with beat and downbeat annotations.

1 papers2 benchmarksAudio

Filosax

48 multitrack jazz recordings with many annotations.

1 papers2 benchmarksMusic

Hainsworth

S. W. Hainsworth and M. D. Macleod, “Particle filtering applied to musical tempo tracking,” EURASIP Journal on Advances in Signal Processing, vol. 2004, pp. 1–11, 2004

1 papers2 benchmarksAudio

Harmonix (The Harmonix Set)

Beats, downbeats, and functional structural annotations for 912 Pop tracks.

1 papers2 benchmarksAudio

HJDB

J. Hockman, M. E. Davies, and I. Fujinaga, “One in the jungle: Downbeat detection in hardcore, jungle, and drum and bass.” in Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), 2012.

1 papers2 benchmarksAudio

JAAH (Jazz Audio-Aligned Harmony)

Eremenko, E. Demirel, B. Bozkurt, and X. Serra, “Audio-aligned jazz harmony dataset for automatic chord transcription and corpus-based research,” in Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), 2018

1 papers2 benchmarksAudio

SIMAC

F. Gouyon, “A computational approach to rhythm description — Audio features for the computation of rhythm periodicity functions and their use in tempo induction and music content processing,” Ph.D. dissertation, Universitat Pompeu Fabra, 2006

1 papers1 benchmarksAudio
PreviousPage 517 of 1000Next