TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Turn-Level Goals Dataset

This dataset is a record of the active learning data collected from interacting with PersonaGPT to fine-tune its actions toward turn-level goals, which are text descriptions of decoding goals for each response in a conversation.

1 papers0 benchmarks

The Berka Dataset

The Berka dataset is a collection of financial information from a Czech bank. The dataset deals with over 5,300 bank clients with approximately 1,000,000 transactions. Additionally, the bank represented in the dataset has extended close to 700 loans and issued nearly 900 credit cards, all of which are represented in the data.

1 papers0 benchmarks

Carbon Intensity 2020 (Carbon Intensity Data of Germany, Great Britain, France, and California in 2020)

Energy production and carbon intensity datasets for the regions Germany, Great Britain, France (all via the ENTSO-E Transparency Platform) and California (via California ISO) for the entire year 2020 +-10 days.

1 papers0 benchmarks

SMLM CEP152-Complex FITS Images

The following files comprise 19 sets of 40,000 images, each set corresponding to a different rendering sigma as described in the paper.

1 papers0 benchmarks

SALAMI (Structural Analysis of Large Amounts of Music Information)

Comes from https://ddmal.music.mcgill.ca/research/SALAMI/:

1 papers0 benchmarks

Lower-limb Kinematics and Kinetics During Continuously Varying Human Locomotion

This dataset reports the lower-limb kinematics and kinetics of ten able-bodied subjects walking at multiple inclines (± 0°, 5°, and 10°) and speeds (0.8 m/s, 1 m/s, and 1.2 m/s), running over level-ground at multiple speeds (1.8 m/s, 2 m/s, 2.2 m/s, and 2.4 m/s), walking and running with constant acceleration and deceleration (± 0.2 m/s2, and 0.5 m/s2), and stair ascent/descent with multiple stair inclines (± 20°, 25°, 30°, and 35°). This dataset also includes sit-stand transitions, walk-run transitions, and walk-stairs transitions. Data were recorded by a Vicon motion capture system and, for applicable tasks, a Bertec instrumented treadmill. This dataset can aid in the development of kinematic models of multi-activity human locomotion and the design and control of agile wearable robots.

1 papers0 benchmarks

TMBuD

TMBuD is a dataset for building recognition and 3D reconstruction of human made structures in urban scenarios. The dataset features 160 images of buildings from Timişoara, Romania, with a resolution of 768 x 1024 pixels each. The proposed dataset will allow proper evaluation of salient edges and semantic segmentation of images focusing on the street view perspective

1 papers0 benchmarksImages

Surgical Hands

Surgical Hands is a dataset that provides multi-instance articulated hand pose annotations for in-vivo videos. The dataset contains 76 video clips from 28 publicly available surgical videos and over 8.1k annotated hand pose instances.

1 papers0 benchmarksVideos

A Dataset of Multispectral Potato Plants Images

The dataset contains aerial agricultural images of a potato field with manual labels of healthy and stressed plant regions. The images were collected with a Parrot Sequoia multispectral camera carried by a 3DR Solo drone flying at an altitude of 3 meters. The dataset consists of RGB images with a resolution of 750×750 pixels, and spectral monochrome red, green, red-edge, and near-infrared images with a resolution of 416×416 pixels, and XML files with annotated bounding boxes of healthy and stressed potato crop.

1 papers10 benchmarksImages

Pyxis

Pyxis is a performance dataset for specialized accelerators on sparse data. Pyxis collects accelerator designs and real execution performance statistics. Currently, there are 73.8 K instances in Pyxis.

1 papers0 benchmarks

Persian Reverse Dictionary Dataset

The Persian Reverse Dictionary Dataset is a collection of 855217 words along with the phrases describing them. The phrases were extracted from the top three most well-known Persian dictionaries (including Amid, Moeen, and Dehkhoda), Persian Wikipedia, and a Persian Wordnet (called Farsnet).

1 papers0 benchmarks

RWanda Built-up Region Segmentation

We create Rwanda built-up regions dataset, a different and versatile in nature from previously available datasets. The varying structure size and formation, irregular patterns of construction, buildings in forests and deserts, and the existence of mud houses make it very challenging. A total of 787 satellite images of size 256 × 256 are collected at a high resolution (HR) of 1.193 meters per pixel and hand tagged for built-up region segmentation using an online tool Label-Box.

1 papers0 benchmarks

Mouse Grooming Behavior

This dataset was generated to characterize mouse grooming behavior. Mouse grooming serves many adaptive functions such as coat and body care, stress reduction, de-arousal, social functions, thermoregulation, nociception, as well as other functions. Alteration of this behavior is measured and used for mouse pre-clinical models of human psychiatric illnesses.

1 papers0 benchmarksVideos

PQ-decaNLP (Paraphrase Questions - decaNLP)

Multitask learning has led to significant advances in Natural Language Processing, including the decaNLP benchmark where question answering is used to frame 10 natural language understanding tasks in a single model. PQ-decaNLP is a crowd-sourced corpus of paraphrased questions, annotated with paraphrase phenomena. This enables analysis of how transformations such as swapping the class labels and changing the sentence modality lead to a large performance degradation.

1 papers0 benchmarksTexts

DrugProt

DrugProt corpus, where domain experts have exhaustively labeled:(a) all chemical and gene mentions, and (b) all binary relationships between them corresponding to a specific set of biologically relevant relation types (DrugProt relation classes).

1 papers1 benchmarks

Building air quality and pandemic risk simulation

The original paper contains a high-level explanation of the dataset characteristics, and potential use cases of the dataset. ArchABM can help to quantify the impact of some of these building- and company policy-related measures.

1 papers0 benchmarksGraphs, Time series

DriverMHG

Driver Micro Hand Gestures (DriverMHG) is a dataset for dynamic recognition of driver micro hand gestures, which consists of RGB, depth and infrared modalities.

1 papers0 benchmarksVideos

AVASpeech-SMAD (AVASpeech-SMAD: A Strongly Labelled Speech and Music Activity Detection Dataset with Label Co-Occurrence)

We propose a dataset, AVASpeech-SMAD, to assist speech and music activity detection research. With frame-level music labels, the proposed dataset extends the existing AVASpeech dataset, which originally consists of 45 hours of audio and speech activity labels. To the best of our knowledge, the proposed AVASpeech-SMAD is the first open-source dataset that features strong polyphonic labels for both music and speech. The dataset was manually annotated and verified via an iterative cross-checking process. A simple automatic examination was also implemented to further improve the quality of the labels. Evaluation results from two state-of-the-art SMAD systems are also provided as a benchmark for future reference.

1 papers0 benchmarksAudio, Music, Speech

Earth’s Mantle Convection

The dataset, generated from a scientific simulation, consists of a time series (251 steps) of 3D scalar fields on a spherical 180x201x360 grid covering 500 Myr of geological time. Each time step is 2 Myrs, and the fields are:

1 papers0 benchmarks3D, Time series

GO21

GO21 is a biomedical knowledge graph that models genes, proteins, drugs, and the hierarchy of the biological processes they participate in. It consists of 806,136 triples with 21 relations and 89127 entities. GO21 can be used for knowledge graph completion tasks (link prediction) as well as hierarchical reasoning tasks, such as ancestor-descendant prediction task proposed in the paper.

1 papers4 benchmarksBiology, Graphs
PreviousPage 410 of 1000Next