TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Gambling Address Dataset

Gambling Address Dataset is a collection of 10,423 gambling addresses that have transactions with gambling contracts. Moreover, 51,004 non-gambling addresses are also selected (such as exchanges, wallet addresses, etc.), making the gambling address dataset more complete. In the dataset, accounts are used to refer to addresses (e.g. 0xd1ce...edec95), where 1, 0, and -1 represent the gamble, non-gamble, and other types, respectively.

1 papers0 benchmarksTexts

Gambling Contract Dataset

Gambling Contract Dataset is a collection of 260 gambling smart contracts from decentralized gambling websites, such as Dicether, Degens. At the same time, in order to construct the negative samples required for training, 1040 smart contracts that are not involved in gambling (e.g., erc20, erc721, mixer, etc.) are selected . In the dataset, accounts are used to refer to contracts (e.g. 0x3fe2b...f8a33f), where 1, 0, and -1 to represent the gamble, non-gamble, and other types, respectively.

1 papers0 benchmarksTexts

BeGin

BeGin provides 23 benchmark scenarios for graph from 14 real-world datasets, which cover 12 combinations of the incremental settings and the levels of problem. In addition, BeGin provides various basic evaluation metrics for measuring the performances and final evalution metrics designed for continual learning.

1 papers0 benchmarksGraphs

DeepParliament

DeepParliament is a legal domain Benchmark Dataset that gathers bill documents and metadata and performs various bill status classification tasks. The dataset text covers a broad range of bills from 1986 to the present and contains richer information on parliament bill content. There are a total of 5329 documents where 4223 are in the train and 1106 are in the test dataset. Each bill document contains many sentences in both cases, and the document’s length varies greatly.

1 papers0 benchmarksTexts

NEREL-BIO

NEREL-BIO is an annotation scheme and corpus of PubMed abstracts in Russian and English. It contains annotations for 700+ Russian and 100+ English abstracts. All English PubMed annotations have corresponding Russian counterparts. NEREL-BIO comprises the following specific features: annotation of nested named entities, it can be used as a benchmark for cross-domain (NEREL -> NEREL-BIO) and cross-language (English -> Russian) transfer.

1 papers0 benchmarksTexts

KurdishInterdialect (Kurdish parallel corpus in Kurmanji, Sorani and English)

a parallel corpus of Sorani (ckb or Central Kurdish) and Kurmanji (kmr or Northern Kurdish) dialects of Kurdish along with English (eng). The parallel corpus contains three manually-aligned corpus in Sorani-Kurmanji, Sorani-English and Kurmanji-English in various formats, namely Translation Memory eXchange file format (.tmx), parallel annotated text useful for ParaConc and raw parallel texts (.txt). This corpus contains 12,327 translation pairs in the two major dialects of Kurdish, Sorani and Kurmanji. We also provide 1,797 and 650 translation pairs in English-Kurmanji and English-Sorani.

1 papers0 benchmarks

ZazaGoraniCorpus (A corpus of Zazaki and Gorani corpus)

A corpus for two endangered languages of the Zaza-Gorani language family: Zazaki and Gorani.

1 papers0 benchmarks

SSF_dataset

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

A Bi-atrial Statistical Shape Model and 100 Volumetric Anatomical Models of the Atria

This dataset is part of the publication "A bi-atrial statistical shape model for large-scale in silico studies of human atria: Model development and application to ECG simulations" by Nagel et al. (https://doi.org/10.1016/j.media.2021.102210). It includes a bi-atrial statistical shape model built based on 47 MR and CT images (Left atrium segmentation challenge (Tobon-Gomez, 2015), Left atrium fibrosis and scar segmentation challenge (Karim, 2013), Left atrial wall thickness challenge (Karim, 2018)). ScalismoLab (https://scalismo.org) was used for parts of the model generation. Further Details are explained in the paper. The SSM is available as an h5 file including information about the mean shape's vertex locations and their triangulation as well as the eigenvectors and -values.

1 papers0 benchmarks

MedalCare-XL

Mechanistic cardiac electrophysiology models allow for personalized simulations of the electrical activity in the heart and the ensuing electrocardiogram (ECG) on the body surface. As such, synthetic signals possess precisely known ground truth labels of the underlying disease (model parameterization) and can be employed for validation of machine learning ECG analysis tools in addition to clinical signals. Recently, synthetic ECG signals were used to enrich sparse clinical data for machine learning or even replace them completely during training leading to good performance on real-world clinical test data.

1 papers0 benchmarks

weather4cast 2022

Weather4Cast 2022 satellite images for weather prediction.

1 papers0 benchmarks

Apron Dataset

The Apron Dataset focuses on training and evaluating classification and detection models for airport-apron logistics. In addition to bounding boxes and object categories the dataset is enriched with meta parameters to quantify the models’ robustness against environmental influences.

1 papers0 benchmarksImages

jaCappella

jaCappella is a corpus of Japanese a cappella vocal ensembles (jaCappella corpus) for vocal ensemble separation and synthesis. It consists of 35 copyright-cleared vocal ensemble songs and their audio recordings of individual voice parts. These songs were arranged from out-of-copyright Japanese children's songs and have six voice parts (lead vocal, soprano, alto, tenor, bass, and vocal percussion). They are divided into seven subsets, each of which features typical characteristics of a music genre such as jazz and enka.

1 papers0 benchmarksMusic

MTC

MTC is a financial-domain dataset of the multi-label topic classification task. It aims to identify the topics of the spoken dialogue.

1 papers0 benchmarksTexts

PSM

PSM is a financial-domain dataset of the pairwise search matching task. It aims to identify the semantic similarity of a sentence pair in the search scenario.

1 papers0 benchmarksTexts

IEE

IEE is a financial-domain dataset of the Insurance-entity extraction task. Its goal is to locate named entities mentioned in the input sentence.

1 papers0 benchmarksTexts

UMD-i Affrodance Dataset

One-Shot Affordance Part Segmentation variant of the UMD dataset. Each object instance in the dataset contains a single image.

1 papers0 benchmarksImages, RGB-D

XBT Snapshot (XBT Snapshot from World Ocean Database)

The World Ocean Database (WOD) is world's largest collection of uniformly formatted, quality controlled, publicly available ocean profile data. This dataset is a snapshot of the XBT observations which have been preprocessed for use in a machine learning pipeline.

1 papers0 benchmarks

PIZZA

PIZZA is a dataset for parsing pizza and drink orders, whose semantics cannot be captured by flat slots and intents.

1 papers0 benchmarksTexts

MatSim (MatSim dataset for materials similarity recognition from images)

MatSim is a synthetic dataset, and natural image benchmark for computer vision-based recognition of similarities and transitions between materials and textures, focusing on identifying any material under any conditions using one or a few examples (one-shot learning), including materials states and subclasses.

1 papers0 benchmarksImages
PreviousPage 448 of 1000Next