TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

SuperLim

The SuperLim dataset is a Swedish version of the English benchmarking platform (Super)GLUE. It forms the basis of a national testbed for Swedish language models. The project aims to provide a standardized collection of benchmarking tests for Swedish language models, supporting the development of trustworthy and robust Natural Language Processing (NLP) applications.

0 papers0 benchmarks

DanishPoliticalComments

The DanishPoliticalComments dataset is a collection of sentences that are labeled with fine-grained polarity in the range from -2 to 2 (negative to positive). The dataset is primarily used for tasks related to text classification, specifically multi-class classification. The dataset is monolingual and is in the Danish language. It falls under the size category of 1K<n<10K. The language creators are listed as 'other' and the annotations are expert-generated.

0 papers0 benchmarks

LCC

This dataset is suitable for sentiment analysis, which consists of Danish data from the Leipzig Collection. The collection and annotation of the dataset are solely due to Finn Årup Nielsen. It was originally annotated as a score between -5 and +5, but the labels in this version have been converted to negative, neutral, and positive labels.

0 papers0 benchmarks

bioRxiv

bioRxiv is a free online archive for unpublished preprints in the life sciences. It allows researchers to share their findings with the scientific community and receive feedback before it undergoes peer review for formal publication. Datasets on bioRxiv are typically associated with these preprints.

0 papers0 benchmarks

Shared Task on Hierarchical Classification of Blurbs - GermEval 2019

This dataset can be used as a benchmark for clustering word embeddings for German. It contains 18'084 unique samples, 28 splits with 177 to 16'425 samples, and 4 to 93 unique classes.

0 papers0 benchmarks

medRxiv

medRxiv (pronounced "med-archive") is an Internet site distributing unpublished e-prints about health sciences. It distributes complete but unpublished manuscripts in the areas of medicine, clinical research, and related health sciences without charge to the reader. Such manuscripts have yet to undergo peer review and the site notes that preliminary status and that the manuscripts should not be considered for clinical application, nor relied upon for news reporting as established information.

0 papers0 benchmarks

SemEval-2015 Task-1 (Paraphrase and Semantic Similarity in Twitter)

Given two sentences, the participants are asked to determine whether they express the same or very similar meaning and optionally a degree score between 0 and 1. Following the literature on paraphrase identification, we evaluate system performance primarily by the F-1 score and Accuracy against human judgments. We also provide additional evaluations by Pearson correlation and PINC (Chen and Dolan, 2011), which measure lexical dissimilarity between sentence pairs.

0 papers0 benchmarks

LinkSO

The LinkSO dataset is a resource for learning to retrieve similar question-answer pairs on Stack Overflow. It consists of three datasets corresponding to three popular programming languages (Python, Java, JavaScript), 690K question pairs, and 26K linked question pairs (i.e., positive examples). The dataset was extracted from Stack Overflow's data dump in April 2018 and was cleaned and pre-processed to remove non-ASCII characters, email addresses, URLs, and code blocks. The dataset was designed to help propose new models, such as neural network models, to improve community-based question-answer retrieval in the software engineering domain.

0 papers0 benchmarks

SweFAQ

The SweFAQ dataset is a collection of frequently asked questions from Swedish authorities' websites with shuffled answers. It was created by Aleksandrs Berdicevskis and is published by Språkbanken Text. The dataset is a part of the SuperLim collection.

0 papers0 benchmarks

cMedQA

The cMedQA dataset is designed for Chinese community medical question answering. It has two versions:

0 papers0 benchmarks

QBQTC (QQ Browser Query Title Corpus)

The QQ Browser Query Title Corpus (QBQTC) is a large-scale dataset constructed for search scenarios by the QQ Browser search engine. It integrates dimensions such as relevance, authority, content quality, and timeliness, and is widely used in search engine business scenarios.

0 papers0 benchmarks

Multilingual Sentiment Datasets

A collection of multilingual sentiment datasets grouped into 3 classes -- positive, neutral, and negative.

0 papers0 benchmarks

SUC (Stockholm-Umeå Corpus)

The Stockholm-Umeå Corpus (SUC) is a collection of Swedish texts from the 1990s, consisting of one million words in total. The corpus is balanced, meaning that it contains various text types and stylistic levels. The texts are annotated with part-of-speech tags, morphological analysis, and lemma (all that can be considered gold standard data), as well as some structural and functional information.

0 papers0 benchmarks

WMT 2021 - Machine Translation of News

WMT21 (Workshop on Machine Translation 2021) Translation Task focuses on news text translation. It includes language pairs such as English to/from various languages like Chinese, Czech, German, Hausa, Icelandic, Japanese, Russian, and more. Goals include investigating current MT techniques for languages other than English, challenges in translating between language families, translation of low-resource languages, and creating publicly available corpora for MT evaluation. The dataset provides parallel corpora for all languages and additional resources for download, with a focus on machine translation of news.

0 papers0 benchmarks

TICO-19

The TICO-19 dataset is a translation initiative focused on COVID-19 content, created by academic and industry partners along with Translators without Borders. It includes translation memories, translated terminologies for COVID-19-related terms, and a benchmark dataset. The benchmark comprises 30 documents, translating 3071 sentences (69.7k words) from English into 36 languages. The effort aims to assist professional translators and machine translation research, emphasizing emergency and crisis-related content availability in multiple languages.

0 papers0 benchmarks

UWB_Data

Dataset for range data gathered using the Ultra-wide Band (UWB) MDEK 1001 Dev. kit in 3 different experimental scenarios. The dataset description document provides the details of experimentation and the dataset.

0 papers0 benchmarks

global-weather-repository

This dataset provides daily weather information for capital cities around the world. Unlike forecast data, this dataset offers a comprehensive set of features that reflect the current weather conditions around the world. Starting from August 29, 2023. It provides over 40+ features , including temperature, wind, pressure, precipitation, humidity, visibility, air quality measurements and more. The dataset is valuable for analyzing Global weather patterns, exploring climate trends, and understanding the relationships between different weather parameters.

0 papers0 benchmarks

Events classification - Biotech news

A dataset specifically tailored to the biotech news sector, aiming to transcend the limitations of existing benchmarks. This dataset is rich in complex content, comprising various biotech news articles covering various events, thus providing a more nuanced view of information extraction challenges.

0 papers0 benchmarksTexts

CLEF-TAR

Technologically Assisted Reviews in Empirical Medicine.

0 papers0 benchmarks

Micro-Ultrasound Prostate Segmentation Dataset

This dataset comprises micro-ultrasound scans and human prostate annotations of 75 patients who underwent micro-ultrasound guided prostate biopsy at the University of Florida. All images and segmentations have been fully de-identified in the NIFTI format.

0 papers0 benchmarks
PreviousPage 656 of 1000Next