TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Moving Symbols

A parameterized synthetic dataset called Moving Symbols to support the objective study of video prediction networks.

1 papers0 benchmarks

MultiSenti

MultiSenti presents a labeled dataset called MultiSenti for sentiment classification of code-switched informal short text, (2) explore the feasibility of adapting resources from a resource-rich language for an informal one, and (3) propose a deep learning-based model for sentiment classification of code-switched informal short text.

1 papers0 benchmarks

Multi-species fruit flower detection datasets

This dataset consists of four sets of flower images, from three different species: apple, peach, and pear, and accompanying ground truth images. The images were acquired under a range of imaging conditions. These datasets support work in an accompanying paper that demonstrates a flower identification algorithm that is robust to uncontrolled environments and applicable to different flower species. While this data is primarily provided to support that paper, other researchers interested in flower detection may also use the dataset to develop new algorithms. Flower detection is a problem of interest in orchard crops because it is related to management of fruit load.

1 papers0 benchmarks

MultiWOZ-coref

MultiWOZ-coref, (or MultiWOZ 2.3) is an extension of the MultiWOZ dataset that adds co-reference annotations in addition to corrections of dialogue acts and dialogue states.

1 papers0 benchmarksTexts

MVS1K

Contains about 1, 000 videos from 10 queries and their video tags, manual annotations, and associated web images.

1 papers0 benchmarksVideos

MWE-CWI

Multiword expressions (MWEs) represent lexemes that should be treated as single lexical units due to their idiosyncratic nature. MWE-CWI is a dataset for MWE detection based on the Complex Word Identification Shared Task 2018 dataset.

1 papers0 benchmarksTexts

NAIST COVID

NAIST COVID is a multilingual dataset of social media posts related to COVID-19, consisting of microblogs in English and Japanese from Twitter and those in Chinese from Weibo. The data cover microblogs from January 20, 2020, to March 24, 2020.

1 papers0 benchmarks

NatCat

A general purpose text categorization dataset (NatCat) from three online resources: Wikipedia, Reddit, and Stack Exchange. These datasets consist of document-category pairs derived from manual curation that occurs naturally by their communities.

1 papers0 benchmarks

NavigationNet

NavigationNet is a computer vision dataset and benchmark to allow the utilization of deep reinforcement learning on scene-understanding-based indoor navigation.

1 papers0 benchmarksImages

PSU NRTDB (PSU Near-Regular Texture Database)

The PSU Near-Regular Texture Database is a texture dataset. It covers the spectrum of textures from completely regular to near-regular to irregular. It also includes video of near-regular textures in motion. The database also contains, or will include, test image sets with ground-truth for translation, rotation, reflection/glide-reflection symmetry detection algorithms.

1 papers0 benchmarksImages

NEEQ Annual Reports

Business taxonomies automatically constructed from the content of corporate annual reports.

1 papers0 benchmarks

Neural Conversational QA

A modified data-set that has fewer spurious patterns than the original data-set, consequently allowing models to learn better.

1 papers0 benchmarks

New Brown Corpus

A new dataset for training and evaluating grounded language models.

1 papers0 benchmarks

NLI-PT

The first Portuguese dataset compiled for Native Language Identification (NLI), the task of identifying an author's first language based on their second language writing. The dataset includes 1,868 student essays written by learners of European Portuguese, native speakers of the following L1s: Chinese, English, Spanish, German, Russian, French, Japanese, Italian, Dutch, Tetum, Arabic, Polish, Korean, Romanian, and Swedish. NLI-PT includes the original student text and four different types of annotation: POS, fine-grained POS, constituency parses, and dependency parses. NLI-PT can be used not only in NLI but also in research on several topics in the field of Second Language Acquisition and educational NLP.

1 papers0 benchmarksTexts

NoReC (fine-grained)

NoReC_fine is a dataset for fine-grained sentiment analysis in Norwegian, annotated with respect to polar expressions, targets and holders of opinion.

1 papers0 benchmarks

Noun-Noun Compound Dataset

The noun–noun compounds dataset created by Fares (2016) consists of compounds annotated with two different taxonomies of relations; that is, for each noun–noun compound there are two distinct relations, drawing on different linguistic schools. The dataset was derived from existing linguistic resources, such as NomBank (Meyers et al., 2004) and the Prague Czech-English Dependency Treebank 2.0 (PCEDT; Hajič et al., 2012).

1 papers0 benchmarks

NREC Agricultural Person-Detection

A dataset to encourage research in these environments. It consists of labeled stereo video of people in orange and apple orchards taken from two perception platforms (a tractor and a pickup truck), along with vehicle position data from RTK GPS.

1 papers0 benchmarksImages

NTPairs (News-Tweet Paired Dataset)

The NTPairs dataset consists of the pairs of news articles and their corresponding tweets that were published by eight media outlets in 2018. The eight outlets were selected to consider diverse outlets, which employ a different editing style for news sharing, in terms of publishing channels and political leaning.

1 papers0 benchmarks

Numeric Fused-Head

The Numeric Fused-Head dataset consists of ~10K examples of crowd-sourced classified examples, labeled into 7 different categories, from two types. In the first type, Reference, the missing head is referenced explicitly somewhere else in the discourse, either in the same sentence or in surrounding sentences. In the second type, Implicit, the missing head does not appear in the text and needs to be inferred by the reader or hearer based on the context or world knowledge. This category was labeled into the 6 most common categories of the dataset. Models are evaluated based on accuracy.

1 papers0 benchmarks

O4B (Open4Business)

O4B is a dataset of 17,458 open access business articles and their reference summaries. The dataset introduces a new challenge for summarization in the business domain, requiring highly abstractive and more concise summaries as compared to other existing datasets.

1 papers0 benchmarks
PreviousPage 373 of 1000Next