TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Supplementary Material (Annotation Table of Review)

The file contains an annotated list of papers that are included in the literature survey.

1 papers0 benchmarksTabular

BioFuelQR

BioFuelQR is a dataset consisting of complex reasoning questions related to catalyst discovery in biofuels. This dataset is aimed at benchmarking scientific question answering methods, particularly for search based text generation.

1 papers0 benchmarksTexts

DotPrompts

DotPrompts is a set of testcases derived from PragmaticCode, such that each testcase consists of a prompt to a dereference location (a code location having the "." operator in Java). It is primarily meant as a benchmark for Code LMs.

1 papers1 benchmarks

Quasimodo-GenT

A mapping of Quasimodo to the relations of ConceptNet.

1 papers0 benchmarks

Ascent-GenT

A mapping of Ascent to the relations of ConceptNet.

1 papers0 benchmarks

The EMBO SourceData-NLP dataset (The SourceData-NLP dataset: integrating curation into scientific publishing for training large language models)

We present the SourceData-NLP dataset produced through the routine curation of papers during the publication process. A unique feature of this dataset is its emphasis on the annotation of bioentities in figure legends. We annotate eight classes of biomedical entities (small molecules, gene products, subcellular components, cell lines, cell types, tissues, organisms, and diseases), their role in the experimental design, and the nature of the experimental method as an additional class. SourceData-NLP contains more than 620,000 annotated biomedical entities, curated from 18,689 figures in 3,223 papers in molecular and cell biology. We illustrate the dataset's usefulness by assessing BioLinkBERT and PubmedBERT, two transformers-based models, fine-tuned on the SourceData-NLP dataset for NER. We also introduce a novel context-dependent semantic task that infers whether an entity is the target of a controlled intervention or the object of measurement.

1 papers4 benchmarksBiology, Biomedical, Texts

reader_engagement

Reader eye tracking and engagement scores for two short stories, aggregated by sentence.

1 papers0 benchmarksTexts

Race Against the Machine

A fully-annotated, open-design dataset of autonomous and piloted high-speed flight

1 papers0 benchmarks

WinoMT-Hindi

Test set of sentences in Hindi with complex coreference involving two entities inspired by WinoBias format of sentences in English. Includes grammatical gender cues of Hindi to test gender bias in Hindi-English NMT Systems.

1 papers0 benchmarksTexts

OTSC-Hindi (Occupation Test Set with Simple Context - Hindi)

Test set of sentences in Hindi with simple gender-specific context used to measure gender bias in NMT systems for Hindi-English.

1 papers0 benchmarksTexts

Alpaca Data Galician

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksTexts

https://ieee-dataport.org/open-access/urb3dcd-urban-point-clouds-simulated-dataset-3d-change-detection

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

Polyp ASH

This dataset was built with data acquired at the Hospital Clinic of Barcelona, Spain. It is composed of a total of 1126 HD polyp images. There are a total of 473 unique polyps, with a variable number of different shots per polyp (minimum: 2, maximum: 24, median: 10). Special attention was paid to ensure that images from the same polyp show different conditions. An external frame-grabber and a white light endoscope were used to capture raw images. The dataset contains images with two different resolutions: 1920 x 1080 and 1350 x 1080.

1 papers0 benchmarksImages

FoodSG-233 (Localized Singaporean Food Image Dataset)

The FoodSG-233 dataset contains 209,861 images, covering 13 food groups and 233 food categories.

1 papers0 benchmarks

WebRTC QoE Estimation Without Using Application Layer Headers

This dataset contains network traces collected in-lab and in a real-world setting. We also collected ground truth Quality of Experience (QoE) logs from the browser. Researchers can use this dataset to solve a wide range of problems that involve understanding video conferencing quality from a network perspective.

1 papers0 benchmarks

Glot500-c (Glot500 Corpus)

A dataset of natural language data collected by putting together more than 150 existing mono-lingual and multilingual datasets together and crawling known multilingual websites. The focus of this dataset is on 500 extremely low-resource languages.

1 papers0 benchmarks

Unsynchronized Dynamic Blender Dataset

Unsynchronized dynamic blender dataset for multi-view dynamic NeRFs for evaluating MAE between the predicted time offsets and the ground truth. It contains three unsyncrhonized scenes; box, deer, and fox.

1 papers0 benchmarks

Randman

Versatile synthetic classification dataset based on precise input spike timings drawn from smooth random manifolds as previously described

1 papers0 benchmarks

Faces Through Time

Faces Through Time (FTT) features 26,247 images of notable people from the 19th to 21st centuries, with roughly 1,900 images per decade on average. It is sourced from Wikimedia Commons, a crowdsourced and open-licensed collection of 50M images.

1 papers0 benchmarksImages

Turkish Punctuation Restoration

we have prepared a dataset using publicly available TED Talks transcripts [27] and selected the Turkish corpus. The resulting Turkish punctuation restoration dataset currently consists of 146K sentences and 1.8M tokens. The ratio of the train, validation, and test splits are 0.8, 0.1, and 0.1, respectively. Data files contain two columns. The first column has the tokens separated by white space. The second column includes tags for each token.

1 papers0 benchmarks
PreviousPage 478 of 1000Next