TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Bangla Word Analogy

We provide a Mikolov-style word-analogy evaluation set specifically for Bangla, with a sample size of 16678, as well as a translated and curated version of the Mikolov dataset, which contains 10594 samples for cross-lingual research.

1 papers0 benchmarks

WebBrain-Raw

WebBrain-Raw is a large-scale dataset built from English Wikipedia articles and their crawlable Wikipedia references. It comprises 153 zipped data chunks in which each line is a Wikipedia page with its reference articles.

1 papers0 benchmarks

IAW Dataset (Ikea Assembly In The Wild Dataset)

The IAW dataset contains 420 Ikea furniture pieces from 14 common categories e.g. sofa, bed, wardrobe, table, etc. Each piece of furniture comes with one or more user instruction manuals, which are first divided into pages and then further divided into independent steps cropped from each page (some pages contain more than one step and some pages do not contain instructions). There are 8568 pages and 8263 steps overall, on average 20.4 pages and 19.7 steps for each piece of furniture. We crawled YouTube to find videos corresponding to these instruction manuals and as such the conditions in the videos are diverse on many aspects e.g. duration, resolution, first- or third-person view, camera pose, background environment, number of assemblers, etc. The IAW dataset contains 1005 raw videos with a length of around 183 hours in total. Among them, approximately 114 hours of content are labeled as 15649 actions to match the corresponding step in the corresponding manual.

1 papers0 benchmarksImages, Videos

FewDR

FewDR is a dataset for Few-shot dense retrieval (DR). FewDR aims to effectively generalize to novel search scenarios by learning a few samples. Specifically, FewDR employs class-wise sampling to establish a standardized "few-shot" setting with finely-defined classes, reducing variability in multiple sampling rounds.

1 papers0 benchmarksTexts

FollowMe Vehicle Behaviour Prediction Dataset

This dataset is a result of a study that was created to assess drivers behaviors when following a lead vehicle. The driving simulator study used a simulated suburban environment for collecting driver behavior data while following a lead vehicle driving through various unsignalized intersections. The driving environment had two lanes in each direction and a dedicated left-turn lane for the intersection. The experiment was deployed on a miniSim Driving Simulator. We programmed the lead vehicle ran- domly turn left, right or go straight through the intersections. In total we had 2(traffic density) × 2(speed level) × 3 = 12 scenarios for each participant to be tested on. We split the data into train, validation and test sets. The setup for the task is to observe 1 second of trajectories and predict the next 3,5 and 8 seconds.

1 papers0 benchmarks

L1BSR (L1BSR dataset)

The Sentinel-2 satellite carries 12 CMOS detectors for the VNIR bands, with adjacent detectors having overlapping fields of view that result in overlapping regions in level-1 B (L1B) images. This dataset includes 3740 pairs of overlapping image crops extracted from two L1B products. Each crop has a height of around 400 pixels and a variable width that depends on the overlap width between detectors for RGBN bands, typically around 120-200 pixels. In addition to detector parallax, there is also cross-band parallax for each detector, resulting in shifts between bands. Pre-registration is performed for both cross-band and cross-detector parallax, with a precision of up to a few pixels (typically less than 10 pixels).

1 papers0 benchmarksImages, Stereo

PGDataset (Profile Generation Dataset)

PGDataset (Profile Generation Dataset) is a dataset created for the PGTask (Profile Generation Task), where the goal is to extract/generate a profile sentence given a dialogue utterance.

1 papers8 benchmarksDialog, Texts

XWikiRef

We provide a new data set XWikiRef for the task of Cross-lingual Multi-document Summarization. This task aims at generating Wikipedia style text in Low Resource languages by taking reference text as input. Overall, the data set contains 8 different languages: bengali (bn), english (en), hindi (hi), marathi (mr), malayalam (ml), odia (or), punjabi (pa) and tamil (ta). It also contains 5 domains: books, films, politicians, sportsman and writers.

1 papers0 benchmarksTexts

YIM Dataset (Yeast Cells in Microstructures Dataset)

An instance segmentation dataset of yeast cells in microstructures. The dataset includes 493 densely annotated microscopy images. For more information see the paper "An Instance Segmentation Dataset of Yeast Cells in Microstructures".

1 papers0 benchmarksBiology, Images, Medical

Trajectory calibration experiments

Data and experiments for motion-based extrinsic calibration using trajectory_calibration. The data was generated for evaluating hand-eye calibration algorithms in extrinsic sensor-to-sensor calibration.

1 papers0 benchmarks6D

Brazilian E-Commerce Public Dataset by Olist

See https://www.kaggle.com/datasets/olistbr/brazilian-ecommerce .

1 papers0 benchmarks

iiwa Robotic Arm Reconstruction Dataset

Please see our website and code repository for detailed description.

1 papers0 benchmarks3d meshes, Videos

CKBP v2

CKBP v2 is a new CSKB Population benchmark, which addresses the two mentioned problems by using experts instead of crowd-sourced annotation and by adding diversified adversarial samples to make the evaluation set more representative.

1 papers0 benchmarksTexts

MLRegTest (A Benchmark for the Machine Learning of Regular Languages)

MLRegTest is a benchmark for sequence classification, containing training, development, and test sets from 1,800 regular languages. Regular languages are formal languages, which are sets of sequences definable with certain kinds of formal grammars, including regular expressions, finite-state acceptors, and monadic second-order logic with either the successor or precedence relation in the model signature for words. This benchmark was designed to help identify those factors, specifically the kinds of long-distance dependencies, that can make it difficult for ML systems to generalize successfully in learning patterns over sequences. MLRegTest organizes its languages according to their logical complexity (monadic second-order, first-order, propositional, or monomial expressions) and the kind of logical literals (string, tier-string, subsequence, or combinations thereof). The logical complexity and choice of literal provides a systematic way to understand different kinds of long-distance de

1 papers0 benchmarks

MIMIC-IV ICD-9

MIMIC-IV ICD-9 contains 209,326 discharge summaries—free-text medical documents—annotated with ICD-9 diagnosis and procedure codes. It contains data for patients admitted to the Beth Israel Deaconess Medical Center emergency department or ICU between 2008-2019. All codes with fewer than ten examples have been removed, and the train-val-test split was created using multi-label stratified sampling. The dataset is described further in Automated Medical Coding on MIMIC-III and MIMIC-IV: A Critical Review and Replicability Study, and the code to use the dataset is found here.

1 papers18 benchmarksTexts

KnowledJe

We introduce KnowledJe, an English-language knowledge graph of antisemitic history and language from the 20th century to the present. Structured as a JSON file, KnowledJe currently contains 618 entries, which consist of 210 event names, 137 place names, 95 person names, 80 dates (years), 38 publication names, 27 organization names, and 1 product name. Each entry is associated with its own dictionary, which contains descriptions, locations, authors, and dates as applicable. We obtain the entries through four Wikipedia articles: “Timeline of antisemitism in the 20th century,” “Timeline of antisemitism in the 21stcentury,” the “Jews” section of “List of religious slurs,” and “Timeline of the Holocaust.” To obtain descriptions for each applicable key, we used the following general rules: 1. If the concept associated with the key is a slur, the description is the entry in the “Meaning, origin, and notes” column of the “List of religious slurs” article. 2. Otherwise, if the concept associate

1 papers0 benchmarksTexts

EchoKG

Echo Corpus (Arviv et al, 2021) infused with information from KnowledJe (Halevy, 2023). Algorithm detailed in Algorithm 1 in Section 3.2 of Halevy (2023) (https://arxiv.org/pdf/2304.11223.pdf). 4,630 total tweet samples, 380 labeled as antisemitic hate speech. Files included in https://github.com/enscma2/knowledje and described in its README.md. Potential use cases: detection of antisemitic hate speech.

1 papers0 benchmarksTexts

Echo Corpus

A large dataset of over 18,000,000 English tweets posted by ∼7K echo users was constructed in the following manner: 1. Base Corpus We have obtained access to a random sample of 10% of all public tweets posted in May and June 2016 – the peak use of the echo. 2. Raw Echo Corpus Searching the base corpus, we extracted all tweets containing the echo symbol, resulting in 803,539 tweets posted by 418,624 users. Filtering out non-English Tweets and users who used the echo less than three times we were left with ∼7K users. 3. Echo Corpus We used Twitter API to obtain the most recent tweets (up to 3.2K) of each of the users remainingin the English list. This process resulted in ∼18M tweets posted by 7,073 users. Some of the accounts we found using the echo were already suspended or deleted at the time of collection, thus their tweets were not retrievable. Relevant footnotes: - The echo is found in tweets written in multiple languages, particularly in East-Asian languages of which the user base

1 papers0 benchmarksTexts

MJFF Levodopa Response Study

The data generated from this study are grouped into 3 main types: (1) participant demographic and clinical data, (2) sensor data from the different devices, as well as clinical scores and metadata related to the tasks performed, and (3) participant diaries collected during the in-clinic and at-home phases of the study. Throughout the data tables, timestamps are provided as UNIX epoch/POSIX time.

1 papers0 benchmarksMedical, Time series

Dataset for analyzing the impact of gamification in software testing education

Data collected for the controlled experiment performed to analyze whether gamification can help in software testing education. The results are reported in "Can gamification help in software testing education? Findings from an empirical study" (DOI: 10.1016/j.jss.2023.111647).

1 papers0 benchmarks
PreviousPage 459 of 1000Next