TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

VISO (VIdeo Satellite Objects)

This dataset is a large-scale dataset for moving object detection and tracking in satellite videos, which consists of 40 satellite videos captured by Jilin-1 satellite platforms. Each image has a resolution of 12000x5000 and contains a great number of objects with different scales. Four common types of vechicles, including plane, car, ship, and train, are manually-labeled. A total of 853,911 instances are labeled by axis-aligned bounding boxes. https://paperswithcode.com/paper/detecting-and-tracking-small-and-dense-moving

1 papers0 benchmarks

Bistatic MIMO Radar Sensing of Specularly Reflecting Surfaces for Wireless Power Transfer

The measurement data <b>VNA_20220722_232002_XETS_reduced.mat</b> includes a data matrix $\mathbf{R}$ acquired with a synthetic aperture measurement testbed described in [2] and [3]. Measured were $N_f=1000$ frequency steps in a band from $3-10$GHz of the scattering parameter $S_{21}$ between a synthetic $51$-ULA with antenna positions saved in the file <b>ULA.mat</b> and a synthetic $(13\times 13)$-URA with antenna positions saved in the file <b>URA.mat</b>. The file <b>XETSantennaCharacterization.mat</b> holds antenna gains of an XETS antenna [4] characterized in an anechoic chamber. XETS antennas were used on both the ULA (oriented towards the negative $x$-direction) and on the URA (oriented towards the positive $x$-direction).<br> These data have been used to perform ultra-wideband (UWB) bistatic radar imaging and wireless power transfer (WPT). Our implementation is provided in <b>MAIN_wall_detection.mat</b>.

1 papers0 benchmarks

WikiWeb2M (Wikipedia Webpage 2M)

Wikipedia Webpage 2M (WikiWeb2M) is a multimodal open source dataset consisting of over 2 million English Wikipedia articles. It is created by rescraping the ∼2M English articles in WIT. Each webpage sample includes the page URL and title, section titles, text, and indices, images and their captions.

1 papers0 benchmarksImages, Texts

ParsVQA-Caps

Despite recent advances in vision-and-language tasks, most progress is still focused on resource-rich languages such as English. Furthermore, widespread vision-and-language datasets directly adopt images representative of American or European cultures resulting in bias. Hence we introduce ParsVQA-Caps, the first benchmark in Persian for Visual Question Answering and Image Captioning tasks. We utilize two ways to collect datasets for each task, human-based and template-based for VQA and human-based and web-based for image captioning. The image captioning dataset consists of over 7.5k images and about 9k captions. The VQA dataset consists of almost 11k images and 28.5k question and answer pairs with short and long answers usable for both classification and generation VQA.

1 papers0 benchmarksImages, Texts

Meta Omnium

Meta Omnium is a dataset-of-datasets spanning multiple vision tasks including recognition, keypoint localization, semantic segmentation and regression. Meta Omnium enables meta-learning researchers to evaluate model generalization to a much wider array of tasks than previously possible, and provides a single framework for evaluating meta-learners across a wide suite of vision applications in a consistent manner.

1 papers0 benchmarks

A View From Somewhere (AVFS)

A View From Somewhere (AVFS)—a dataset of 638,180 face similarity judgments over 4,921 faces. Each judgment corresponds to the odd-one-out (i.e., least similar) face in a triplet of faces and is accompanied by both the identifier and demographic attributes of the annotator who made the judgment.

1 papers0 benchmarksImages

titanic5 Dataset

titanic5 Dataset Created by David Beltran del Rio March 2016.

1 papers0 benchmarks

Webis-TLDR-17 Corpus

This corpus contains preprocessed posts from the Reddit dataset, suitable for abstractive summarization using deep learning. The format is a json file where each line is a JSON object representing a post. The schema of each post is shown below: - author: string (nullable = true) - body: string (nullable = true) - normalizedBody: string (nullable = true) - content: string (nullable = true) - content_len: long (nullable = true) - summary: string (nullable = true) - summary_len: long (nullable = true) - id: string (nullable = true) - subreddit: string (nullable = true) - subreddit_id: string (nullable = true) - title: string (nullable = true)

1 papers0 benchmarksTexts

COVIDx CXR-3

COVIDx CXR-3 is an open access benchmark dataset that we generated, comprising 30,882 CXR images across 17,026 patient cases. Images may be added over time to improve the dataset.

1 papers1 benchmarksImages, Medical

Tinto (Tinto: Multisensor Benchmark for 3D Hyperspectral Point Cloud Segmentation in the Geosciences)

The increasing use of deep learning techniques has reduced interpretation time and, ideally, reduced interpreter bias by automatically deriving geological maps from digital outcrop models. However, accurate validation of these automated mapping approaches is a significant challenge due to the subjective nature of geological mapping and the difficulty in collecting quantitative validation data. Additionally, many state-of-the-art deep learning methods are limited to 2D image data, which is insufficient for 3D digital outcrops, such as hyperclouds. To address these challenges, we present Tinto, a multi-sensor benchmark digital outcrop dataset designed to facilitate the development and validation of deep learning approaches for geological mapping, especially for non-structured 3D data like point clouds. Tinto comprises two complementary sets: 1) a real digital outcrop model from Corta Atalaya (Spain), with spectral attributes and ground-truth data, and 2) a synthetic twin that uses latent

1 papers0 benchmarks3D, Hyperspectral images, Point cloud

HAC (Hybrid Adverse Conditions)

HAC is a dataset for learning and benchmarking arbitrary Hybrid Adverse Conditions restoration. HAC contains 31 scenarios composed of an arbitrary combination of five common weather, with a total of 316K adverse-weather/clean pairs.

1 papers0 benchmarksImages

Morphological Classification of Galaxies

Dataset can be used by anyone who is interested to perform morphological classification of galaxies. Originally dataset provided by Kaggle user Jay Lin (https://www.kaggle.com/jay1985) 4 years ago. Dataset was used in conference paper "Morphological Classification of Galaxies Using SpinalNet"

1 papers0 benchmarksImages

SWS (Smart Word Suggestions Benchmark)

Smart Word Suggestions (SWS) is a task and benchmark. This task involves identifying words or phrases that require improvement and providing substitution suggestions. The benchmark includes human-labeled data for testing, a large distantly supervised dataset for training, and the framework for evaluation. The test data includes 1,000 sentences written by English learners, accompanied by over 16,000 substitution suggestions annotated by 10 native speakers. The training dataset comprises over 3.7 million sentences and 12.7million suggestions generated through rules.

1 papers0 benchmarksTexts

Multi-CrossRE

Multi-CrossRE is a broadest multi-lingual dataset for Relation Extraction (RE) including 26 languages in addition to English, and covering six text domains. It is a machine translated version of CrossRE crossre, with a sub-portion including more than 200 sentences in seven diverse languages checked by native speakers.

1 papers0 benchmarksTexts

Simulated wind farm graph dataset (floris-wind-farm-dataset)

FLORIS farm dataset A dataset for graph neural network modeling of wind farms. The current version of the dataset contains two farms, with very different geometry but similar inter-turbine statistics. The wind farms were simulated with the steady-state wake model FLORIS.

1 papers0 benchmarksGraphs

Honeycombs in Concrete (Honeycombs in Concrete Instance Segmentation)

The directory HiCIS contains two datasets for instance segmentation of honeycombs in concrete in COCO Format. The datasets orginate from images scraped from the internet and the other one is provided by Metis Systems AG. The directory HiCC/web contains the dataset using the images from the internet and HICC/metis contains the dataset using the images provided by Metis Systems AG as part of the research project Smart Design and Construction (SDaC).

1 papers0 benchmarksImages

naab

naab: A ready-to-use plug-and-play corpus for Farsi The biggest cleaned and ready-to-use open-source textual corpus in Farsi. It contains about 130GB of data, 250 million paragraphs, and 15 billion words. The project name is derived from the Farsi word NAAB K which means pure and high grade. We also provide the raw version of the corpus called naab-raw and an easy-to-use preprocessor that can be employed by those who wanted to make a customized corpus.

1 papers0 benchmarks

TFVulFix (TensorFlow Vulnerability Fixes)

TFVulFix is a dataset containing commits from TensorFlow, which is a well-known deep learning library. It contains 290 vulnerability fixing and 1,535 non-vulnerability-fixing commits. In this dataset, no commit is explicitly linked to an issue.

1 papers0 benchmarks

SMIC

I'm writing a description for the SMIC dataset on Papers With Code. SMIC is a dataset for facial expression recognition (FER).

1 papers0 benchmarks

ChatGPT Advice Responses

Taking Advice from ChatGPT is a laboratory study of how student participants incorporate advice generated by ChatGPT. In a survey conducted through the Experimental Social Science Laboratory, 118 students answered 2,828 questions on topics from the MMLU benchmark. The rich dataset includes questions/choices, advice characteristics, participant answers, and participant background. It can be used to explore algorithm aversion, advice-taking, ChatGPT usage, and more.

1 papers0 benchmarksTexts
PreviousPage 461 of 1000Next