TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Wiki5K Hebrew segmentation

Training data for Hebrew morphological word segmentation

1 papers0 benchmarks

Table Tennis Ball Trajectories with Spin

This data set contains real-world table tennis ball trajectories recorded with our custom developed table tennis ball launcher AIMY. The data has been filtered. Faulty samples and ball trajectories have been removed from the data set. The trajectories are stored in the widely supported file format HDF5 (Hierarchical Data Format). For easier usage of the data, we attach a simple Python script.

1 papers0 benchmarks

HM3D-Semantics (Habitat-Matterport 3D Semantics)

Habitat-Matterport 3D Semantics Dataset (HM3D-Semantics v0.1) is the largest-ever dataset of semantically-annotated 3D indoor spaces. It contains dense semantic annotations for 120 high-resolution 3D scenes from the Habitat-Matterport 3D dataset. The HM3D scenes are annotated with the 1700+ raw object names, which are mapped to 40 Matterport categories. On average, each scene in HM3D-Semantics v0.1 consists of 646 objects from 114 categories.

1 papers0 benchmarks3D, Images

GOD-Wiki, DIR-Wiki, ThingsEEG-Text.

brain-image-text trimodal datasets

1 papers0 benchmarks

Domains Project

Domains Project is a public dataset contains freely available sorted list of Internet domains.

1 papers0 benchmarks

SPAVE-28G (Signal Propagation Analyses in V2X Ecosystems (S.P.A.V.E) at 28 GHz on the NSF POWDER testbed)

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksEnvironment, Graphs, Time series, Tracking

FSC-P2 (Fearless Steps Challenge Phase2)

The Fearless Steps Initiative by UTDallas-CRSS led to the digitization, recovery, and diarization of 19,000 hours of original analog audio data, as well as the development of algorithms to extract meaningful information from this multichannel naturalistic data resource. As an initial step to motivate a stream-lined and collaborative effort from the speech and language community, UTDallas-CRSS is hosting a series of progressively complex tasks to promote advanced research on naturalistic “Big Data” corpora. This began with ISCA INTERSPEECH-2019: "The FEARLESS STEPS Challenge: Massive Naturalistic Audio (FS-#1)". This first edition of this challenge encouraged the development of core unsupervised/semi-supervised speech and language systems for single-channel data with low resource availability, serving as the “First Step” towards extracting high-level information from such massive unlabeled corpora. As a natural progression following the successful Inaugural Challenge FS#1, the FEARLESS

1 papers0 benchmarksAudio, Texts

CC-Riddle (Chinese Character Riddle)

CC-Riddle is a Chinese character riddle dataset covering the majority of common simplified Chinese characters by crawling riddles from the Web and generating brand new ones. In the generation stage, the authors provide the Chinese phonetic alphabet, decomposition and explanation of the solution character for the generation model and get multiple riddle descriptions for each tested character. Then the generated riddles are manually filtered and the final dataset, CCRiddle is composed of both human-written riddles and filtered generated riddle.

1 papers0 benchmarksTexts

VizWiz-FewShot

VizWiz-FewShot is a a few-shot localization dataset originating from photographers who authentically were trying to learn about the visual content in the images they took. It includes nearly 10,000 segmentations of 100 categories in over 4,500 images that were taken by people with visual impairments.

1 papers0 benchmarksImages

NAFLD pathology and healthy tissue samples

The dataset contains tiles extracted from Whole Slide Images (WSI) of stained tissue samples.

1 papers0 benchmarks

VISOR - Semi supervised video object segmentation (val)

VISOR is a dataset of pixel annotations and a benchmark suite for segmenting hands and active objects in egocentric video. VISOR annotates videos from EPIC-KITCHENS, and it contains 272K manual semantic masks of 257 object classes, 9.9M interpolated dense masks, and 67K hand-object relations, covering 36 hours of 179 untrimmed videos.

1 papers0 benchmarksImages, Videos

Trailers12k

Trailers12k is a movie trailer dataset comprised of 12,000 titles associated to ten genres. It distinguishes from other datasets by its collection procedure aimed at providing a high-quality publicly available dataset.

1 papers0 benchmarksImages, Texts, Videos

UCSF PDGM (The University of California San Francisco Preoperative Diffuse Glioma MRI Dataset)

MRI-based artificial intelligence (AI) research on patients with brain gliomas has been rapidly increasing in popularity in recent years in part due to a growing number of publicly available MRI datasets. Notable examples include The Cancer Genome Atlas Glioblastoma dataset (TCGA-GBM) consisting of 262 subjects and the International Brain Tumor Segmentation (BraTS) challenge dataset consisting of 542 subjects (including 243 preoperative cases from TCGA-GBM). The public availability of these glioma MRI datasets has fostered the growth of numerous emerging AI techniques including automated tumor segmentation, radiogenomics, and MRI-based survival prediction. Despite these advances, existing publicly available glioma MRI datasets have been largely limited to only 4 MRI contrasts (T2, T2/FLAIR, and T1 pre- and post-contrast) and imaging protocols vary significantly in terms of magnetic field strength and acquisition parameters. Here we present the University of California San Francisco Pre

1 papers0 benchmarks

Handwash Dataset

Hand Wash Dataset consists of 292 videos of hand washes with each hand wash having 12 steps, for a total of 3,504 clips, in different environments to provide as much variance as possible. The variance was important to ensure that the model is robust and can work in more than a few environments. This dataset is designed for action recognition tasks

1 papers0 benchmarks

Multi-domain Image Characteristics Dataset

The Multi-domain Image Characteristic Dataset consists of thousands of images sourced from the internet. Each image falls under one of three domains - animals, birds, or furniture. There are five types under each domain. There are 200 images of each type, summing up the total dataset to 3,000 images. The master file consists of two columns; the image name and the visible characteristics of that image. Every image was manually analyzed and the characteristics for each image were generated, ensuring accuracy.

1 papers0 benchmarksImages, Texts

CORRONA CERTAIN (Comparative Effectiveness Registry to Study Therapies for Arthritis and Inflammatory Conditions)

CERTAIN, or the Comparative Effectiveness Registry to Study Therapies for Arthritis and Inflammatory Conditions, is designed as a prospective nested substudy under our larger RA registry.

1 papers0 benchmarks

ACL Anthology Corpus with Full Text

This repository provides full-text and metadata to the ACL anthology collection (80k articles/posters as of September 2022) also including .pdf files and grobid extractions of the pdfs.

1 papers0 benchmarksTexts

DBP-5L (Spanish)

DPB-5L is a Multilingual KG dataset containing 5 KGs in English, French, Japanese, Greek, and Spanish. The dataset is used for the Knowledge Graph Completion and Entity Alignment task. DPB-5L (Spanish) is a subset of DPB-5L with Spanish KG.

1 papers0 benchmarksGraphs

Replication Data for: On estimating Armington elasticities for Japan's meat imports

Replication Data for: On estimating Armington elasticities for Japan's meat imports contains monthly import values and quantities from Jan 1996 to Dec 2020 for all 78 items.

1 papers0 benchmarks

Mint (Multilingual Intimacy analysis)

Mint is a new Multilingual intimacy analysis dataset covering 13,384 tweets in 10 languages including English, French, Spanish, Italian, Portuguese, Korean, Dutch, Chinese, Hindi, and Arabic. The dataset is released along with the SemEval 2023 Task 9: Multilingual Tweet Intimacy Analysis.

1 papers0 benchmarksTexts
PreviousPage 440 of 1000Next