TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Digital Typhoon Dataset V2

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksImages, Time series

Nepali Text Corpus

Overview Nepali-Text-Corpus is a comprehensive collection of approximately 6.4 million articles in the Nepali language. This dataset is the largest text dataset on Nepali Language. It encompasses a diverse range of text types, including news articles, blogs, and more, making it an invaluable resource for researchers, developers, and enthusiasts in the fields of Natural Language Processing (NLP) and computational linguistics.

1 papers0 benchmarks

SuSy Dataset

The SuSy Dataset combines authentic photographs and AI-generated images designed for training and evaluating synthetic image detection models. It contains over 25,000 images from six different sources, including real-world photographs from COCO and synthetic images created by state-of-the-art diffusion models such as DALL-E 3, Midjourney, and Stable Diffusion.

1 papers0 benchmarksImages

MVImgNet

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

Bengali Curated News Summary Dataset

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksTexts

Datasets with Anomalous Regions

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

AutoTherm

Temporal Dataset for Indoor and In-Vehicle Thermal Comfort Estimation Abstract Thermal comfort estimation is essential for enhancing user experience in static indoor environments and dynamic in-vehicle scenarios. While traditional datasets focus on buildings, their application to fast-changing conditions, such as in vehicles, remains unexplored. We address this gap by introducing two temporal datasets collected from (1) a self-built climatic chamber with 31 sensor signals and user-labeled ratings from 18 participants and (2) in-vehicle studies with 20 participants in a BMW 3 Series.

1 papers0 benchmarksAudio, EEG, Images, Time series, Tracking

OmniLab

In order to evaluate the effectiveness of NToP in real-world scenarios, we collect a new dataset OmniLab with a top-view omnidirectional camera, mounted on the ceiling of two different rooms (bedroom, living room) at 2.5 m height. Five actors (3 males, 2 females) perform 15 actions from CMU-MoCap database (brooming, cleaning windows, down and get up, drinking, fall-on-face, in chair and stand up, pull object, push object, rugpull, turn left, turn right, upbend from knees, upbend from waist, up from ground, walk, walk-old-man) in two rooms with varying clothes. The recorded action length is 2.5 s, which results in 60 images for each scene at a frame rate of 24 FPS. The position of the camera is fixed and the resolution of the images is 1200 by 1200 pixels. A total of 4800 frames are collected. All annotations of 17 keypoints conforming to COCO conventions are estimated through a keypoint detector and subsequently refined by four different humans in two loops to ensure high annotation qu

1 papers0 benchmarksActions, Images

NToP

Human pose estimation (HPE) in the top-view using fisheye cameras presents a promising and innovative application domain. However, the availability of datasets capturing this viewpoint is extremely limited, especially those with high-quality 2D and 3D keypoint annotations. Addressing this gap, we leverage the capabilities of Neural Radiance Fields (NeRF) technique to establish a comprehensive pipeline for generating human pose datasets from existing 2D and 3D datasets, specifically tailored for the top-view fisheye perspective. Through this pipeline, we create a novel dataset NToP (NeRF-powered Top-view human Pose dataset for fisheye cameras) with over 570 thousand images, and conduct an extensive evaluation of its efficacy in enhancing neural networks for 2D and 3D top-view human pose estimation. Extensive evaluations on existing top-view 2D and 3D HPE datasets as well as our new real-world top-view 2D HPE dataset OmniLab prove that our dataset is effective and exceeds previous datase

1 papers0 benchmarks3d meshes, Images

3DO Dataset

3DO Dataset | On the Generalization of WiFi-based Person-centric Sensing in Through-Wall Scenarios

1 papers0 benchmarks

arabic-img2md (Arabic Img2MD)

Click to add a brief description of the dataset (Markdown and LaTe# Arabic Img2MD

1 papers0 benchmarks

BimanGrasp-Dataset

BimanGrasp-Dataset is the first synthesized bimanual dexterous hand grasp pose dataset.

1 papers0 benchmarks

VREM-FL datasets

This dataset collection includes three files used for the experiments. Each file contains 6 columns: {timestep, vehicle ID, x coordinate in the map, y coordinate in the map, real bitrate, estimated bitrate}. The datasets, obtained from REMs with Gaussian estimation and real (https://ieee-dataport.org/open-access/crawdad-romataxi) or simulated (https://eclipse.dev/sumo/) vehicular mobility, are used in the original paper for optimizing the task of federated learning (client scheduling and resource allocation).

1 papers0 benchmarksTables, Time series

https://zenodo.org/records/13495922

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

Augmented Wine Quality

The dataset utilized for this study is the Wine Quality dataset, which comprises 1,599 rows and 11 features related to the chemical properties of wine samples. The goal is to predict the ”quality” of the wine, a target variable that is an ordinal integer value, based on the following 10 features: fixed acidity, volatile acidity, citric acid, residual sugar, chlorides, free sulfur dioxide, density, pH, sulphates, and alcohol.

1 papers0 benchmarks

Equilibrium-Traffic-Networks

This repository contains three graph datasets for the UE traffic assignment problem on Sioux-Falls, Eastern-Massachusetts and Anaheim networks in both dgl and pyg formats. The datasets are generated and used to train and evaluate models for solving the User Equilibrium (UE) problem on three transportation networks:

1 papers0 benchmarksGraphs

USCOCO (Unexpected Situations of Common Objects in Context)

A test set of grammatically correct sentences and layouts (visual “imagined” situations), called Unexpected Situations of Common Objects in Context (USCOCO) describing compositions of entities and relations that are unlikely to be found in MS COCO.

1 papers0 benchmarksTexts

CRSB (Context Retrieval Supervision Benchmark)

The Official dataset proposed int the paper Context Awareness Gate For Retrieval Augmented Generation

1 papers0 benchmarks

UK Key Stage Readability (UK Key Stage Readability for English Texts)

Education is increasingly data-driven, and the ability to analyse and adapt educational materials quickly and effectively is important for keeping materials contemporary and interesting. These approaches also have the potential to personalise learning experiences. One of the challenges in this domain is aligning new literature with the appropriate educational stages. This dataset aims to contribute to alleviating this knowledge gap.

1 papers2 benchmarksTexts

United-Syn-Med

The United-Syn-Med dataset is a specialized medical speech dataset designed to evaluate and improve Automatic Speech Recognition (ASR) systems within the healthcare domain. It comprises English medical speech recordings, with a particular focus on medical terminology and clinical conversations. The dataset is well-suited for various ASR tasks, including speech recognition, transcription, and classification, facilitating the development of models tailored for medical contexts.

1 papers0 benchmarksAudio, Speech
PreviousPage 531 of 1000Next