TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Linked WikiText-2

The Linked Wikitext-2 language modeling dataset contains over 2 million tokens from Wikipedia articles, along with annotations linking mentions to their corresponding entities and relations in Wikidata. It is designed to match as closely as possible the contents of the popular WikiText-2 dataset.

1 papers0 benchmarks

ShortPersianEmo

ShortPersianEmo is a new data set for emotion recognition in Persian short texts. The ShortPersianEmo dataset is a single-label dataset that contains 5472 short Persian texts collected from Twitter and Digikala. Our dataset is annotated according to Rachael Jack’s emotional model in five emotional classes happiness, sadness, anger, fear, and other. Unlike publicly accessible datasets that do not impose any restrictions on text length, ShortPersianEmo specifically focuses on short texts. The average text length in the ShortPersianEmo dataset is 56 words. Table 1 presents a comparison between the introduced ShortPersianEmo dataset and other datasets from the literature for emotion detection in Persian text. For more information on this dataset please read our paper. If you use this dataset in any research work, please cite our paper.

1 papers3 benchmarksTexts

REBUS (A Robust Evaluation Benchmark of Understanding Symbols)

Recent advances in large language models have led to the development of multimodal LLMs (MLLMs), which take both image data and text as an input. Virtually all of these models have been announced within the past year, leading to a significant need for benchmarks evaluating the abilities of these models to reason truthfully and accurately on a diverse set of tasks. When Google announced Gemini (Gemini Team et al., 2023), they showcased its ability to solve rebuses—wordplay puzzles which involve creatively adding and subtracting letters from words derived from text and images. The diversity of rebuses allows for a broad evaluation of multimodal reasoning capabilities, including image recognition, multi- step reasoning, and understanding the human creator’s intent. We present REBUS: a collection of 333 hand-crafted rebuses spanning 13 diverse cate- gories, including hand-drawn and digital images created by nine contributors. Samples are presented in Table 1. Notably, GPT-4V, the most powe

1 papers1 benchmarksImages, Texts

MOSAD (Mobile Sensing Human Activity Data Set)

MOSAD (Mobile Sensing Human Activity Data Set) is a multi-modal, annotated time series (TS) data set that contains 14 recordings of 9 triaxial smartphone sensor measurements (126 TS) from 6 human subjects performing (in part) 3 motion sequences in different locations. The aim of the data set is to facilitate the study of human behaviour and the design of TS data mining technology to separate individual activities using low-cost sensors in wearable devices.

1 papers0 benchmarksTime series

HASCD (Human Activity Segmentation Challenge Dataset)

HASCD (Human Activity Segmentation Challenge Dataset) contains 250 annotated multivariate time series capturing 10.7 h of real-world human motion smartphone sensor data from 15 bachelor computer science students. The recordings capture 6 distinct human motion sequences designed to represent pervasive behaviour in realistic indoor and outdoor settings. The data set serves as a benchmark for evaluating machine learning workflows.

1 papers0 benchmarksTime series

AUT-VI (Amirkabir campus dataset)

AUT-VI is a super-challenging visual inertial dataset with 126 diverse sequences in 17 locations. This dataset contains dynamic objects, challenging loop-closure/map-reuse, different lighting conditions, reflections, and sudden camera movements to cover all extreme navigation scenarios. Moreover, the Android application for data capture is released to the public to support ongoing development efforts. This dataset aims to exploit the remaining challenges in VIO algorithms, in the hope of improving them to facilitate navigation for visually impaired individuals in both indoor and outdoor settings.

1 papers0 benchmarks

MLO-Cn2 (Mauna Loa Seeing Study)

The Mauna Loa Seeing Study was performed by the EOL/Integrated Surface Flux System team, capturing surface meteorology and flux products at the Mauna Loa Observatory in Hawaii.

1 papers3 benchmarksTime series

USNA-Cn2 (long-term) (Unites States Naval Academy Long-term Scintillation Study)

The USNA long-term scintillation study is a continuing effort to characterize and measure optical turbulence in the near-maritime boundary layer.

1 papers2 benchmarksTime series

USNA-Cn2 (short-duration) (Unites States Naval Academy Short-duration Optical Turbulence Dataset)

The USNA long-term scintillation study is a continuing effort to characterize and measure optical turbulence in the near-maritime boundary layer.

1 papers3 benchmarksTime series

multi-view-3DGPR

If you want to known more about this dataset and new method, please read our paper link.

1 papers0 benchmarksImages

SICKLE (Satellite Imagery for Cropping annotated with Keyparameter LabEls)

The availability of well-curated datasets has driven the success of Machine Learning (ML) models. Despite greater access to earth observation data in agriculture, there is a scarcity of curated and labelled datasets, which limits the potential of its use in training ML models for remote sensing (RS) in agriculture. To this end, we introduce a first-of-its-kind dataset called SICKLE, which constitutes a time-series of multi-resolution imagery from 3 distinct satellites: Landsat-8, Sentinel-1 and Sentinel-2. Our dataset constitutes multi-spectral, thermal and microwave sensors during January 2018 - March 2021 period. We construct each temporal sequence by considering the cropping practices followed by farmers primarily engaged in paddy cultivation in the Cauvery Delta region of Tamil Nadu, India; and annotate the corresponding imagery with key cropping parameters at multiple resolutions (i.e. 3m, 10m and 30m). Our dataset comprises 2, 370 season-wise samples from 388 unique plots, having

1 papers1 benchmarksEnvironment, Images, Time series

VGS (VideoGazeSpeech)

Gaze following attracts much attention recently while existing databases commonly lack audio information. In this work, we collect the first gaze following dataset containing audios, the VideoGazeSpeech Dataset. The dataset is used to evaluate our method and also encourage future research in multi-modal gaze following. Our dataset comprises a total of $35,231$ frames of $29$ videos. Each video in the dataset has an average duration of approximately $20$ seconds and is recorded at a frame rate of $25$ frames per second (fps). The resolution of each video is $1280 \times 720$ pixels, and the entire dataset occupies a storage space of $7.2$ GB.

1 papers0 benchmarks

ChesapeakeRSC (Chesapeake Roads Spatial Context)

A novel remote sensing dataset for evaluating a geospatial machine learning model's ability to learn long range dependencies and spatial context understanding. We create a task to use as a proxy for this by training models to extract roads which have been broken into disjoint pieces due to tree canopy occluding large portions of the road.

1 papers3 benchmarksImages

Koo Platform Dataset

Full Koo platform Dataset. See https://zenodo.org/records/10476212

1 papers0 benchmarks

Object-Centric Stylized COCO

An object-centric version of Stylized COCO to benchmark texture bias and out-of-distribution robustness of vision models. See the ECCV 22 paper and supplementary material for details.

1 papers0 benchmarksImages

Western Mediterranean Wetlands Bird Dataset

Manually labelled dataset of bird recordings from the species of interest inhabiting in the wetlands of the "Aiguamolls del Empord`{a}" natural park in Girona, Spain. The dataset includes 5,795 annotated audio clips generated from a source of 1,098 recordings retrieved from the Xeno-Canto portal, adding up to a total of 201.6 minutes (12,096 seconds) of vocalizations of different lengths, alongside with their corresponding annotations.

1 papers0 benchmarksAudio

Grounded Image-Text dataset

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

EAPD (Expert-labeled Aesthetics Perception Database)

An expert benchmark aiming to comprehensively evaluate the aesthetic perception capacities of MLLMs.

1 papers0 benchmarksImages, Texts

OpenAlex (OpenAlex: The open catalog to the global research system)

Details of OpenAlex Data can be seen in the official wensite: https://openalex.org/about

1 papers0 benchmarks

TAO-Amodal

Our dataset augments the TAO dataset with amodal bounding box annotations for fully invisible, out-of-frame, and occluded objects. Note that this implies TAO-Amodal also includes modal segmentation masks (as visualized in the color overlays above). Our dataset encompasses 880 categories, aimed at assessing the occlusion reasoning capabilities of current trackers through the paradigm of Tracking Any Object with Amodal perception (TAO-Amodal).

1 papers0 benchmarksImages, Videos
PreviousPage 485 of 1000Next