TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

DA$^2$ (Dual-Arm Dexterity-Aware)

DA2 is a large-scale Dual-Arm Dexterity-Aware (DA2) grasp dataset, including a total of about 9M parallel-jaw grasp pairs for more than 6000 different meshes. The grasp pairs are labeled with multiple grasp dexterity measures by fully analyzing the grasp matrix. The dataset constitutes a standardized data source filling the gap in vision-guided dual-arm grasping of arbitrary objects.

1 papers0 benchmarks

CB-ToF-Extension (Extension of Cornell-Box Time-of-Flight Dataset)

An extension of the Cornell-Box Time-of-Flight Dataset (https://paperswithcode.com/dataset/cb-tof) containing moving objects. It follows the same data structure.

1 papers0 benchmarks

LRDE DBD

This is a dataset is composed of full-document images, groundtruth, and tools to perform an evaluation of binarization algorithms. It allows pixel-based accuracy and OCR-based evaluations.

1 papers0 benchmarks

Acoustic frequency responses in a conventional classroom

Hahmann, Manuel; Verburg Riezu, Samuel Arturo (2021): Acoustic frequency responses in a conventional classroom. Technical University of Denmark. Dataset. https://doi.org/10.11583/DTU.13315286

1 papers0 benchmarks

TAT (Taiwanese Across Taiwan)

Taiwanese Across Taiwan (TAT) corpus is a Large-Scale database of Native Taiwanese Article/Reading Speech collected across Taiwan. This corpus contains native Taiwanese speech of various accent across Taiwan. The corpus is annotated twice for use in voice recognition research. The corpus contains recording from 100 native speakers, each with length of 30 minutes making a total of 100 hours of speech data.

1 papers2 benchmarksSpeech

HOWS (HOWS-CL-25)

HOWS-CL-25 (Household Objects Within Simulation dataset for Continual Learning) is a synthetic dataset especially designed for object classification on mobile robots operating in a changing environment (like a household), where it is important to learn new, never seen objects on the fly. This dataset can also be used for other learning use-cases, like instance segmentation or depth estimation. Or where household objects or continual learning are of interest.

1 papers1 benchmarksImages, RGB-D

detection_of_IoT_botnet_attacks_N_BaIoT Data Set

This dataset addresses the lack of public botnet datasets, especially for the IoT. It suggests real traffic data, gathered from 9 commercial IoT devices authentically infected by Mirai and BASHLITE.

1 papers0 benchmarks

Lipogram-e

This is a dataset of 3 English books which do not contain the letter "e" in them. This dataset includes all of "Gadsby" by Ernest Vincent Wright, all of "A Void" by Georges Perec, and almost all of "Eunoia" by Christian Bok (except for the single chapter that uses the letter "e" in it)

1 papers2 benchmarksTexts

Lowest Common Ancestor Generations (LCAG) Phasespace Particle Decay Reconstruction Dataset

This dataset contains simulated synthetic particle decays, simulated using the PhaseSpace library. All simulated decay topologies have a common root particle of mass 100 (arbitrary units). Intermediate particles are selected at random with replacement from the following masses: [90, 80, 70, 50, 25, 20, 10]. Final state particles, which make up the leaf nodes of generated topologies, are drawn with replacement from the following masses: [1, 2, 3, 5, 12]. For each intermediate particle (including the root), we limit the minimum number of children to two, and the maximum five. The dataset contains the resulting simulated particle physics decays, with information about the detected particle (leaves) to be used as input, and Lowest Common Ancestor Generations (LCAGs) to be used as training targets.

1 papers0 benchmarksPhysics

HYPERVIEW (Seeing Beyond the Visible)

The dataset comprises 2886 patches in total (2 m GSD), of which 1732 patches for training and 1154 patches for testing. The patch size varies (depending on agricultural parcels) and is on average around 60x60 pixels. Each patch contains 150 contiguous hyperspectral bands (462-942 nm, with a spectral resolution of 3.2 nm), which reflects the spectral range of the hyperspectral imaging sensor deployed on-board Intuition-1.

1 papers1 benchmarks3D, Images

HuTu 80 (HuTu 80 cell populations)

The image set contains 180 high-resolution color microscopic images of human duodenum adenocarcinoma HuTu 80 cell populations obtained in an in vitro scratch assay (for the details of the experimental protocol, we refer to (Liang et al., 2007)). Briefly, cells were seeded in 12-well culture plates ($20 \times 10^3$ cells per well) and grown to form a monolayer with 85\% or more confluency. Then the cell monolayer was scraped in a straight line using a pipette tip ($200 \mu L$). The debris was removed by washing with a growth medium and the medium in wells was replaced. The scratch areas were marked to obtain the same field during the image acquisition. Images of the scratches were captured immediately following the scratch formation, as well as after 24, 48 and 72 h of cultivation.

1 papers2 benchmarksBiomedical, Images

SFpark (San Francisco Park Evaluation)

The San Francisco Municipal Transportation Agency (SFMTA) website provides data collected during the SFpark pilot project. On-street occupancy rate data contain per-block hourly occupancy rates and meter prices for seven parking districts.

1 papers0 benchmarks

Demosthenes

Corpus for argument mining in legal documents, composed of 40 decisions of the Court of Justice of the European Union on matters of fiscal state aid

1 papers0 benchmarksTexts

High-cardinality Geometrically Shaped Constellation for the AWGN channel and optical fibre channel

Optimised constellation for the paper High-Cardinality Geometrical Constellation Shaping for the Nonlinear Fibre Channel. Each file is a constellation optimised for the SNR in dB mentioned in the filename, containing the coordinates of the constellation points as comma-separated values. Each column represents a dimension and each row is a separate constellation point. The bit labels for the generalised mutual information (GMI) are implied and follow natural mapping, the first row is 0,..,0,0 the second 0,...0,1 the third 0,..,1,0 the fourth 0,...,1,1 etc and the last 1,...,1,1. The file named gmi.txt is the GMI for the resulting constellations.

1 papers0 benchmarks

MovieCLIP

MovieCLIP is a movie-centric taxonomy of 179 scene labels derived from movie scripts and auxiliary web-based video datasets designed for visual scene recognition.

1 papers0 benchmarksImages

XiaChuFang Recipe Corpus

XiaChuFang Recipe Corpus contains recipes are from 下厨房 (XiaChuFang), a popular Chinese recipe sharing website. The full recipe corpus contains 1,520,327 Chinese recipes. Among them, 1,242,206 recipes belong to 30,060 dishes. A dish has 41.3 recipes on average.

1 papers0 benchmarksTexts

Wikipedia Knowledge Graph dataset

Wikipedia is the largest and most read online free encyclopedia currently existing. As such, Wikipedia offers a large amount of data on all its own contents and interactions around them, as well as different types of open data sources. This makes Wikipedia a unique data source that can be analyzed with quantitative data science techniques. However, the enormous amount of data makes it difficult to have an overview, and sometimes many of the analytical possibilities that Wikipedia offers remain unknown. In order to reduce the complexity of identifying and collecting data on Wikipedia and expanding its analytical potential, after collecting different data from various sources and processing them, we have generated a dedicated Wikipedia Knowledge Graph aimed at facilitating the analysis, contextualization of the activity and relations of Wikipedia pages, in this case limited to its English edition. We share this Knowledge Graph dataset in an open way, aiming to be useful for a wide range

1 papers0 benchmarksTabular

Cross-institution Male Pelvic Structures

The data set includes 589 T2-weighted images acquired from the same number of patients collected by seven studies, INDEX, the SmartTarget Biopsy Trial, PICTURE, TCIA Prostate3T, Promise12, TCIA ProstateDx (Diagnosis) and the Prostate MR Image Database. Further details are reported in the respective study references.

1 papers0 benchmarksMedical

The Reddit Climate Change Dataset

The Reddit Climate Change Dataset is a dataset of 620K Reddit posts and 4.6M comments - all mentions of the terms "climate" and "change" until 2022-09-01 across the entire Reddit social network. Both were procured with SocialGrep's export feature and released as part of SocialGrep Reddit datasets. The posts are labeled with their subreddit, title, creation date, domain, selftext, and score. The comments are labeled with their subreddit, body, creation date, sentiment (calculated for you using a VADER pipeline), and score.

1 papers0 benchmarksTabular, Texts

RGZ EMU: Semantic Taxonomy (Radio Galaxy Zoo EMU: Towards a Semantic Radio Galaxy Morphology Taxonomy)

The data used in - "Radio Galaxy Zoo EMU: Towards a Semantic Radio Galaxy Morphology Taxonomy" (Bowles et al. submitted) - "A New Task: Deriving Semantic Class Targets for the Physical Sciences" (Bowles et al. 2022: https://arxiv.org/abs/2210.14760) accepted at the Fifth Workshop on Machine Learning and the Physical Sciences, Neural Information Processing Systems 2022.

1 papers0 benchmarksImages, Tabular, Texts
PreviousPage 441 of 1000Next