TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

CANDOR Corpus (CANDOR = Conversation: A Naturalistic Dataset of Online Recordings)

The CANDOR corpus is a large, novel, multimodal corpus of 1,656 recorded conversations in spoken English. This 7+ million word, 850 hour corpus totals over 1TB of audio, video, and transcripts, with moment-to-moment measures of vocal, facial, and semantic expression, along with an extensive survey of speaker post conversation reflections.

1 papers0 benchmarksImages, Tabular, Texts, Time series, Videos

ML guided Logic synthesis

Logic synthesis is a challenging and widely-researched combinatorial optimization problem during integrated circuit (IC) design. It transforms a high-level description of hardware in a programming language like Verilog into an optimized digital circuit netlist, a network of interconnected Boolean logic gates, that implements the function. Spurred by the success of ML in solving combinatorial and graph problems in other domains, there is growing interest in the design of ML-guided logic synthesis tools. Yet, there are no standard datasets or prototypical learning tasks defined for this problem domain. Here, we describe OpenABC-D,a large-scale, labeled dataset produced by synthesizing open source designs with a leading open-source logic synthesis tool and illustrate its use in developing, evaluating and benchmarking ML-guided logic synthesis. OpenABC-D has intermediate and final outputs in the form of 870,000 And-Inverter-Graphs (AIGs) produced from 1500 synthesis runs plus labels such a

1 papers0 benchmarks

NELA-GT-2021

NELA-GT-2021 is the fourth installment of the NELA-GT datasets, NELA-GT-2021. The dataset contains 1.8M articles from 367 outlets between January 1st, 2021 and December 31st, 2021. Just as in past releases of the dataset, NELA-GT-2021 includes outlet-level veracity labels from Media Bias/Fact Check and tweets embedded in collected news articles.

1 papers0 benchmarksTexts

ILPC22-Small

A small dataset from the Inductive Link Prediction Challenge 2022. Training graph contains 10K entities, 96 relations, 78K triples. Inference graph contains 7K entities, 96 relations, 21K triples. Validation and test triples to predict belong to the inference graph.

1 papers7 benchmarksGraphs

ILPC22-Large

A large dataset from the Inductive Link Prediction Challenge 2022. Training graph contains 46K entities, 130 relations, 202K triples. Inference graph contains 30K entities, 130 relations, 77K triples. Validation and test triples to predict belong to the inference graph.

1 papers7 benchmarksGraphs

BBAI Dataset (Black-box Agent Integration)

This dataset is for evaluating the task of Black-box Multi-agent Integration which focuses on combining the capabilities of multiple black-box conversational agents at scale. It provides data to explore two main frameworks of exploration: question agent pairing and question response pairing.

1 papers1 benchmarksTexts

Thermal focus image database

The database was acquired using a thermographic camera TESTO 880-3. This camera is equipped with an uncooled detector and has a spectral sensitivity range from 8 to 14 μm. It has a removable German optic lens. It provides the following main features:

1 papers0 benchmarks

Slovenian Twitter dataset 2018-2020

A comprehensive set of all Slovenian tweets posted in the 2018-2020 period, with retweet links and assigned hate speech classes. Available at a public language resource repository CLARIN.SI.

1 papers0 benchmarks

AIT-LDSv2.0 (AIT Log Data Set V2.0)

Synthetic log data suitable for evaluation of intrusion detection systems, federated learning, and alert aggregation. Each of the 8 datasets corresponds to a testbed representing a small enterprise network including mail server, file share, WordPress server, VPN, firewall, etc. Normal user behavior is simulated to generate background noise over a time span of 4-6 days. At some point, a sequence of attack steps are launched against the network. Log data is collected from all hosts and includes Apache access and error logs, authentication logs, DNS logs, VPN logs, audit logs, Suricata logs, network traffic packet captures, horde logs, exim logs, syslog, and system monitoring logs. Attacks include scans (nmap, WPScan, dirb), webshell upload, password cracking, privilege escalation, remote command execution, and data exfiltration.

1 papers0 benchmarks

eVED (Extended Vehicle Energy Dataset)

Extended Vehicle Energy Dataset (eVED) is an extended version of the Vehicle Energy Dataset (VED), which is a large-scale dataset for vehicle energy consumption analysis. Compared with its original version, the extended VED (eVED) dataset is enhanced with accurate vehicle trip GPS coordinates, serving as a reliable basis to associate the VED trip records with external information e.g., road speed limit and intersections, from accessible map services to accumulate attributes that is relevant and essential in analyzing vehicle energy consumption.

1 papers0 benchmarks

Multi-focus thermal database

The database was acquired using a thermographic camera TESTO 882-3 equipped with an uncooled detector and a spectral sensitivity range from 8 to 14 μm. It has a removable German optic lens with these main features:

1 papers0 benchmarks

i3DMM Test Dataset

A new dataset consisting of 64 people with different expressions and hairstyles.

1 papers0 benchmarks

NEMO Sea Surface Temperature Dataset

This dataset contains spatiotemporal sequences of SST generated by the NEMO ocean engine. The observations correspond to 250 randomly selected data sites within a [0, 550] × [100, 650] square cropped from the area between 50 deg N − 65 deg N and 75W deg − 10W deg starting from 01-01-2016 to 12-31-2017. The data is divided into 24 sequences, each lasting 30 days (extra days in each month are truncated). Data corresponding to 2016 are used for training and the rest is used for validation and testing, in the equal sequential split.

1 papers0 benchmarks

Full-Spectral Autofluorescence Lifetime Microscopic Images

The dataset contains full-spectral autofluorescence lifetime microscopic images (FS-FLIM) acquired on unstained ex-vivo human lung tissue, where 100 4D hypercubes of 256x256 (spatial resolution) x 32 (time bins) x 512 (spectral channels from 500nm to 780nm). This dataset associates with our paper "Deep Learning-Assisted Co-registration of Full-Spectral Autofluorescence Lifetime Microscopic Images with H&E-Stained Histology Images" (https://arxiv.org/abs/2202.07755) and "Full spectrum fluorescence lifetime imaging with 0.5 nm spectral and 50 ps temporal resolution" (https://doi.org/10.1038/s41467-021-26837-0). The FS-FLIM images provide transformative insights into human lung cancer with extra-dimensional information. This will enable visual and precise detection of early lung cancer. With the methodology in our co-registration paper, FS-FLIM images can be registered with H&E-stained histology images, allowing characterisation of tumour and surrounding cells at a celluar level with abs

1 papers0 benchmarksBiomedical, Hyperspectral images

Code Smells in Elixir

Dataset used in research submitted to ICPC ERA 2022

1 papers0 benchmarks

IntHarmony

This newly curated synthetic dataset specifies an additional reference region to guide image harmonization. There are 118,287 training images and 959 test images. The dataset consists of objects, backgrounds, and people.

1 papers0 benchmarksImages

EGDB

This dataset contains transcriptions of the electric guitar performance of 240 tablatures, rendered with different tones. The goal is to contribute to automatic music transcription (AMT) of guitar music, a technically challenging task.

1 papers0 benchmarksAudio

ChildCIdb (ChildCIdbv1)

A large-scale, first-of-its-kind database aimed at generating a better understanding of the way children interact with mobile devices during their development process. ChildCIdbv1 comprises data collected from 438 children, from 18 months to 8 years old, encompassing the first three development stages of Piaget's theory. Data collected spans interaction with screens using both finger and pen stylus, information regarding the previous experience of the child with mobile devices, the child’s grade level, and whether attention-deficit/hyperactivity disorder (ADHD) is present.

1 papers0 benchmarksInteractive

VidHarm

VidHarm is a professionally annotated dataset for detection of harmful content in video. Include 3589 annotate video clips from a variety of film trailers. In contrast to previous approaches which mostly use meta data from long sequences, it uses the raw video and focus on short clips.

1 papers0 benchmarksVideos

Spatial Commonsense Graph Dataset (Spatial Commonsense Graph for Object Localisation in Partial Scenes)

Dataset built from partial reconstructions of real-world indoor scenes using RGB-D sequences from ScanNet, aimed at estimating the unknown position of an object (e.g. where is the bag?) given a partial 3D scan of a scene. The dataset mostly consists of bedrooms, bathrooms, and living rooms. Some room types like closet and gym only have a few instances.

1 papers0 benchmarks3D, Images
PreviousPage 422 of 1000Next