TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

BNLP-Resources

Datasets for Bangla Natural Language Processing tasks.

0 papers0 benchmarksTexts

BEHAVIOR

BEHAVIOR is a benchmark with the 100 household activities that represent a new challenge for embodied AI solutions.

0 papers0 benchmarksEnvironment

UAV-based multispectral vineyards

UAS-based Multispectral othomosaics of vineyards from central Portugal

0 papers0 benchmarksImages

DAHLIA (DAily Human Life Activity)

DAHLIA dataset [1] is devoted to human activity recognition, which is a major issue for adapting smart-home services such as user assistance. DAHLIA has been realized in Mobile Mii Platform by CEA LIST, and has been partly supported by ITEA 3 Emospaces Project (https://itea3.org/project/emospaces.html)

0 papers0 benchmarksRGB-D, Videos

ESPADA (Extended Synthetic and Photogrammetric Aerial-Image Dataset)

We present a new aerial image dataset, named ESPADA, intended for the training of deep neural networks for depth image estimation from a single aerial image. Given the difficulty of creating aerial image datasets containing image pairs of chromatic images related to their depth images, simulators such as AirSim have been proposed to generate synthetic images from photorealistic scenes. The latter enables the generation of thousands of images that can be used to train and evaluate neural models. However, we argue that synthetic photorealistic aerial image datasets can be improved by adding images generated from photogrammetric models imported into the simulator, thus enabling a less artificial generation of both chromatic and depth images. To assess the quality of these images, we compare the performance of 4 deep neural networks whose pre-trained models and code for re-training are publicly available. We also use ORB-SLAM, in its RGB-D version, to indirectly assess the estimated depth

0 papers0 benchmarks

Frustrated Legislators: Replication data and code

Description: Replication data and code for Aref, S., and Neal, Z.P., "Identifying hidden coalitions in the US House of Representatives by optimally partitioning signed networks based on generalized balance" (2021) Scientific Reports. http://dx.doi.org/10.1038/s41598-021-98139-w

0 papers0 benchmarks

Hocalarim: Turkish Student Reviews

We have constructed our dataset by five fields available on the website that were found convenient for the study of student expectations and experience. This includes out-of-five star ratings on easiness, understandability, recitation, accessibility and helpfulness. Average rating was calculated based on these given five fields. Overall sentiment of the review was determined based on the average rating where any score higher than 3.5 (>=) was labeled as a positive review, and anything lower than 2.5 (<) was labeled as a negative review. The five main aspects students needed to rate was given below.

0 papers0 benchmarks

COVID-19 Disinfo (COVID-19 Disinformation Twitter Dataset)

With the emergence of the COVID-19 pandemic, the political and the medical aspects of disinformation merged as the problem got elevated to a whole new level to become the first global infodemic. Fighting this infodemic has been declared one of the most important focus areas of the World Health Organization, with dangers ranging from promoting fake cures, rumors, and conspiracy theories to spreading xenophobia and panic. Addressing the issue requires solving a number of challenging problems such as identifying messages containing claims, determining their check-worthiness and factuality, and their potential to do harm as well as the nature of that harm, to mention just a few. To address this gap, we release a large dataset of 16K manually annotated tweets for fine-grained disinformation analysis that focuses on COVID-19, combines the perspectives and the interests of journalists, fact-checkers, social media platforms, policy makers, and society, and covers Arabic, Bulgarian, Dutch, and

0 papers0 benchmarksTexts

EU-ADR

The EU-ADR corpus is a biomedical relation extraction dataset that contains 100 abstracts, with relations between drug, disorder, and targets.

0 papers0 benchmarksTexts

LEAKAGE-PERSONA Dataset

This is the synthetic dataset used for training a model which alerts users for potential leakages of personal information.

0 papers0 benchmarks

Corn Seeds Dataset

This dataset is the images of corn seeds considering the top and bottom view independently (two images for one corn seed: top and bottom). There are four classes of the corn seed (Broken-B, Discolored-D, Silkcut-S, and Pure-P) 17802 images are labeled by the experts at the AdTech Corp. and 26K images were unlabeled out of which 9k images were labeled using the Active Learning (BatchBALD)

0 papers0 benchmarksImages

COVID-19 Contact Tracing Survey

A survey of Israelis about their attitudes towards COVID-19 contact tracing apps

0 papers0 benchmarks

DCCW

| Name | Purpose | |------|---------| | FM100P | Evaluation of the single palette sorting | | KHTP | Evaluation of the palette pair sorting | | LHSP | Evaluation of the palette similarity measurement | | Perceptual Study | Perceptual Study |

0 papers0 benchmarks

Genome-wide miRNA detection (Genome-wide hairpins datasets of animals and plants for novel miRNA prediction)

We've made available several genome-wide datasets, which can be used for training microRNA (miRNA) classifiers. The hairpin sequences available are from the genomes of: Homo sapiens, Arabidopsis thaliana, Anopheles gambiae, Caenorhabditis elegans and Drosophila melanogaster. Hairpin.s are small RNA sequences that naturaly folds into a hairpin-structure. However, not all hairpins have clear function (they are not miRNAs).

0 papers0 benchmarksBiology, Biomedical

Supporting data for "Multi-Stage Malaria Parasites Recognition by Deep Learning"

Malaria, a mosquito-borne infectious disease affecting humans and other animals, is widespread in the tropical and subtropical regions. Microscopy is the most common method in diagnosing the malaria parasite from stained blood smears. However, this procedure is time-consuming, error-prone, and requires a well-trained professional. Moreover, the recognition of a malaria parasite through a microscope is still a challenging process, especially in distinguishing multiple stages of parasites.

0 papers0 benchmarksImages

fNIRS2MW (The Tufts fNIRS to Mental Workload Dataset)

The Tufts fNIRS to Mental Workload (fNIRS2MW) open-access dataset is a new dataset for building machine learning classifiers that can consume a short window (30 seconds) of multivariate fNIRS recordings and predict the mental workload intensity of the user during that window.

0 papers0 benchmarksTime series

Lemons quality control dataset

Lemon dataset has been prepared to investigate the possibilities to tackle the issue of fruit quality control. It contains 2690 annotated images (1056 x 1056 pixels). Raw lemon images have been captured using the procedure described in the following blogpost and manually annotated using CVAT.

0 papers0 benchmarksImages

Pan-STARRS (Panoramic Survey Telescope and Rapid Response System (Pan-STARRS))

Pan-STARRS is a system for wide-field astronomical imaging developed and operated by the Institute for Astronomy at the University of Hawaii. Pan-STARRS1 (PS1) is the first part of Pan-STARRS to be completed and is the basis for both Data Releases 1 and 2 (DR1 and DR2). The PS1 survey used a 1.8 meter telescope and its 1.4 Gigapixel camera to image the sky in five broadband filters (g, r, i, z, y).

0 papers0 benchmarks

CEAHB2021-5 (Chinese Ethnic Ancient Handwritten Books database)

Ancient books script identification of Chinese ethnic minorities with deep convolutional neural networks via multi-branch and spatial pyramid pooling

0 papers0 benchmarks

TLHDIBD2021 (Tai Le historical document image binarization dataset)

Hybrid-CBF: A hybrid classification and binarization framework for historical Tai Le document image binarization The binarization of historical documents is very important and more challenging than the binarization of ordinary documents. As a result of the serious noise pollution found on the historical Tai Le documents, a new hybrid classification and binarization framework (Hybrid-CBF) is proposed for the binarization of historical Tai Le document images. The Tai Le historical document image binarization dataset (TLHDIBD2021) containing 2,780 image pairs is constructed. Due to the different degrees of document background pollution, the single method has a poor effect on the binarization of historical Tai Le documents. First, Hybrid-CBF clusters the historical Tai Le document images according to the noise level estimation to obtain document images with different noise levels. Second, the corresponding optimal binarization method is used for historical Tai Le documents with different n

0 papers0 benchmarks
PreviousPage 643 of 1000Next