TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

CBCT Walnut (Cone-Beam X-Ray CT Data Collection Designed for Machine Learning)

The scans are performed using a custom-built, highly flexible X-ray CT scanner, the FleX-ray scanner, developed by XRE nvand located in the FleX-ray Lab at the Centrum Wiskunde & Informatica (CWI) in Amsterdam, Netherlands. The general purpose of the FleX-ray Lab is to conduct proof of concept experiments directly accessible to researchers in the field of mathematics and computer science. The scanner consists of a cone-beam microfocus X-ray point source that projects polychromatic X-rays onto a 1536-by-1944 pixels, 14-bit flat panel detector (Dexella 1512NDT) and a rotation stage in-between, upon which a sample is mounted. All three components are mounted on translation stages which allow them to move independently from one another.

0 papers0 benchmarks3D, Medical

ADFI (Anomaly Detection Datasets for Visual Inspection)

ADFI Dataset is an image dataset for anomaly detection methods with a focus on industrial inspection. Each category sub dataset comprises a training set of images and a test set of images with various kinds of defects as well as images without defects.

0 papers0 benchmarksImages

Roman Republican Coin Dataset

Based on Crawford’s work, we collect the most diverse and extensive image dataset of the reverse sides. For most of the Roman Republic coin classes, the obverse side depicts more discriminative information than the observe side. Our dataset has 228 motif classes, including 100 classes that are the main classes for training and testing, which we call the main dataset RRCD-Main. The images of the additional 128 classes constitute the disjoint test set, RRCD-Disjoint, which we allocate to assess the generalization ability of our models. Therefore, the training and testing can be evaluated on completely disjoint datasets. To the best of our knowledge, RRCD is the most diverse dataset proposed while it is the largest dataset of the Roman Republican coins

0 papers0 benchmarks

Supplementary material (Funding Covid-19 research: Insights from an exploratory analysis using open data infrastructures - Supplementary material)

Funding Covid-19 research: Insights from an exploratory analysis using open data infrastructures - Supplementary material

0 papers0 benchmarks

SMCOVID19-CT (Contact Tracing Data (from Italian SM-COVID-19 App))

We present a real data analysis of a CT experiment that was conducted in Italy for 8 months and involved more than 100,000 CT app users.

0 papers0 benchmarksTabular, Texts

Dafonts Free

This is a dataset of 18624 fonts labeled as 100% Free and Public domain / GPL / OFL on https://www.dafont.com/ with .ttf and .otf extensions.

0 papers0 benchmarks

FloW

Marine wastes are severely threatening marine animals and their habitat, also causing an impact on human life through toxic substances transportation and accumulation. To prevent the wastes especially the plastic trash from getting into the ocean, it is essential to detect and clean the floating wastes in inland waters efficiently like in rivers, lakes, and canals.

0 papers0 benchmarks

Moroccan Monay dataset

A dataset of all Moroccan money

0 papers0 benchmarks

SportsSum

SportsSum is a Chinese sports game summarization dataset that contains 5,428 soccer games of live commentaries and the corresponding news articles.

0 papers0 benchmarksTexts

IEEE-CIS Fraud Detection

Can you detect fraud from customer transactions? Imagine standing at the check-out counter at the grocery store with a long line behind you and the cashier not-so-quietly announces that your card has been declined. In this moment, you probably aren’t thinking about the data science that determined your fate.

0 papers0 benchmarksTabular

STEM-ECR

Grounding Scientific Entity References in STEM Scholarly Content to Authoritative Encyclopedic and Lexicographic Sources The STEM ECR v1.0 dataset has been developed to provide a benchmark for the evaluation of scientific entity extraction, classification, and resolution tasks in a domain-independent fashion. It comprises annotations for scientific entities in scientific Abstracts drawn from 10 disciplines in Science, Technology, Engineering, and Medicine. The annotated entities are further grounded to Wikipedia and Wiktionary, respectively.

0 papers0 benchmarksTexts

endless forams

This dataset was built based on a subset of foraminifer samples from the Yale Peabody Museum (YPM) Coretop Collection and the Natural History Museum, Lon- don (NHM) Henry A. Buckley Collection.

0 papers0 benchmarks

Dataset of Distribution Transformers at Cauca Department (Colombia)

Dataset contains 16.000 electric power distribution transformers from Cauca Department (Colombia). They are distributed in rural and urban areas of 42 municipalities. The information covers 2019 and 2020 years, has 6 categorical variables and 5 continuous variables. First ones correspond to: location, self-protected, removable connector, criticality according to ceraunic level, client and installation type. Second ones are transformer power, burn rate, users number, unsupplied electricity and secondary lines length.

0 papers0 benchmarks

TRECVID-AVS21 (V3C1)

The dataset has been designed to represent true web videos in the wild, with good visual quality and diverse content characteristics, The test video collection for TRECVID-AVS2019-TRECVID-AVS2021, which contains 1,082,649 web video clips, with even more diverse content, no predominant characteristics and low self-similarity.

0 papers0 benchmarksTexts, Videos

SEN (Sentiment analysis of Entities in News headlines)

SEN is a novel publicly available human-labelled dataset for training and testing machine learning algorithms for the problem of entity level sentiment analysis of political news headlines.

0 papers0 benchmarksTexts

Pistachio Image Dataset

Citation Request : 1. OZKAN IA., KOKLU M. and SARACOGLU R. (2021). Classification of Pistachio Species Using Improved K-NN Classifier. Progress in Nutrition, Vol. 23, N. 2, pp. DOI:10.23751/pn.v23i2.9686. (Open Access) https://www.mattioli1885journals.com/index.php/progressinnutrition/article/view/9686/9178

0 papers0 benchmarksImages, Texts

Anshita (Anshita Dhoot)

Potential cases

0 papers0 benchmarks

Four Shapes

This dataset contains 16,000 images of four shapes; square, star, circle, and triangle. Each image is 200x200 pixels.

0 papers0 benchmarks

Fluo-N3DL-TRIC

Developing Tribolium Castaneum embryo (3D cartographic projection)

0 papers0 benchmarks

Fluo-C3DH-A549-SIM

Simulated GFP-actin-stained A549 Lung Cancer cells embedded in a Matrigel matrix

0 papers0 benchmarks
PreviousPage 645 of 1000Next