19,997 machine learning datasets
19,997 dataset results
The Modified Swiss Dwellings (MSD) dataset is an ML-ready dataset for floor plan generation and analysis at building-level scale. The MSD dataset is completely derived from the Swiss Dwellings database (v3.0.0). The MSD dataset contains highly-detailed 5372 floor plans of single- as well as multi-unit building complexes across Switzerland, hence extending the building scale w.r.t. of other well know floor plan datasets like the RPLAN dataset.
PPED: Periodic Phenomena Event-based Dataset The dataset features 12 one-second sequences of periodic phenomena (rotation - 01-06, flicker - 07-08, vibration - 09-10 and movement - 11-12) with GT frequencies ranging from 3.2Hz up to 2000Hz in file formats .raw and .hdf5.
TinyChirp dataset for model training, validation and testing
Datasets are listed in the repository's readme file. This one is extra and yields 20K+ items after filtering with a fuzzy parser.
OHLCVT stands for Open, High, Low, Close, Volume and Trades and represents the following trading information within each time frame (such as one minute, five minute, hourly, daily, etc.):
Dataset Card for ESG/DLT Named Entity Recognition Dataset This dataset contains named entities related to Distributed Ledger Technology (DLT) and Environmental, Social, and Governance (ESG) topics created to support research in these areas and at the intersection of these domains.
This dataset is a "part I" extension of the "Engineered cardiac microbundle time-lapse microscopy image dataset" and contains 732 experimental time-lapse image sequences of beating hiPSC-based cardiac microbundles using microbundle strain gauge platforms [1] ("Type1"). In "part II" extension, we include 808 experimental time-lapse image sequences of beating hiPSC-based cardiac microbundles using FibroTUG platforms [2] ("Type2").
This dataset is a "part II" extension of the "Engineered cardiac microbundle time-lapse microscopy image dataset" and contains 808 experimental time-lapse image sequences of beating hiPSC-based cardiac microbundles using FibroTUG platforms [1] ("Type2"). In "part I" extension, we include 732 experimental time-lapse image sequences of beating hiPSC-based cardiac microbundles using microbundle strain gauge platforms [2] ("Type1").
In the realm of document engineering and Natural Language Processing (NLP), the integration of digitally born catalogs into product design processes presents a novel avenue for enhancing information extraction and interoperability. This paper introduces CatalogBank, a dataset developed to bridge the gap between textual descriptions and other data modalities related to engineering design catalogs. We utilized existing information extraction methodologies to extract product information from PDF-based catalogs to use in downstream tasks to generate a baseline metric. Our approach not only supports the potential automation of design workflows but also overcomes the limitations of manual data entry and non-standard metadata structures that have historically impeded the seamless integration of textual and other data modalities. Through the use of DocumentLabeler, an open-source annotation tool adapted for our dataset, we demonstrated the potential of CatalogBank in supporting diverse documen
This dataset comprises 1-minute fingertip video recordings collected from 150 anemic patients, ranging from 6 months to 32 years of age, with hemoglobin levels between 4.3 gm/dL and 12.4 gm/dL. The videos were recorded using a smartphone’s camera and flashlight, designed to capture PPG (Photoplethysmography) signals, which are essential for non-invasive hemoglobin level estimation. Key Features:
This dataset includes User Story (or Issue) text descriptions, User Story titles, and Story Points from 33 software development projects, comprising a total of 20,479 User Stories (or issues) extracted from GitLab repositories, amounting to 12,262.7 Story Points. The mining process focused on GitLab’s top open-source projects that use agile software development methodologies and record task sizes in Story Points. Only tasks with the State attribute set to Closed and with the Weight attribute filled in were collected. The Weight field in GitLab is used to record the effort in Story Points. The data was mined between January 2023 and April 2023. The projects in the dataset have diverse characteristics, covering different programming languages, business domains, and geographic locations of the teams.
Science Journal for Kids Data This repository contains a dataset of abstracts from the Science Journal for Kids website and the original academic papers. It includes metadata such as titles, URLs, reading levels, and links to the full academic papers. The dataset is designed to support research and analysis of educational content tailored for young learners.
Science Journal for Kids Data This repository contains a dataset of abstracts from the Science Journal for Kids website and the original academic papers. It includes metadata such as titles, URLs, reading levels, and links to the full academic papers. The dataset is designed to support research and analysis of educational content tailored for young learners.
This R package, documented in a very similar way to the book R4DS, provides functions to replicate the original Stata results from the book An Advanced Guide to Trade Policy Analysis.
The Room environment - v2
Mapping urban large-area advertising structures using drone imagery and deep learning-based spatial data analysis.
A configurable synthetic dataset of simple shapes with ground truth concepts and known causal relationships between concepts and classes.
Data files with the information required to replicate all the experiments reported in the paper:
An adaption of the MVTec Anomaly Detection dataset, presented in the paper "Domain-independent detection of known anomalies".
Image corruptions modelling primary optical aberrations.