19,997 machine learning datasets
19,997 dataset results
The CANDOR corpus is a large, novel, multimodal corpus of 1,656 recorded conversations in spoken English. This 7+ million word, 850 hour corpus totals over 1TB of audio, video, and transcripts, with moment-to-moment measures of vocal, facial, and semantic expression, along with an extensive survey of speaker post conversation reflections.
Logic synthesis is a challenging and widely-researched combinatorial optimization problem during integrated circuit (IC) design. It transforms a high-level description of hardware in a programming language like Verilog into an optimized digital circuit netlist, a network of interconnected Boolean logic gates, that implements the function. Spurred by the success of ML in solving combinatorial and graph problems in other domains, there is growing interest in the design of ML-guided logic synthesis tools. Yet, there are no standard datasets or prototypical learning tasks defined for this problem domain. Here, we describe OpenABC-D,a large-scale, labeled dataset produced by synthesizing open source designs with a leading open-source logic synthesis tool and illustrate its use in developing, evaluating and benchmarking ML-guided logic synthesis. OpenABC-D has intermediate and final outputs in the form of 870,000 And-Inverter-Graphs (AIGs) produced from 1500 synthesis runs plus labels such a
NELA-GT-2021 is the fourth installment of the NELA-GT datasets, NELA-GT-2021. The dataset contains 1.8M articles from 367 outlets between January 1st, 2021 and December 31st, 2021. Just as in past releases of the dataset, NELA-GT-2021 includes outlet-level veracity labels from Media Bias/Fact Check and tweets embedded in collected news articles.
A small dataset from the Inductive Link Prediction Challenge 2022. Training graph contains 10K entities, 96 relations, 78K triples. Inference graph contains 7K entities, 96 relations, 21K triples. Validation and test triples to predict belong to the inference graph.
A large dataset from the Inductive Link Prediction Challenge 2022. Training graph contains 46K entities, 130 relations, 202K triples. Inference graph contains 30K entities, 130 relations, 77K triples. Validation and test triples to predict belong to the inference graph.
This dataset is for evaluating the task of Black-box Multi-agent Integration which focuses on combining the capabilities of multiple black-box conversational agents at scale. It provides data to explore two main frameworks of exploration: question agent pairing and question response pairing.
The database was acquired using a thermographic camera TESTO 880-3. This camera is equipped with an uncooled detector and has a spectral sensitivity range from 8 to 14 μm. It has a removable German optic lens. It provides the following main features:
A comprehensive set of all Slovenian tweets posted in the 2018-2020 period, with retweet links and assigned hate speech classes. Available at a public language resource repository CLARIN.SI.
Synthetic log data suitable for evaluation of intrusion detection systems, federated learning, and alert aggregation. Each of the 8 datasets corresponds to a testbed representing a small enterprise network including mail server, file share, WordPress server, VPN, firewall, etc. Normal user behavior is simulated to generate background noise over a time span of 4-6 days. At some point, a sequence of attack steps are launched against the network. Log data is collected from all hosts and includes Apache access and error logs, authentication logs, DNS logs, VPN logs, audit logs, Suricata logs, network traffic packet captures, horde logs, exim logs, syslog, and system monitoring logs. Attacks include scans (nmap, WPScan, dirb), webshell upload, password cracking, privilege escalation, remote command execution, and data exfiltration.
Extended Vehicle Energy Dataset (eVED) is an extended version of the Vehicle Energy Dataset (VED), which is a large-scale dataset for vehicle energy consumption analysis. Compared with its original version, the extended VED (eVED) dataset is enhanced with accurate vehicle trip GPS coordinates, serving as a reliable basis to associate the VED trip records with external information e.g., road speed limit and intersections, from accessible map services to accumulate attributes that is relevant and essential in analyzing vehicle energy consumption.
The database was acquired using a thermographic camera TESTO 882-3 equipped with an uncooled detector and a spectral sensitivity range from 8 to 14 μm. It has a removable German optic lens with these main features:
A new dataset consisting of 64 people with different expressions and hairstyles.
This dataset contains spatiotemporal sequences of SST generated by the NEMO ocean engine. The observations correspond to 250 randomly selected data sites within a [0, 550] × [100, 650] square cropped from the area between 50 deg N − 65 deg N and 75W deg − 10W deg starting from 01-01-2016 to 12-31-2017. The data is divided into 24 sequences, each lasting 30 days (extra days in each month are truncated). Data corresponding to 2016 are used for training and the rest is used for validation and testing, in the equal sequential split.
The dataset contains full-spectral autofluorescence lifetime microscopic images (FS-FLIM) acquired on unstained ex-vivo human lung tissue, where 100 4D hypercubes of 256x256 (spatial resolution) x 32 (time bins) x 512 (spectral channels from 500nm to 780nm). This dataset associates with our paper "Deep Learning-Assisted Co-registration of Full-Spectral Autofluorescence Lifetime Microscopic Images with H&E-Stained Histology Images" (https://arxiv.org/abs/2202.07755) and "Full spectrum fluorescence lifetime imaging with 0.5 nm spectral and 50 ps temporal resolution" (https://doi.org/10.1038/s41467-021-26837-0). The FS-FLIM images provide transformative insights into human lung cancer with extra-dimensional information. This will enable visual and precise detection of early lung cancer. With the methodology in our co-registration paper, FS-FLIM images can be registered with H&E-stained histology images, allowing characterisation of tumour and surrounding cells at a celluar level with abs
Dataset used in research submitted to ICPC ERA 2022
This newly curated synthetic dataset specifies an additional reference region to guide image harmonization. There are 118,287 training images and 959 test images. The dataset consists of objects, backgrounds, and people.
This dataset contains transcriptions of the electric guitar performance of 240 tablatures, rendered with different tones. The goal is to contribute to automatic music transcription (AMT) of guitar music, a technically challenging task.
A large-scale, first-of-its-kind database aimed at generating a better understanding of the way children interact with mobile devices during their development process. ChildCIdbv1 comprises data collected from 438 children, from 18 months to 8 years old, encompassing the first three development stages of Piaget's theory. Data collected spans interaction with screens using both finger and pen stylus, information regarding the previous experience of the child with mobile devices, the child’s grade level, and whether attention-deficit/hyperactivity disorder (ADHD) is present.
VidHarm is a professionally annotated dataset for detection of harmful content in video. Include 3589 annotate video clips from a variety of film trailers. In contrast to previous approaches which mostly use meta data from long sequences, it uses the raw video and focus on short clips.
Dataset built from partial reconstructions of real-world indoor scenes using RGB-D sequences from ScanNet, aimed at estimating the unknown position of an object (e.g. where is the bag?) given a partial 3D scan of a scene. The dataset mostly consists of bedrooms, bathrooms, and living rooms. Some room types like closet and gym only have a few instances.