19,997 machine learning datasets
19,997 dataset results
A dataset of books for very young children.
Probes to evaluate commonsense in language models.
Repository for UML-English data This repository contains the data used for "Extraction of UML Class Diagrams from Natural Language Specification" (Yang et al. 2022)
Checkpoints, generated EMA representations, audio outputs, and annotations for paper titled "Articulation GAN: Unsupervised modeling of articulatory learning"
Panoramic Video Panoptic Segmentation Dataset is a large-scale dataset that offers high-quality panoptic segmentation labels for autonomous driving. The dataset has labels for 28 semantic categories and 2,860 temporal sequences that were captured by five cameras mounted on autonomous vehicles driving in three different geographical locations, leading to a total of 100k labeled camera images.
Are Large Pre-Trained Language Models Leaking Your Personal Information? We analyze whether Pre-Trained Language Models (PLMs) are prone to leaking personal information. Specifically, we query PLMs for email addresses with contexts of the email address or prompts containing the owner's name.
6000 French user reviews from three applications on Google Play (Garmin Connect, Huawei Health, Samsung Health) are labelled manually. We selected four labels: rating, bug report, feature request and user experience.
none
Annotated Earth Observation dataset of extreme events
Unpaired dataset: The dataset is built by ourselves, and there are all real haze images from websites.
Equilibrium structures of the tautobase(reference) optimized at the level of theory of popular quantum chemical databases (QM9,PC9 and ANI-E). The structures were generated from the SMILES structures of the original publication and then optimized using Gaussian09. For simplicity, structures are divided on type 'A' and 'B'. The database consists of 1257 pairs (2514 molecules) for each database evaluated. The format of the database is .xyz
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
The dataset for this task is TAU Audio-Visual Urban Scenes 2021. The dataset contains synchronized audio and video recordings from 12 European cities in 10 different scenes.
111
Financial Language Understanding Evaluation is an open-source comprehensive suite of benchmarks for the financial domain. It contains benchmarks across 5 NLP tasks in financial domain as well as common benchmarks used in the previous research. The tasks are financial sentiment analysis, news headline classification, named entity recognition, structure boundary detection and question answering.
E2E Refined is a dataset for sentence classification. It consists of 40,560 examples for training, 4,489 for validation, and 4,555 for test. It is a refined version of the well-known MR-to-text E2E dataset where many deletion/insertion/substitution errors has been fixed.
The SNS data (Valente et al., 2013) is a four-wave survey conducted in Los Angeles county, the United States, which features a sample of 1,795 high-school students. The survey collected information about high-school students between grades 10 to 12, a majority of them self-identified as Hispanic. Among the collected information we have socio-economic status, demographics, social networks, and consumption of alcohol, tobacco, and marijuana–substance use.
Source: Text mining methodologies with R: An application to central bank texts
Images collected on an LED array microscope (also known as a Fourier ptychographic microscope) on 172 fields-of-view of frog blood smears. Two of the fields-of-view ( example_000000 and example_000001) have 85 intensity images under single LED illumination, and all fields-of-view have 8 intensity images, 4 taken with uniformly random patterns and 4 taken with pseudo-Dirichlet random patterns.
Virtual-PedCross-4667 is a dataset for pedestrian crossing prediction. It consists of 4667 video sequences, 2862 pedestrian crossing sequences and 1804 not-crossing sequences. Totally, 745k video frames with the resolution of 1280×720 are saved.