19,997 machine learning datasets
19,997 dataset results
Benchmark to evaluate the capability of LMs to consolidate and recall information from multiple training documents.
Genre annotations for movies The file genre2movies.csv contains genre-movie tuples based on Wikidata annotations (https://www.wikidata.org/).
This dataset contains a collection of papers retrieved by using a PRISMA systematic review of Open Data and Public Domain data in Agriculture. This collection of papers uses, creates, or discusses about Open Data and Public Domain.
Multilingual human VALUE(MVALUE) is a multilingual dataset covering 7 concepts of human values: morality, deontology, utilitarianism, fairness, truthfulness, toxicity and harmfulness, each concept subset of it includes positive and negative texts that represent the two opposing directions of the concept. We performed translation on collected human value datasets from English into 15 non-English languages using Google Translate. These languages belong to various language families, including Indo-European (Catalan, French, Indonesian, Portuguese, Spanish), NigerCongo (Chichewa, Swahili), Dravidian (Tamil, Telugu), Uralic (Finnish, Hungarian), Sino-Tibetan (Chinese), Japonic (Japanese), Koreanic (Korean) and Austro-Asiatic (Vietnamese).
IUPAC Standards Online is a database built from IUPAC’s standards and recommendations, extracted from the journal Pure and Applied Chemistry (PAC).
FreeMan is the first large-scale multi-view human motion dataset under real scenarios. FreeMan was captured by synchro- nizing 8 smartphones across diverse scenarios. It comprises 11M frames from 8000 sequences, viewed from different perspectives. These sequences cover 40 subjects across 10 different scenarios, each with varying lighting conditions.
Mix of Minimal Optimal Sets (MMOS) of dataset has two advantages for two aspects, higher performance and lower construction costs on math reasoning.
ASCAD (ANSSI SCA Database) is a set of databases that aims at providing a benchmarking reference for the SCA community: the purpose is to have something similar to the MNIST database that the Machine Learning community has been using for quite a while now to evaluate classification algorithms performance.
ASCAD database version 2. This database contained the power consumption of a STM32 Cortex M4 microcrontroller (STM32F303RCT7) during 800.000 random AES encryptions. The AES encryptions are protected with shuffling and affine masking, and the implementation is available on https://github.com/ANSSI-FR/SecAESSTM32. The raw dataset is split into 8 files of 100.000 encryptions, and the extracted dataset contained the 800.000 preprocessed traces with additional metadata.
For the purpose of training and evaluating our intent classification model for electric automation, we curated a dataset consisting of intent-based user instructions. The dataset comprises a total of 14 intents, each associated with approximately 10 user instructions, resulting in a total of 140 instructions for electric automation. The intents were carefully selected to cover a diverse range of control commands and actions commonly encountered in electric automation scenarios. These intents include commands for turning on/off electrical appliances. Each user instruction in the dataset is labeled with its corresponding intent, allowing the model to learn the mapping between input instructions and their intended actions.
We propose ChaosBench, a large-scale, multi-channel, physics-based benchmark for subseasonal-to-seasonal (S2S) climate prediction. It is framed as a high-dimensional video regression task that consists of 45-year, 60-channel observations for validating physics-based and data-driven models, and training the latter. Physics-based forecasts are generated from 4 national weather agencies with 44-day lead-time and serve as baselines to data-driven forecasts. Our benchmark is one of the first to incorporate physics-based metrics to ensure physically-consistent and explainable models. We establish two tasks: full and sparse dynamics prediction.
Description This Dataset contains review information on Google map (ratings, text, images, etc.), business metadata (address, geographical info, descriptions, category information, price, open hours, and MISC info), and links (relative businesses) up to Sep 2021 in the United States.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
A new dataset for portrait harmonization based on the FFHQ. It contains real images, foreground masks, and synthesized composites.
This dataset is composed of 69 individual objects and 57 meaningful pairs. The objects cover a wide range of categories, including decor item, food, furniture, instrument, jewelry, luggage, person, pet, plant, plushie, scene, thing, toy, transportation, and wearable item.
A biomedical dataset supporting ontology enrichment from texts, by concept discovery and placement, adapting the MedMentions dataset (PubMed abstracts) with SNOMED CT of versions in 2014 and 2017 under the Diseases (disorder) sub-category and the broader categories of Clinical finding, Procedure, and Pharmaceutical / biologic (CPP) product.
This is the synthetic dataset that is introduced in the paper https://arxiv.org/abs/2403.03375. It allows fine-grain control on the spurious correlation strength, core and spurious feature hardness/complexity. Parity and staircase function are included in the codebase.
This is the static test data from the study "Global Geolocated Realtime Data of Interfleet Urban Transit Bus Iding" collected by GRD-TRT-BUF-4I. test-data-a.csv was collected from December 31, 2023 00:01:30 UTC to January 1, 2024 00:01:30 UTC. test-data-b.csv was collected from January 4, 2024 01:30:30 UTC to January 5, 2024 01:30:30 UTC. test-data-c.csv was collected from January 10, 2024 16:05:30 UTC to January 11, 2024 16:05:30 UTC.