19,997 machine learning datasets
19,997 dataset results
SentimentArcs’ reference corpus for novels consists of 25 narratives selected to create a diverse set of well recognized novels that can serve as a benchmark for future studies. The composition of the corpora was limited by the effect of copyright laws as well as historical imbalances. Most works were obtained from US and Australian Gutenberg Projects. The corpora is expected to grow in size and diversity over time.
The TREC Fair Ranking track evaluates systems according to how well they fairly rank documents. The 2020 focuses on scholarly search and fairly ranking academic abstracts and papers from authors belonging to different groups.
This dataset consists of six columns. The first four columns represent the input features (i.e., area, sensing range, transmission range, and the number of sensors). The last two columns represent the response variable or target variable (i.e., number of barriers (Gaussian) and number of barriers (Uniform)).
Wikipedia users activity for two language editions, Portuguese and Italian, for up to 8 January 2020.
This dataset contains data of 125 1-hour simulations of ship motion during various sea states performing random maneuvers in 4 degrees of freedom (surge-sway-yaw-roll). The original ship is a patrol ship developed by Perez et al. 1. We have extended it with a set of two symmetrically placed rudder propellers. Additionally, we simulate wind forces according to Isherwood's wind model 2. Wind-induced waves are generated with the JONSWAP spectrum 3 and the corresponding wave forces are then computed using wave force response amplitude operators (ROA).
The Herbarium 2022: Flora of North America is a part of a project of the New York Botanical Garden funded by the National Science Foundation to build tools to identify novel plant species around the world. The dataset strives to represent all known vascular plant taxa in North America, using images gathered from 60 different botanical institutions around the world.
This is a dataset of audiens comment for each KOL that uses Instagram as their campaign platform. The comments are scrapped and generated as csv through apify.com
This dataset contains samples of CTI (Cyber Threat Intelligence) data in natural language, labeled with the corresponding adversarial techniques from the MITRE ATT&CK framework.
This dataset consists of charge densities of individual snapshots from a molecular dynamics trajectory (DFT simulations?). We insert 8 ethylene carbonate molecules in the simulation box. To quickly explore a large part of the configurational space we put Hookean constraints on the molecular bonds (to maintain molecular identity such that molecules are not torn apart at such high temperature) and run Langevin molecular dynamics with thermostat temperature of 3000 K. The simulation was run for 12380 steps of 0.5 fs.
The datasets of "Reinforcement Learning-enhanced Shared-account Cross-domain Sequential Recommendation" (TKDE 2022)
A large-scale dataset of measurements of ETSI ITS-G5 Dedicated Short Range Communications (DSRC) is presented. Our dataset consists of network interactions happening between two On-Board Units (OBUs) and four Road Side Units (RSUs). Each OBU was fitted onto a vehicle. The two vehicles have been driven across the Innovate UK-funded FLOURISH Test Track encompassing key roads in the center of Bristol, UK. As for the RSUs, they were located at fixed locations around the track. Each RSU and OBU is equipped with two transceivers operating at different frequencies. During our experiments, each transceiver broadcast Cooperative Awareness Messages (CAMs) every 10ms to the neighboring RSUs and or OBUs.
This dataset consists of ~350k JPEG images of streetlight columns installed on a public road infrastructure located in the city of Bristol, UK.
To evaluate the performance on 4K burst images/video, we collect several clips from website. The dataset can be download from : https://drive.google.com/file/d/1YDljUONvyKUO24smTx__CUH_4Zxhle09/view?usp=sharing
Introduction The Niko Chord Progression Dataset is used in AccoMontage2. It contains 5k+ chord progression pieces, labeled with styles. There are four styles in total: Pop Standard, Pop Complex, Dark and R&B. Some progressions have an 'Unknown' style. Some statistics are provided below.
EMC Dutch clinical corpus contains four types of anonymized clinical documents: entries from general practitioners, specialists’ letters, radiology reports, and discharge letters. The identified UMLS terms in the corpus are annotated for negation, temporality, and experiencer properties.
A dataset for online novel recommendation.
The PKU Sketch Re-ID dataset is constructed by National Engineering Laboratory for Video Technology (NELVT), Peking University.
A new Actor-identified A-AVA dataset based on the existing AVA dataset and the TAO dataset, by assigning the unique actor identity and actions to each actor.
The datasets of "Time Interval-enhanced Graph Neural Network for Shared-account Cross-domain Sequential Recommendation" (TNNLs 2022)
We present the Single-dish PARKES data sets for finding the uneXpected (SPARKESX), a compilation of real and simulated high-time resolution observations. SPARKESX comprises three mock surveys from the Parkes ''Murriyang'' radio telescope. A broad selection of simulated and injected expected signals (such as pulsars, fast radio bursts), poorly known signals (such as the features expected from flare stars) and unknown unknowns are generated for each survey. We provide a baseline by presenting how successful a typical pipeline based on the standard pulsar search software, PRESTO, is at finding the injected signals.