19,997 machine learning datasets
19,997 dataset results
Here we release the dataset (Multi_Channel_Grid, abbreviated as MC_Grid) used in our paper LIMUSE: LIGHTWEIGHT MULTI-MODAL SPEAKER EXTRACTION.
APEACH is the first crowd-generated Korean evaluation dataset for hate speech detection. Sentences of the dataset are created by anonymous participants using an online crowdsourcing platform DeepNatural AI.
This dataset contain online reviews gathered from google reviews written by north american casino users. explain motivations and summary of its content. Can be used to study user experience and relative research directions such as cultural impacts on latency of aspects, domain importance, sentiment analysis, opinion mining, aspect-based sentiment analysis, etc.
WHAMR_ext is an extension to the WHAMR corpus with larger RT60 values (between 1s and 3s)
The dataset contains 60,000 Stack Overflow questions from 2016-2020, classified into three categories:
The lists of Tweet IDs for the experiments of the article: This Sample seems to be good enough! Assessing Coverage and Temporal Reliability of Twitter's Academic API by Juergen Pfeffer, Angelina Mooseder, Luca Hammer, Oliver Stritzel, David Garcia.
The OU-ISIR Gait Database, Multi-View Large Population Database with Pose Sequence (OUMVLP-Pose) is meant to aid research efforts in the general area of developing, testing and evaluating algorithms for model-based gait recognition.
HeLa cells on a flat glass Dr. G. van Cappellen. Erasmus Medical Center, Rotterdam, The Netherlands
Simulated nuclei of HL60 cells stained with Hoescht
MDA231 human breast carcinoma cells infected with a pMSCV vector including the GFP sequence, embedded in a collagen matrix
This is the dataset used for classifying Gene-Disease relationship types from sentences. The dataset consists of 3 files:
DeePore is a deep learning workflow for rapid estimation of a wide range of porous material properties based on the binarized micro–tomography images. By combining naturally occurring porous textures we generated 17,700 semi–real 3–D micro–structures of porous geo–materials with the size of $256^3$ voxels and 30 physical properties of each sample are calculated using physical simulations on the corresponding pore network models.
To form the collection of nighttime RAW samples, we first selected a total of 150 images with the spatial resolution at 3464×5202 from the training and validation sets provided by the night image challenge. And then these RAW images are pre-processed to best produce noise-free samples using a notable CNN based denoiser. This is because nighttime imaging experiences a very challenging situation with heavy noises incurred by high ISO setting under poor illumination condition (e.g., underexposure).
Data-set from "PEEK-An LSTM Recurrent Network for Motion Classification from Sparse Data"
Using Council Data Project infrastructures (https://councildataproject.org), we assemble longitudinal municipal council meeting transcript data. This initial release of the Councils in Action dataset includes over 350 meetings of the city councils of Seattle Washington and Portland Oregon, and the county council of King County Washington.
Periodic Tic sounds (T0=1s) sampled at 16kHz with duration of nearly 10s.
This dataset consists of raw EEG data from 48 subjects who participated in a multitasking workload experiment utilizing the SIMKAP multitasking test. The subjects’ brain activity at rest was also recorded before the test and is included as well. The Emotiv EPOC device, with sampling frequency of 128Hz and 14 channels was used to obtain the data, with 2.5 minutes of EEG recording for each case. Subjects were also asked to rate their perceived mental workload after each stage on a rating scale of 1 to 9 and the ratings are provided in a separate file.
EEG signals from 60 users have been recorded whose age range lies between 6 and 55 years. Among all, there were 25 females and 35 male users. In general, all the participants were either school children or belonged to the socioeconomic cross section of the population with no medical history. The EEG recordings were acquired from all 14 electrodes operating at a sampling rate of 128 Hz. During recording, the participants were asked to comfortably sit on the chair with clear thoughts and a relaxed state.
Bitcoin is a peer-to-peer electronic payment system that popularized rapidly in recent years. Usually, we need to query the complete history of bitcoin blockchain data to acquire variables of economic meaning. This becomes increasingly difficult now with over 1.6 billion historical transactions on the Bitcoin blockchain. It is thus important to query Bitcoin transaction data in a way that is more efficient and provides economic insights. We apply cohort analysis that interprets bitcoin blockchain data using methods developed for population data in social science. Specifically, we query and process the Bitcoin transaction input and output data within each daily cohort. With this, we then create datasets and visualizations for some key indicators of bitcoin transactions, including the daily lifespan distributions of accumulated spent transaction output (STXO) and the daily age distributions of accumulated unspent transaction output (UTXO). We provide a computationally feasible approach t
USC-GRAD-STDdb comprises 115 video segments containing more than 25,000 annotated frames of HD 720p resolution (≈1280x720) with small objects of interest from 16 (≈4x4) to 256 (≈16x16) as pixel area. The length of the videos changes from 150 up to 500 frames. The size of every object is determined through the bounding box, so that a good annotation is of utmost importance for reliable performance metrics. As it may seem obvious, the smaller the object, the harder the annotation. The annotation has been carried out with the ViTBAT tool, adjusting the boxes as much as possible to the objects of interest in each video frame. In total, more than 56,000 ground truth labels have been generated.