19,997 machine learning datasets
19,997 dataset results
Minsk2019 ALS database is a dataset collected in Republican Research and Clinical Center of Neurology and Neurosurgery (Minsk, Belarus). A total of 54 speakers were recorded, with 39 healthy speakers (23 males, 16 females) and 15 ALS patients with signs of bulbar dysfunction (6 males, 9 females). It is designed for the task of ALS Detection.
Dataset included measuring static tension under 2 kg load in different points of the CB and measurements in dynamic conditions. The latter conditions presumed the range of the linear belt speeds between nu_1 = 0.5 and nu_max = 1.7 m/s. 400 Hz unified sampling frequency for the experiments. It corresponded with 140 samples.
MENYO-20k is the first multi-domain parallel corpus with a special focus on clean orthography for Yorùbá--English with standardized train-test splits for benchmarking.
The National Health and Nutrition Examination Survey (NHANES) provides data on the health and environmental exposure of the non-institutionalized US population. Such data have considerable potential to understand how the environment and behaviors impact human health. These data are also currently leveraged to answer public health questions such as prevalence of disease. However, these data need to first be processed before new insights can be derived through large-scale analyses. NHANES data are stored across hundreds of files with multiple inconsistencies. Correcting such inconsistencies takes systematic cross examination and considerable efforts but is required for accurately and reproducibly characterizing the associations between the exposome and diseases (e.g., cancer mortality outcomes). Thus, we developed a set of curated and unified datasets and accompanied code by merging 614 separate files and harmonizing unrestricted data across NHANES III (1988-1994) and Continuous (1999-20
Quality of Experience Evaluation of Interactive Virtual Environments with Audiovisual Scenes (QoEVAVE) provides an initial audiovisual database consiting of 12 sequences capturing real-life nature and urban scenes. The maximum video resolution is 7680x3840 (8k) at 60 frames-per-second, with 4th-order Ambisonics spatial audio (4OA). All video sequences are recorded with a minumum target duration of 60 seconds and designed to represent real-life settings for systematically evaluating various dimensions of uni-/multimodal perception, cognition, behavior, and quality of experience (QoE) in a controlled virtual environment. This database serves as a novel high-quality reference material with an equal focus on auditory and visual sensory information within the QoE community.
HPointLoc is a dataset designed for exploring capabilities of visual place recognition in indoor environment and loop detection in simultaneous localization and mapping. It is based on the popular Habitat simulator from 49 photorealistic indoor scenes from the Matterport3D dataset and contains 76,000 frames.
Fine-Grained Vehicle Detection (FGVD) is a dataset for fine-grained vehicle detection captured from a moving camera mounted on a car. The FGVD dataset is challenging as it has vehicles in complex traffic scenarios with intra-class and inter-class variations in types, scale, pose, occlusion, and lighting conditions.
MTNeuro is a multi-task neuroimaging benchmark built on volumetric, micrometer-resolution X-ray microtomography images spanning a large thalamocortical section of mouse brain, encompassing multiple cortical and subcortical regions.
The dataset of images is built upon a collection of 454 samples kindly provided by the urology department of the Hospital Universitary de Bellvitge (Barcelona, Spain) in the time span of several years. They cover all the main 9 classes but cystine, for which just 4 samples were available so we discarded this class as mentioned above. As for the rest, we tried to get all second scheme classes balanced and, at the same time, to record as much examples as possible to account for intraclass.
This dataset tests the capabilities of language models to correctly capture the meaning of words denoting probabilities (WEP), e.g. words like "probably", "maybe", "surely", "impossible".
This dataset provides wireless measurements from two industrial testbeds: iV2V (industrial Vehicle-to-Vehicle) and iV2I+ (industrial Vehicular-to-Infrastructure plus sensor).
Dataset Summary The dataset used to train and evaluate TunesFormer is collected from two sources: The Session and ABCnotation.com. The Session is a community website focused on Irish traditional music, while ABCnotation.com is a website that provides a standard for folk and traditional music notation in the form of ASCII text files. The combined dataset consists of 285,449 ABC tunes, with 99\% (282,595) of the tunes used as the training set and the remaining 1\% (2854) used as the evaluation set.
AviationQA is introduced in the paper titled- There is No Big Brother or Small Brother: Knowledge Infusion in Language Models for Link Prediction and Question Answering
Hand-disambiguation of a sample of U.S. patents inventor mentions from PatentsView.org.
We've all been there: a light turns green and the car in front of you doesn't budge. Or, a previously unremarkable vehicle suddenly slows and starts swerving from side-to-side.
We have prepared a dataset, ParagraphOrdreing, which consists of around 300,000 paragraph pairs. We collected our data from Project Gutenberg. We have written an API for gathering and pre-processing in order to have the appropriate format for the defined task. Each example contains two paragraphs and a label that determines whether the second paragraph comes really after the first paragraph (true order with label 1) or the order has been reversed.
-Tab 1 (Carboxylase table): This expanded table contains additional information for carboxylase classes and splits them into individual examples.
-Tab 1 Rubisco forms: This excel sheet contains one row for every rubisco form considered in this review (some forms like IAq and IAc from [31] are not considered separately because they are only phylogenetically separated in the small subunits). For each form we include the pathway in which it functions, the chemical reaction catalyzed - for known types of Form IV RLPs (rubisco-like proteins), an example sequence of the protein and a citation. We also include an example structure if available.
This .csv file contains all of the sequences used in the phylogenetic analysis (see above). For form annotation sequences under 360aa were excluded and no upper limit was set. Sequences removed by trimAL (using a gap threshold of 0.1) are labeled as “Unannotated” - many of them may not be actual rubiscos. Sequences that are too short are labeled as such. Some sequences will have a form indicated but not a subform, for instance, some sequences are labeled as Form III but with no indicated subform because they do not fit into an established subclade. We used a tree made from a 65% identity dereplication (using CD-HIT with standard parameters, File S4). The tree was produced as described above using IQTree with the following parameters: -bb 1000 -m MFP -safe. The tree was rooted just past the Form IIIA clade so that all bona fide rubiscos form one clade and all RLPs form another clade. There are a few branches in between that we consider to be RLPs.
This tree was generated as indicated above in the methods. The model chosen by the algorithm was LG_F_R10.