19,997 machine learning datasets
19,997 dataset results
Includes co-referent name string pairs along with their similarities.
The sample EEG dataset consists of the newborn EEG data recorded for the work published as:
The data contains the following attributes for Korea Stock Price Index (KOSPI) for January 2000–December 2016: 1. Date (YYYY.M(M).D(D)) 2. Opening Price for the date, PX_OPEN 3. Highest Price for the date, PX_HIGH 4. Lowest Price for the date, PX_LOW 5. Closing Price for the date, PX_LAST 6. Total volume traded on the date, PX_VOLUME
Raw StarCraft II data is subject to processing under the Blizzard end user license agreement (EULA), and in special cases Blizzard AI and Machine Learning License may be applied. Please refer to the materials listed below.
SC2EGSet: StarCraft II Esport Game State Dataset
Our TGRDB dataset was collected with a 180 fisheye RGB camera on-board of a moving tour-guide robot. A first dataset in tour-guide scenario. Statistical comparisons between TGRDB and existing datasets are refferred to https://arxiv.org/abs/2207.03726. We hope this dataset will drive the progress of research in service robotics, long-term multi-person tracking, and fine-grained or clothes-inconsistency person re-identification.
The CareerCoach 2022 gold standard is available for download in the NIF and JSON format, and draws upon documents from a corpus of over 99,000 education courses which have been retrieved from 488 different education providers.
This dataset contains dialogue lines from the games Knights of the Old Republic 1 & 2 and Neverwinter Nights 1. Some of the dialogue lines are marked as persuasive (which is when the player character is attempting a Persuade skill check.)
PDDL dataset of Rearrangement tasks in large-scale 3D scene graphs.
Highlights
MatriVasha the largest dataset of handwritten Bangla compound characters for research on handwritten Bangla compound character recognition. The proposed dataset contains 120 different types of compound characters that consist of 306,464 images written where 152,950 male and 153,514 female handwritten Bangla compound characters. This dataset can be used for other issues such as gender, age, district base handwriting research because the sample was collected that included district authenticity, age group, and an equal number of men and women.
The Mafia Dataset was created to model the behavior of deceptive actors in the context of the Mafia game, as described in the paper “Putting the Con in Context: Identifying Deceptive Actors in the Game of Mafia”. We hope that this dataset will be of use to others studying the effects of deception on language use.
ANTILLES is a part-of-speech tagging corpus based on UD_French-GSD which was originally created in 2015 and is based on the universal dependency treebank v2.0.
This dataset consists of Winograd schemas that test coreference resolution systems' ability to differentiate singular vs plural they/them pronouns. It consists of 4077 templates, each with a group of people, a singular person (which can be filled with a name or a generic "someone") and a single they/them pronoun to resolve.
Summary:
Article-Bias-Prediction Dataset The articles crawled from www.allsides.com are available in the ./data folder, along with the different evaluation splits.
This dataset concentrates on the activities of the crowd for a fine-grained image classification task, named as Crowd Activity dataset, as automatically understanding crowd activity is meaningful for social security. This dataset is newly collected, where the images are mainly searched on the Internet and collected from streets by mobile phones. All images in this dataset contain at least one text instance. The categories come from activities of daily living and demonstrations stimulated by hot events in recent years. Specifically, this dataset consists of 21 categories and 8785 images in total. The 21 categories broadly fall into two types: activities of daily living(i.e., celebrating Christmas, holding sport meeting, holding concert, celebrating birthday party, celebrity speech, teaching, graduation ceremony, picnic, press briefing, shopping, celebrating Thanks giving day) and demonstrations (i.e., protecting animals, protecting environment, appealing for peace, Brexit, COVID-19, ele
Hypertention Disease Medication dataset.
Loucount is a retail object detection and and counting dataset with rich annotations in retail stores, which consists of 50, 394 images with more than 1.9 million object instances in 140 categories
Audiogram data based on a "gold standard" audiometer and the uHear iOS application of 163 participants