19,997 machine learning datasets
19,997 dataset results
J. W. Pennebaker and L. A. King, “Linguistic styles: Language use as an individual difference,” J. Pers. Soc. Psychol., vol. 77, no. 6, pp. 1296–1312, Dec. 1999, doi: 10.1037/0022-3514.77.6.1296.
ECTF is a dataset for Twitter fake news detection in the Covid-19 domain.
The Extended UCF Crime extends the UCF Crime data set that consists of 13 anomaly classes. The extension adds two different anomaly classes to the data set, which are ”molotov bomb” and ”protest” classes. It also adds 33 videos to the fighting class. In total, the extension adds 216 videos to the training set, 17 videos to the test set.
This dataset was constructed from an analysis of about 1.5 million apps from Google Play to identify a set of common libraries, to facilitate Android app analysis. It contains 1,113 libraries supporting common functionalities and 240 libraries for advertisement.
The Workflow Trace Archive (WTA) is an open-access archive of workflow traces from diverse computing infrastructures. The WTA includes >48 million workflows captured from >10 computing infrastructures, representing a broad diversity of trace domains and characteristics.
The BLEBeacon dataset is a collection of Bluetooth Low Energy (BLE) advertisement packets/traces generated from BLE beacons carried by people following their daily routine inside a university building for a whole month. A network of Raspberry Pi 3 (RPi)-based edge devices were deployed inside a multi-floor facility continuously gathering BLE advertisement packets and storing them in a cloud-based environment. The focus is on presenting a real-life realization of a location-aware sensing infrastructure, that can provide insights for smart sensing platforms, crowd-based applications, building management, and user-localization frameworks.
AdobeIndoorNav is a dataset collected in real-world to facilitate the research in DRL based visual navigation. The dataset includes 3D reconstruction for real-world scenes as well as densely captured real 2D images from the scenes. It provides high-quality visual inputs with real-world scene complexity to the robot at dense grid locations.
This dataset contains 17,090 audio clips of length 30 seconds sampled from archives collected from 6 Guinean radio stations. The broadcasts consist of news and various radio shows in languages including French, Guerze, Koniaka, Kissi, Kono, Maninka, Mano, Pular, Susu, and Toma. Some radio shows include phone calls, background and foreground music, and various noise types. We collected this dataset for the purpose of unsupervised speech representation learning. A validation set of 300 tagged audio clips is also included.
This dataset contains 10,083 recorded utterances in French, Maninka, Pular and Susu from 49 speakers (16 female and 33 male) ranging from 5 to 76 years old on a variety of devices.
BuGL is a large-scale cross-language dataset for bug localization in code. BuGL constitutes of more than 10,000 bug reports drawn from open-source projects written in four programming languages, namely C, C++, Java, and Python. The dataset consists of information which includes Bug Reports and Pull-Requests. BuGL aims to unfold new research opportunities in the area of bug localization.
BIRD (Big Impulse Response Dataset) is an open dataset that consists of 100,000 multichannel room impulse responses (RIRs) generated from simulations using the Image Method, making it the largest multichannel open dataset currently available. These RIRs can be used to perform efficient online data augmentation for scenarios that involve two microphones and multiple sound sources.
The dataset consists of images of 158 filled out bank checks containing various complex backgrounds, and handwritten text and signatures in the respective fields, along with both pixel-level and patch-level segmentation masks for the signatures on the checks. Please visit the dataset homepage for more details.
The dataset contains 7,601 Gab posts classified on three different aspects: abuse presence or not, abuse severity and abuse target.
This is a dataset of 22.5 hours of synthesized audio using the open-source learnfm clone of the DX7 FM synthesizer, based upon 31K presets from Bobby Blue. These represent "natural'' synthesis sounds---i.e.presets devised by humans.
This is a large-scale dataset collected from WhatsApp public groups. It has been created from 178 public groups containing around 45K users and 454K messages. This dataset allows researchers to ask questions like (i) Are WhatsApp groups a broadcast, multicast or unicast medium? (ii) How interactive are users, and how do these interactions emerge over time? (iii) What geographical span do WhatsApp groups have, and how does geographical placement impact interaction dynamics? (iv) What role does multimedia content play in WhatsApp groups, and how do users form interaction around multimedia content? (v) What is the potential of WhatsApp data in answering further social science questions, particularly in relation to bias and representability?
THÖR is a dataset with human motion trajectory and eye gaze data collected in an indoor environment with accurate ground truth for position, head orientation, gaze direction, social grouping, obstacles map and goal coordinates. THOR also contains sensor data collected by a 3D lidar and involves a mobile robot navigating the space.
FacebookVideosLive18 dataset includes 1,000,000 Facebook live videos with their metadata (title, source, length, creation time, description, etc.), broadcasters locations and viewers locations. We are using a set of synchronised scripts that allow to have a global view of the real time streaming system every 3 minutes. We believe that our dataset is the first that tracks the locations and behaviors of live viewers. We expect FacebookVideosLive18 to support various trending research areas such as cloud computing, multimedia data allocation, multi-cloud allocation, edge computing, edge caching and transcoding, data analytics, etc.
FTR-18 is a multilingual rumour dataset on football transfer news. Transfer rumours are continuously published by sports media. They can both harm the image of player or a club or increase the player's market value. The proposed dataset includes transfer articles written in English, Spanish and Portuguese. It also comprises Twitter reactions related to the transfer rumours. FTR-18 is suited for rumour classification tasks and allows the research on the linguistic patterns used in sports journalism.
This is a dataset of 10.6 million GitHub projects that are copies of others, and link each record with the project's ultimate parent. The ultimate parents were derived from a ranking along six metrics. The related projects were calculated as the connected components of an 18.2 million node and 12 million edge denoised graph created by directing edges to ultimate parents. The graph was created by filtering out more than 30 hand-picked and 2.3 million pattern-matched clumping projects. Projects that introduced unwanted clumping were identified by repeatedly visualizing shortest path distances between unrelated important projects.
PersianQA: a dataset for Persian Question Answering Persian Question Answering (PersianQA) Dataset is a reading comprehension dataset on Persian Wikipedia. The crowd-sourced the dataset consists of more than 9,000 entries. Each entry can be either an impossible-to-answer or a question with one or more answers spanning in the passage (the context) from which the questioner proposed the question. Much like the SQuAD2.0 dataset, the impossible or unanswerable questions can be utilized to create a system which "knows that it doesn't know the answer".