19,997 machine learning datasets
19,997 dataset results
This dataset contains 44,001 Bengali comments, curated to detect cyberbullying using Natural Language Processing (NLP) techniques. Each comment is labeled by experts, categorizing different forms of harassment and offensive behavior. The dataset enables the identification of inappropriate content, ranging from mild to severe harassment, facilitating precise classification and analysis. This resource is designed for researchers and developers working on cyberbullying detection, sentiment analysis, and content moderation in Bengali text.
Description The F1 Regulations, Safety, and Racing Performance dataset provides an overview of key factors influenced by the evolving Fédération Internationale de l'Automobile (FIA) regulations from 1990 to 2023. This dataset includes metrics such as the number of teams, drivers, races, fatalities, car weight, DRS implementation, and overtakes. It tracks the introduction of new regulations each season, especially those impacting aerodynamics, making it an essential resource for analyzing the long-term effects of regulatory changes on safety, racing dynamics, and overall spectacle in Formula 1.
SimNICT is the first dataset for training universal non-ideal measurement CT (NICT) enhancement models.
A high-quality dataset forms the foundation for machine learning-based predictions of structural load capacity. Therefore, this study collected 222 sets of rectangular section reinforced UHPC beams flexural capacity test data,The key factors that affect the UHPC beam’s flexural capacity include geometric parameters, material properties, and reinforcement details. Eight fundamental structural features that significantly impact the flexural capacity of UHPC beams were selected as input parameters for machine learning: beam section width (b), beam section height (h), compressive strength of UHPC cubes (fcu), tensile strength of UHPC (ft), volume fraction of steel fibers (Vf), aspect ratio of steel fibers (Lf / Df), longitudinal reinforcement ratio (ρs), and yield strength of longitudinal reinforcement (fy). The output is the flexural capacity (Mu).
该数据集是一个全面而多样化的驾驶员行为监测数据集,其中包括来自美洲,亚洲和非洲的26名不同种族,肤色和性别(13名男性和13名女性)的参与者。数据集中的所有图像都是由固定在汽车仪表板上的摄像头拍摄的,所有图像都是RGB像素。该数据集共包含22424张图像。
This dataset contains 4606 articles from 1996 to 2024 that were presented in MIE (Medical Informatics Europe Conference) conferences. This data was extracted from PubMed and topic extraction and affiliation parsing were done on it.
This set consists of a longitudinal collection of 150 subjects aged 60 to 96. Each subject was scanned on two or more visits, separated by at least one year for a total of 373 imaging sessions. For each subject, 3 or 4 individual T1-weighted MRI scans obtained in single scan sessions are included. The subjects are all right-handed and include both men and women. 72 of the subjects were characterized as nondemented throughout the study. 64 of the included subjects were characterized as demented at the time of their initial visits and remained so for subsequent scans, including 51 individuals with mild to moderate Alzheimer’s disease. Another 14 subjects were characterized as nondemented at the time of their initial visit and were subsequently characterized as demented at a later visit.
This dataset is part of my bachelor thesis project. It was created by combining multiple open-source datasets from RoboFlow Universe as well as manual annotation.
Saarbruecken Voice Database contains voice and EGG recordings of patients diagnosed with voice disorder, as well as healthy persons.
Sentinel-1 SAR samples from a set of manually chosen points across the world from different times polarization and orbital pass. These set of images can help to extract time, polarization and orbital pass invariant features from SAR images. For each single file there is a numpy saved file specifying it's detailed information like location polarization and etc. A list of all locations are also available in the training folder. The other two folder are just for test and playing around. They don't have the information files.
We introduce low-light image enhancement benchmark dataset “Low-light Images of Streets (LoLI-Street),” which contains three subsets: train, validation, and test. The train and validation sets consist of 30k and 3k paired low-light and high-light images, respectively, and the real low-light test set (RLLT) contains 1k images under real-world low-light conditions, totaling 33k images.
This dataset is from Hartmann and Kemmerzell (2010), who, among other things, analyze the causes of the emergence of party ban provisions in sub-Saharan Africa.
The CICIoMT2024 dataset is a comprehensive dataset designed for cybersecurity research focused on the Internet of Medical Things (IoMT). Developed by the Canadian Institute for Cybersecurity, it simulates realistic IoMT network traffic, representing the diverse and evolving threats faced by connected healthcare devices. The dataset comprises network traffic data from various IoMT devices, with labeled instances for 18 distinct types of cyberattacks, alongside benign traffic data.
Media-Text dataset comprising images of banners, posters, covers and another images characterised for media industry.
The Turkish Scene Text Recognition (TS-TR) dataset was primarily developed to fill the gap in non-English text recognition resources, specifically addressing the unique challenges presented by the Turkish language, such as special characters and diacritics. This dataset mirrors real-world conditions with texts displayed in various fonts, sizes, orientations, and complex backgrounds from multiple urban and rural environments. Such diversity ensures the training of models that can generalize across different scenarios, including varying lighting conditions and complex visual layouts.
Dota 2 is a popular computer game with two teams of 5 players. At the start of the game, each player chooses a unique hero with different strengths and weaknesses. Predict the winning team.
Of the 12,330 sessions in the dataset, 84.5% (10,422) were negative class samples that did not end with shopping, and the rest (1908) were positive class samples ending with shopping.
TURSpider is a novel Turkish Text-to-SQL dataset that includes complex queries, akin to those in the original Spider dataset. TURSpider dataset comprises two main subsets: a dev set and a training set, aligned with the structure and scale of the popular Spider dataset. The dev set contains 1034 data rows with 1023 unique questions and 584 distinct SQL queries. In the training set, there are 8659 data rows, 8506 unique questions, and corresponding SQL queries.
trying
A real-world dataset for multi-image super-resolution that matches low-resolution Sentinel-2 images with high-resolution WorldView-2 images. The dataset has been introduced with the paper: - Kowaleczko, P., Tarasiewicz, T., Ziaja, M., Kostrzewa, D., Nalepa, J., Rokita, P., & Kawulok, M. (2023). A real-world benchmark for Sentinel-2 multi-image super-resolution. Scientific Data, 10(1), 644.