19,997 machine learning datasets
19,997 dataset results
This collection consists of DICOM images and DICOM Segmentation Objects (DSOs) for 197 patients with Colorectal Liver Metastases (CRLM). The collection consists of a large, single-institution consecutive series of patients that underwent resection of CRLM and matched preoperative computed tomography (CT) scans for quantitative image analysis. Inclusion criteria were (a) pathologically confirmed resected CRLM, (b) available data from pathologic analysis of the underlying non-tumoral liver parenchyma and hepatic tumor, (c) available preoperative conventional portal venous contrast-enhanced multi-detector computed tomography (MDCT) performed within 6 weeks of hepatic resection. Patients with 90-day mortality or that had less than 24 months of follow-up were excluded. Additionally, because pathologic and radiographic alterations of the non-tumoral liver parenchyma caused by hepatic artery infusion (HAI) of chemotherapy are not well described, any patient who received preoperative HAI was e
WTA (Wind Turbine Aerial) and TLA (Transmission Line Aerial) are public datasets which contain a set of RGB images from wind turbine farms and transmission towers and power lines, along with semantic ground truth for relevant classes. This is the official repository of the paper: WTA/TLA: A UAV-captured Dataset for Semantic Segmentation of Energy Infrastructure (url).
The datasets on this page are designed for machine learning-based Network Intrusion Detection Systems (NIDS) and are organised into the following high-level collections:
This dataset presents a novel, multi-variate time series specifically designed for advancing research in spatio-temporal forecasting. The primary goal of this dataset is to facilitate the accurate prediction of traffic throughput volumes across 5G communication networks.
<a href="https://gts.ai/dataset-download/urban-visual-pollution-dataset/" target="_blank">👉 Download the dataset here</a>
Description:
Applications of unmanned aerial vehicle (UAV) in logistics, agricultural automation, urban management, and emergency response are highly dependent on oriented object detection (OOD) to enhance visual perception. Although existing datasets for OOD in UAV provide valuable resources, they are often designed for specific downstream tasks. Consequently, they exhibit limited generalization performance in real flight scenarios and fail to thoroughly demonstrate algorithm effectiveness in practical environments. To bridge this critical gap, we introduce CODrone, a comprehensive oriented object detection dataset for UAVs that accurately reflects real-world conditions. It also serves as a new benchmark designed to align with downstream task requirements, ensuring greater applicability and robustness in UAV-based OOD. Based on application requirements, we identify four key limitations in current UAV OOD datasets-low image resolution, limited object categories, single-view imaging, and restricte
test
The Storytelling Video Dataset is a high-quality, human-reviewed multimodal dataset featuring over 700 full-body video recordings of native Russian speakers. Each video is 10+ minutes long and includes synchronized speech, facial expressions, gestures, and emotional variation. The dataset is ideal for research and development in:
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Artificial Relationships in Fiction Dataset Description Artificial Relationships in Fiction (ARF) is a synthetically annotated dataset for Relation Extraction (RE) in fiction, created from a curated selection of literary texts sourced from Project Gutenberg. The dataset captures the rich, implicit relationships within fictional narratives using a novel ontology and GPT-4o for annotation. ARF is the first large-scale RE resource designed specifically for literary texts, advancing both NLP model training and computational literary analysis.
gggg
Description:
Description:
Data in this study come from western Ecuador's Choco tropical forest, including \textit{Fundación para la Conservación de los Andes Tropicales Reserve and adjacent Reserva Ecológica Mache-Chindul park} (FCAT; 00$^\circ$23'28'' N, 79$^\circ$41'05'' W), \textit{Jama-Coaque Ecological Reserve} (00$^\circ$06'57'' S, 80$^\circ$07'29'' W), \textit{Canande Reserve} (0$^\circ$31'34'' N 79$^\circ$12'47'' W), and \textit{Tesoro Escondido Reserve} (0$^\circ$33'16'' N 79$^\circ$10'31'' W). FCAT is a high diversity humid tropical forest at elevation $\sim$500m, receiving $\sim$3000 mm yr$^{-1}$ precipitation with persistent fog during drier period. Jama-Coaque ranges from the boundary of the tropical moist deciduous/tropical moist evergreen forest at the lower elevations ($\sim$1000 mm precipitation yr$^{-1}$, $\sim$250 m asl) to fog-inundated wet evergreen forests above 580m to 800m. Canande (350–500 m elevation) and Tesoro Escondido ($\sim$200 m elevation) are lowland everwet Choco forests, both
This dataset consists of images of various hand and power tools, specifically designed to aid in training and improving AI-based object recognition systems.
<h1>Huawei University Challenge Competition 2021</h1>
Description:
TURSpider is a novel Turkish Text-to-SQL dataset that includes complex queries, akin to those in the original Spider dataset. TURSpider dataset comprises two main subsets: a dev set and a training set, aligned with the structure and scale of the popular Spider dataset. The dev set contains 1034 data rows with 1023 unique questions and 584 distinct SQL queries. In the training set, there are 8659 data rows, 8506 unique questions, and corresponding SQL queries.
Description: