19,997 machine learning datasets
19,997 dataset results
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
The dataset contains aerial images containing three commonly occurring natural disasters earthquake/collapsed buildings, flood, wildfire/fire, and a normal class; do not reflect any disaster. It consist of 167723 aerial images divided into 4 classes. The dataset is an extension of the AIDER dataset (Aerial Image Dataset for Emergency Response Applications).
Wood plate bark removal processing is critical for ensuring the quality of wood processing and its products. To address the issue of lack of datasets available for the application of deep learning methods to this field, and to fill the research gap of deep learning methods in the application field of wood plate bark removal equipment, a benchmark for wood plate segmentation in bark removal processing is proposed in this study.
The Russian Financial Statements Database (RFSD) The Russian Financial Statements Database (RFSD) is an open, harmonized collection of annual unconsolidated financial statements of the universe of Russian firms.
Source: Linking Datasets on Organizations Using Half-a-Billion Open-Collaborated Records (Description (Markdown and LATEX enabled))
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
This is a dataset for 3-way sentiment classification of reviews (negative, neutral, positive). It is a merge of Stanford Sentiment Treebank (SST-3) and DynaSent Rounds 1 and 2, licensed under Apache 2.0 and Creative Commons Attribution 4.0 respectively. The SST-3, DynaSent R1, and DynaSent R2 datasets were randomly mixed to form a new dataset with 102,097 Train examples, 5,421 Validation examples, and 6,530 Test examples. See Table 1 for the distribution of labels within this merged dataset.
Engagement with the government of Taiwan as part of the vTaiwan participatory process which led to the successful regulation of Uber in Taiwan.
Dataset Description This dataset contains rental property listings scraped from Tonaton.com, one of Ghana's leading online classifieds platforms. It provides valuable information on rental prices across various regions in Ghana, along with other property details. The dataset is designed to support analysis, visualization, and modeling of rental prices in the Ghanaian real estate market.
A well-labeled challenging dataset, to facilitate the research on style recognition on anime images by collecting images from 190 anime and cartoon works covering 93 years from 13 countries and regions, 2D and 3D work into consideration concurrently. We choose at most ten roles for each work. All the images are obtained from the Internet. The images in the LSASRD dataset are mainly from existing anime and cartoons. Moreover, some are from comics or games of the same anime series. Unlike illustration or video datasets, we provide a moderate amount of contextual information in a wide variety of styles. LSASRD requires the ability of context understanding of image models.
CityTopia is the largest synthetic dataset for 3D cities with annotations, offering high-fidelity scenes generated using 3D assets from the Unreal Engine 5 CitySample project.
This dataset contains data scraped from search results of the query #israel and #palestine during the early months of the 2023 crisis.
This dataset comprises 77,175 Reddit posts from 115 subreddit forums, annotated for the presence of 15 topics related to eating disorders and dieting. The dataset includes labels and scores on all 77,175 Reddit posts, determined by 5 Large Language Models: GPT-4o, Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct, Mistral-7B-Instruct-v0.3, Vicuna-7b-v1.5, as well as by the ensemble of the four open-source LLMs. The dataset also includes a subset of 1,080 human-annotated posts for evaluation.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
The TimberVision dataset consists of more than 2k annotated RGB images and contains a total of 51k trunk components including cut and lateral surfaces, thereby surpassing any existing dataset in this domain in terms of both quantity and detail by a large margin. The dataset can be used to train oriented object detection and instance segmentation and evaluate the influence of multiple scene parameters on model performance. Additionally, a generic framework is provided to fuse the components detected by the models for both tasks into unified trunk representations. Furthermore, geometric properties are derived automatically and multi-object tracking is applied to further enhance robustness.
This dataset is the translation of the MS-marco dataset, marking it the first large-scale urdu IR dataset.
Perturbed version of the data used in Van Dijcke, Gunsilius, and Wright (2024).
Perturbed version of the data used in Van Dijcke, Gunsilius, and Wright (2024).