19,997 machine learning datasets
19,997 dataset results
This dataset is a large-scale dataset for moving object detection and tracking in satellite videos, which consists of 40 satellite videos captured by Jilin-1 satellite platforms. Each image has a resolution of 12000x5000 and contains a great number of objects with different scales. Four common types of vechicles, including plane, car, ship, and train, are manually-labeled. A total of 853,911 instances are labeled by axis-aligned bounding boxes. https://paperswithcode.com/paper/detecting-and-tracking-small-and-dense-moving
The measurement data <b>VNA_20220722_232002_XETS_reduced.mat</b> includes a data matrix $\mathbf{R}$ acquired with a synthetic aperture measurement testbed described in [2] and [3]. Measured were $N_f=1000$ frequency steps in a band from $3-10$GHz of the scattering parameter $S_{21}$ between a synthetic $51$-ULA with antenna positions saved in the file <b>ULA.mat</b> and a synthetic $(13\times 13)$-URA with antenna positions saved in the file <b>URA.mat</b>. The file <b>XETSantennaCharacterization.mat</b> holds antenna gains of an XETS antenna [4] characterized in an anechoic chamber. XETS antennas were used on both the ULA (oriented towards the negative $x$-direction) and on the URA (oriented towards the positive $x$-direction).<br> These data have been used to perform ultra-wideband (UWB) bistatic radar imaging and wireless power transfer (WPT). Our implementation is provided in <b>MAIN_wall_detection.mat</b>.
Wikipedia Webpage 2M (WikiWeb2M) is a multimodal open source dataset consisting of over 2 million English Wikipedia articles. It is created by rescraping the ∼2M English articles in WIT. Each webpage sample includes the page URL and title, section titles, text, and indices, images and their captions.
Despite recent advances in vision-and-language tasks, most progress is still focused on resource-rich languages such as English. Furthermore, widespread vision-and-language datasets directly adopt images representative of American or European cultures resulting in bias. Hence we introduce ParsVQA-Caps, the first benchmark in Persian for Visual Question Answering and Image Captioning tasks. We utilize two ways to collect datasets for each task, human-based and template-based for VQA and human-based and web-based for image captioning. The image captioning dataset consists of over 7.5k images and about 9k captions. The VQA dataset consists of almost 11k images and 28.5k question and answer pairs with short and long answers usable for both classification and generation VQA.
Meta Omnium is a dataset-of-datasets spanning multiple vision tasks including recognition, keypoint localization, semantic segmentation and regression. Meta Omnium enables meta-learning researchers to evaluate model generalization to a much wider array of tasks than previously possible, and provides a single framework for evaluating meta-learners across a wide suite of vision applications in a consistent manner.
A View From Somewhere (AVFS)—a dataset of 638,180 face similarity judgments over 4,921 faces. Each judgment corresponds to the odd-one-out (i.e., least similar) face in a triplet of faces and is accompanied by both the identifier and demographic attributes of the annotator who made the judgment.
titanic5 Dataset Created by David Beltran del Rio March 2016.
This corpus contains preprocessed posts from the Reddit dataset, suitable for abstractive summarization using deep learning. The format is a json file where each line is a JSON object representing a post. The schema of each post is shown below: - author: string (nullable = true) - body: string (nullable = true) - normalizedBody: string (nullable = true) - content: string (nullable = true) - content_len: long (nullable = true) - summary: string (nullable = true) - summary_len: long (nullable = true) - id: string (nullable = true) - subreddit: string (nullable = true) - subreddit_id: string (nullable = true) - title: string (nullable = true)
COVIDx CXR-3 is an open access benchmark dataset that we generated, comprising 30,882 CXR images across 17,026 patient cases. Images may be added over time to improve the dataset.
The increasing use of deep learning techniques has reduced interpretation time and, ideally, reduced interpreter bias by automatically deriving geological maps from digital outcrop models. However, accurate validation of these automated mapping approaches is a significant challenge due to the subjective nature of geological mapping and the difficulty in collecting quantitative validation data. Additionally, many state-of-the-art deep learning methods are limited to 2D image data, which is insufficient for 3D digital outcrops, such as hyperclouds. To address these challenges, we present Tinto, a multi-sensor benchmark digital outcrop dataset designed to facilitate the development and validation of deep learning approaches for geological mapping, especially for non-structured 3D data like point clouds. Tinto comprises two complementary sets: 1) a real digital outcrop model from Corta Atalaya (Spain), with spectral attributes and ground-truth data, and 2) a synthetic twin that uses latent
HAC is a dataset for learning and benchmarking arbitrary Hybrid Adverse Conditions restoration. HAC contains 31 scenarios composed of an arbitrary combination of five common weather, with a total of 316K adverse-weather/clean pairs.
Dataset can be used by anyone who is interested to perform morphological classification of galaxies. Originally dataset provided by Kaggle user Jay Lin (https://www.kaggle.com/jay1985) 4 years ago. Dataset was used in conference paper "Morphological Classification of Galaxies Using SpinalNet"
Smart Word Suggestions (SWS) is a task and benchmark. This task involves identifying words or phrases that require improvement and providing substitution suggestions. The benchmark includes human-labeled data for testing, a large distantly supervised dataset for training, and the framework for evaluation. The test data includes 1,000 sentences written by English learners, accompanied by over 16,000 substitution suggestions annotated by 10 native speakers. The training dataset comprises over 3.7 million sentences and 12.7million suggestions generated through rules.
Multi-CrossRE is a broadest multi-lingual dataset for Relation Extraction (RE) including 26 languages in addition to English, and covering six text domains. It is a machine translated version of CrossRE crossre, with a sub-portion including more than 200 sentences in seven diverse languages checked by native speakers.
FLORIS farm dataset A dataset for graph neural network modeling of wind farms. The current version of the dataset contains two farms, with very different geometry but similar inter-turbine statistics. The wind farms were simulated with the steady-state wake model FLORIS.
The directory HiCIS contains two datasets for instance segmentation of honeycombs in concrete in COCO Format. The datasets orginate from images scraped from the internet and the other one is provided by Metis Systems AG. The directory HiCC/web contains the dataset using the images from the internet and HICC/metis contains the dataset using the images provided by Metis Systems AG as part of the research project Smart Design and Construction (SDaC).
naab: A ready-to-use plug-and-play corpus for Farsi The biggest cleaned and ready-to-use open-source textual corpus in Farsi. It contains about 130GB of data, 250 million paragraphs, and 15 billion words. The project name is derived from the Farsi word NAAB K which means pure and high grade. We also provide the raw version of the corpus called naab-raw and an easy-to-use preprocessor that can be employed by those who wanted to make a customized corpus.
TFVulFix is a dataset containing commits from TensorFlow, which is a well-known deep learning library. It contains 290 vulnerability fixing and 1,535 non-vulnerability-fixing commits. In this dataset, no commit is explicitly linked to an issue.
I'm writing a description for the SMIC dataset on Papers With Code. SMIC is a dataset for facial expression recognition (FER).
Taking Advice from ChatGPT is a laboratory study of how student participants incorporate advice generated by ChatGPT. In a survey conducted through the Experimental Social Science Laboratory, 118 students answered 2,828 questions on topics from the MMLU benchmark. The rich dataset includes questions/choices, advice characteristics, participant answers, and participant background. It can be used to explore algorithm aversion, advice-taking, ChatGPT usage, and more.