19,997 machine learning datasets
19,997 dataset results
SciCo is an expert-annotated dataset for hierarchical CDCR (cross-document coreference resolution) for concepts in scientific papers, with the goal of jointly inferring coreference clusters and hierarchy between them.
The model forecasts for the sub-seasonal forecasting application considered in the Online Learning under Optimism and Delay paper experiments. This dataset consists of a single ZIP archive (919MB) that contains 1) a "models" folder that contains, for each model the forecasts for the Precip. 3-4w, Precip. 5-6w, Temp. 3-4w, Temp. 5-6w tasks on the western United States geography, and 2) a "data" folder that contains supporting geographic data. The data should be used to reproduce the PoolD experiments in https://github.com/geflaspohler/poold as described in the README. (2021-06-10)
This data is for the Mis2-KDD 2021 under review paper: Dataset of Propaganda Techniques of the State-Sponsored Information Operation of the People’s Republic of China
This is a set of signals-pairs, univariate and multivariate, that can be used to test alignment algorithms. Signals are morphologically different.
Since robust foreground/background separation and segmentation of cellular objects (i.e.,identification of which pixels below to which objects) strongly depends on image quality, focus artifacts are detrimental to data quality. This image set provides examples of in- and out-of-focus synthetic images, which can be used for validation of focus metrics.
A large-scale training dataset suffering from the defocus spread effect (DSE) is synthesized by applying an $\alpha$-matte boundary defocus model to the VOC 2012 dataset.
We release both the processed data and evaluation results from our own experiments, and the underlying raw data that can be used for future experiments and schemes in the domain of Zero-Interaction Security. Find more details in the dataset description on Zenodo.
The SurfaceGrid dataset contains nearly a million 512x512 images for use in training neural networks on shape-fron-surface contour task.
Chinese Medical Information Extraction, a dataset that is also released in CHIP2020, is used for CMeIE task. The task is aimed at identifying both entities and relations in a sentence following the schema constraints. There are 53 relations defined in the dataset, including 10 synonymous sub-relationships and 43 other sub-relationships.
CHIP Clinical Trial Classification, a dataset aimed at classifying clinical trials eligibility criteria, which are fundamental guidelines of clinical trials defined to identify whether a subject meets a clinical trial or not, is used for the CHIP-CTC task. All text data are collected from the website of the Chinese Clinical Trial Registry (ChiCTR) , and a total of 44 categories are defined. The task is like text classification; although it is not a new task, studies and corpus for the Chinese clinical trial criterion are still limited, and we hope to promote future researches for social benefits.
The Oxford Road Boundaries is a dataset designed for training and testing machine-learning-based road-boundary detection and inference approaches.
Probing cross-modal capabilities of Vision & Language models with a counting task.
Survey instrument, analysis code, and anonymized responses for the paper on review practices in SE.
This case surveillance public use dataset has 12 elements for all COVID-19 cases shared with CDC and includes demographics, any exposure history, disease severity indicators and outcomes, presence of any underlying medical conditions and risk behaviors, and no geographic data.
The PEDC is a corpus of 14 episodes of This American Life podcast transcripts that have been annotated for events. The corpus contains excerpts from these episodes (listed in Tabe 1) that are dialogue. The granularity of annotation in this corpus is the token; each token is either annotated as an event, or a nonevent. For more information please download the corpus, and see the annotation guide for more specifics on how we define event, and the README for how the annotations are encoded. Also, much more information regarding the corpus, and its use is in the Automatic extraction of personal events from dialogue paper.
Calliar is a dataset for Arabic calligraphy. The dataset consists of 2500 json files that contain strokes manually annotated for Arabic calligraphy.
HICRD (Heron Island Coral Reef Dataset) is a large-scale real underwater image dataset for underwater image restoration. There are 2000 reference restored images and 6003 original underwater images in the unpaired training set.
Hi-Phy is a benchmark for physical reasoning that allows researchers to test individual physical reasoning capabilities. Inspired by how humans acquire these capabilities, the benchmark proposes a general hierarchy of physical reasoning capabilities with increasing complexity. this benchmark tests capabilities according to this hierarchy through generated physical reasoning tasks in the video game Angry Birds.
Indian Masked faces in the wild Database is collected into three sets:(i) Indian Celebrity, (ii) Instagram and (iii) Indian Crowd. The Indian Celebrity contains 40 Indian celebrities with 435 images, including Bollywood actors/actresses, television stars, sports personalities, and politicians. The Instagram set contains 377 images of 40 subjects downloaded from Instagram. We collected masked and non-masked images of Indian people with a public profile. The Indian Crowd set is collected from the common people who volunteered to contribute to the dataset. This set contains 120 subjects with 562 images. All the Images are collected in both constrained and unconstrained environments with variation in pose, illumination, background and masks worn by the people.