19,997 machine learning datasets
19,997 dataset results
A labeled dataset that presents fake news surrounding the conflict in Syria. The dataset consists of a set of articles/news labeled by 0 (fake) or 1 (credible). Credibility of articles are computed with respect to a ground truth information obtained from the Syrian Violations Documentation Center (VDC). In particular, for each article, we crowdsource the information extraction (e.g., date, location, Number of casualties) job using the crowdsourcing platform Figure Eight (formally CrowdFlower). Then, we match those articles against the VDC database to be able to deduce whether an article is fake or not. The dataset can be used to train machine learning models to detect fake news.
A total of 80 real material samples were captured in a dark room. For each material, multiple captures were collected at different distances from the camera (between 250 and 650 mm) to observe both macro- and micro-level details. The dataset is mostly comprised of planar specimens but also includes non-planar objects such as mugs, globes, crumpled paper, etc. As shown above, it contains a rich diversity of materials, including diffuse or specular wrapping papers, fabrics, anisotropic metals, plastics, rugs, ceramic and wood flooring samples, etc. Each capture set includes 12 LDR (8 bpp) RGB-D images at 4K pixel resolution. Each set is captured at 50% and 100% of maximum light intensity. In total, we captured 462 such image sets (combinations of light intensities, distances to the camera, and material sample).
Real and simulated lidar data of indoor and outdoor scenes, before and after geometric scene changes have occurred. Data include lidar scans from multiple viewpoints with provided coordinate transforms, and manually annotated ground-truth regarding which parts of the scene have changed between subsequent scans.
ACFR Orchard Fruit Dataset is an agricultural dataset containing images and annotations for different fruits, collected at different farms across Australia. The dataset was gathered by the agriculture team at the Australian Centre for Field Robotics, The University of Sydney, Australia.
The University of Washington/Northwestern University (UW/NU) Corpus contains recordings and textgrids of Pacific Northwest and Northern Cities speakers reading a subset of the IEEE "Harvard" sentences. The UW/NU Corpus Version 1.0 has been used to study the effects of dialectal variation on speech intelligibility, while version 2.0 is being used in ongoing research in speech intelligibility and gender interaction. Development is supported by the National Institutes of Health, National Institute on Deafness and Other Communication Disorders grant R01-DC006014. The PN/NC Corpus is well suited for both clinical and research studies where high-fidelity recordings and regional accent control are desirable.
Source: Cell fate inclination within 2-cell and 4-cell mouse embryos revealed by single-cell RNA sequencing
Source: Heterogeneity in Oct4 and Sox2 Targets Biases Cell Fate in 4-Cell Mouse Embryos
Source: Single-cell RNA-Seq profiling of human preimplantation embryos and embryonic stem cells
Source: Single-cell RNA-seq reveals dynamic, random monoallelic gene expression in mammalian cells
TPM values together with cell type annotations that were obtained from Alex Pollen on 15/10/15
Source: Reconstructing lineage hierarchies of the distal lung epithelium using single-cell RNA-seq
Using the Experience-Sampling Method (ESM), participants are asked to report TV consumption multiple times each day for a five week period. Through self-reported data, authors decrease uncertainty of exposure to content, and allow collection of non-trivial information, such as how much attention is paid to the TV. The data is structured to accommodate quantitative analyses, e.g. in the CARS community, and is publicly available under the name Contextual TV (CTV) dataset.
Dataset based on Twitter usernames of American politicians. Data extracted from Wikidata.
A medium-scale synthetic 4D Light Field video dataset for depth (disparity) estimation. From the open-source movie Sintel. The dataset consists of 24 synthetic 4D LFVs with 1,204x436 pixels, 9x9 views, and 20–50 frames, and has ground-truth disparity values, so that can be used for training deep learning-based methods. Each scene was rendered with a clean pass after modifying the production file of Sintel with reference to the MPI Sintel dataset.
A dataset for flying honeybee detection introduced in "A Method for Detection of Small Moving Objects in UAV Videos".
Metric-Type of Numerical Tables is a dataset extracted from scientific papers (ACL anthology website) consisting of header tables, captions, and metric-types.
Each file contains a specific dataset described in the paper "On Automatic Parsing of Log Records". For example, T_E.txt contains the data for the dataset $T_E$.
The Ubuntu Chat Corpus (UCC) is composed of archived chat logs from Ubuntu's Internet Relay Chat technical support channels. Ubuntu uses IRC as one of many modes of technical support -- it offers real-time problem solving. The authors have taken some of the archived messages (which are in the public domain), reorganized the file structure, removed some unnecessary system messages, and compressed them to make it easier to obtain.
The Liu et al. Corpus is a pretraining dataset for large language models. It consists of 160Gb of news, books, stories, and web text.
To validate the generalization abilities of SOD models, we create a small-scale dataset by collecting the most challenging images with varying brightness and contrast, background and foreground colors overlap, among many others. We conclude that the current models, including ours, are not trust-worthy for real-world practice, demanding extensive future research for more efficient and generalized SOD models.