19,997 machine learning datasets
19,997 dataset results
A Chinese sign language dataset that includes dialogue information.
We introduce HSIRS, a large scale dataset of hyper-spectral images along with corresponding manually annotated segmentation maps for material characterization and classification based on spectral signature. Such data can be used to simulate any type of spectrometer and to train DNNs end-to-end for spectral reconstruction and image segmentation tasks. HSIRS features scenes containing real and fake (made of polyester, plastic or ceramic) food items with different backgrounds and scene layouts, some scenes contain also color checkers. Spectral bands are sequentially captured using a VariSpecTM tunable color filter and the scene is illuminated with 4 Halogen light sources.
Dataset of cross-layer Radio Access Network (RAN) Key Performance Measurements (KPMs) and protocol stack logs collected on an Open RAN deployment instantiated on Colosseum with traffic twinned from that of commercial cellular traces. The dataset includes Base Station (BS)- and User Equipment (UE)-level KPMs from PHY, MAC, and App layers under different RAN configurations representative of AI/ML control policies, number of UEs, and traffic demand. The fine-grained metrics of the dataset make it possible to understand the connection between PHY and MAC KPMs measured at the BS and UEs, control policies, and end-to-end and App-layer KPMs that reflect user experience.
Diverse guitar-playing motions about 1 hour long, including: • 12 major scales, • chromatic scales, • diverse chords, • arpeggios, • strumming and picking, • bends, • sliding, • vibrato, • palm mute, • natural harmonics, • artificial harmonics, • hammer-ons and pull-offs.
GTA-UAV dataset provides a large continuous area dataset (covering 81.3km<sup>2</sup>) for UAV visual geo-localization, expanding the previously aligned drone-satellite pairs to arbitrary drone-satellite pairs to better align with real-world application scenarios. Our dataset contains:
This Dataset contains pairs off textual natural language questions and SPARQL queries on a small subset of the CoyPu KnowledgeGraph (https://coypu.org/ergebnisse/knowledge-graph)
This Dataset contains pairs off textual natural language questions and SPARQL queries on a small organizational graph(https://github.com/AKSW/AI-Tomorrow-2023-KG-ChatGPT-Experiments/blob/main/FoafVcardOrg/foaf-vcard-org-data.ttl) which was introduced in "LLM-assisted knowledge graph engineering: Experiments with chatgpt" by L.-P. Meyer et al. 2023 (DOI 10.1007/978-3-658-43705-3_8)
Parallel version of annotations in GUM RST v9.1.
RST corpus for Russian.
Diderot’s Encyclopédie is a reference work from XVIIIth century in Europe that aimed at collecting the knowledge of its era. This repository hosts an annotated dataset of more than 10,400 of the Encyclopédie entries with Wikidata identifiers enabling us to connect these entries to the Wikidata graph. The dataset can serve to train and evaluate named entity solvers.
Orchid2024 is a fine-grained classification dataset specifically designed for Chinese Cymbidium orchid cultivars. It includes data collected from 20 cities across 12 provincial administrative regions in China and encompasses 1,269 cultivars from 8 Chinese Cymbidium orchid species and 6 additional categories, totaling 156,630 images. The dataset covers nearly all common Chinese Cymbidium cultivars currently found in China, with its fine granularity and focus on the real world making it a unique and practical resource for researchers and practitioners.
This is the official dataset collected for to test the sim-to-real transfer. It contains 6 articulated object instances, each captured from 20 camera views under 5 states in scenarios with and without background, as well as presence or absence of distractors.
A benchmark + dataset for evaluating multimodal models on business process management (BPM) tasks.
FUSU dataset covers 5 whole urban areas, 847 km^2 located in the north and south of China, with 17 land use and land cover (LULC) classes and over 170K images and 30 billion pixels of annotations, supporting segmentation, change detection and domain adaptation tasks. This data comprises 2 parts:
A small-scale, real-world Project Aria dataset with high quality static 3D oriented bounding boxs annotations.
Around 90k different RL runs: 256 hyperparameter configurations for 10 seeds each across a total of 3 algorithms (PPO, SAC, DQN) and 22 environments.
The Guided Lexrank algorithm is applied to dataset special_appeal.csv to summarize the texts of legal documents. The obtained summary and the texts of the topics contained in dataset themes.csv are submitted to the BM25 algorithm for similarity assessment. From a list of topics, the GLARE method produces a ranking with suggested topics for a given document.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
ALLO is an anomaly detection and localization dataset for space stations in lunar orbit. Synthetically rendered using Blender, ALLO provides realistic images of what a robotic manipulator on a space station will encounter including possible anomalies.
URL: https://sparse.tamu.edu/LAW