19,997 machine learning datasets
19,997 dataset results
Multilingual explainable fact-checking dataset on Russia-Ukraine Conflict 2022
The uniD dataset is an innovative collection of naturalistic road user trajectories, captured within the RWTH Aachen University campus using drone technology to address common challenges such as occlusions found in traditional traffic data collection methods. It meticulously documents the movement and classifies each road user by type. Employing cutting-edge computer vision algorithms, the dataset ensures high positional accuracy. Its utility spans various applications, from predicting road user behavior and modeling driver actions to conducting scenario-based safety checks for automated driving systems and facilitating the data-driven design of Highly Automated Driving (HAD) system components.
We present a multilingual test set for conducting speech intelligibility tests in the form of diagnostic rhyme tests. The materials currently contain audio recordings in 5 languages and further extensions are in progress. For Mandarin Chinese, we provide recordings for a consonant contrast test as well as a tonal contrast test. Further information on the audio data, test procedure and software to set up a full survey which can be deployed on crowdsourcing platforms is provided in our paper [arXiv preprint] and GitHub repository. We welcome contributions to this open-source project.
Raw negotiation transcripts generated for the paper "Evaluating Language Model Agency through Negotiations". The data includes transcripts from self-play (a model plays against an independent version of itself; corresponding to Section 4.1 of the paper) and cross-play (a model plays against another model; Section 4.2). This dataset encompasses 2926 transcripts (942 self-play, 1984 cross-play).
Abstract Forecasting methods from averaging regression analysis lines, reversed and direct lines. With both lines intercepting laid the turning points in direction of original line. The average regression method.
SemEval-2016 Task 6, titled "Stance Detection in Tweets," provides a specialized dataset for the computational linguistics and natural language processing (NLP) communities to explore and analyze users' positions towards certain targets, based solely on the content of their tweets. Stance detection aims to determine whether the author of a piece of text is in favor of, against, or neutral towards a specified target, such as a political figure, policy, or product.
The PRONTO heterogeneous benchmark dataset is based on an industrial-scale multiphase flow facility. It includes data from heterogeneous sources, including process measurements, alarm records, high frequency ultrasonic flow and pressure measurements, an operation log and video recordings. The study collected data from various operational conditions with and without induced faults to generate a multi-rate, multi-modal dataset. The dataset is suitable for developing and validating algorithms for fault detection and diagnosis (FDD) and data fusion.
Recorded with a Husky A200 wheeled UGV, BorealTC contains 116 min of Inertial Measurement Unit (IMU), motor current, and wheel odometry data, focusing on typical boreal forest terrains, notably snow, ice, and silty loam. The dataset also includes experiments on asphalt and flooring. All runs were recorded in Forêt Montmorency and on the main campus of Université Laval, Quebec City, Québec, Canada
Recorded with a Husky A200 wheeled UGV, the Vulpi 2021 dataset contains 13 min of Inertial Measurement Unit (IMU), motor current, and wheel odometry data, focusing on agricultural terrains. The dataset includes experiments on concrete, a dirt road, a ploughed terrain and an unploughed terrain that were all recorded on an experimental farm in San Cassiano, Lecce, Italy.
SatBird is a dataset and benchmark for the task of predicting bird species encounter rates jointly at a specific location using remote sensing data. The dataset was obtained from publicly available eBird bird observation records, Sentinel-2 satellite data, and WorldClim and SoilGrids environmental data. SatBird is composed of 3 sub-datasets: (i) USA summer dataset, generally corresponding to the breeding season, (ii) USA winter dataset, the nonbreeding season, (iii) Kenya dataset, as an example of a low-data regime. 670 and 1054 bird species are considered for the USA and Kenya respectively.
The dataset contains two few-shot chemical fine-grained entity extraction datasets, based on human-annotated ChemNER+ and CHEMET. For each dataset, we randomly sample a subset based on the frequency of each type class. Specifically, given a dataset, we first set the number of maximum entity mentions $k$ for the most frequent entity type in the dataset. We then randomly sample other types and ensure that the distribution of each type remains the same as in the original dataset. We choose the values $6, 9, 12, 15, 18$ as the potential maximum entity mentions for $k$. The ChemNER+ and CHEMET few-shot datasets contain 52 and 28 types respectively.
The Food Recall Incidents dataset consists of 7,546 short texts (from 5 to 360 characters each), which are the titles of food recall announcements (therefore referred to as title), crawled from 24 public food safety authority websites by Agroknow. The texts are written in 6 languages, with English (6,644) and German (888) being the most common, followed by French (8), Greek (4), Italian (1) and Danish (1). Most of the texts have been authored after 2010 and they describe recalls of specific food products due to specific hazards. Experts manually classified each text to four groups of classes describing hazards and products on two levels of granularity:
Dataset Overview vanilla.csv: Represents the interactions without specific role-play instructions. boss.csv: Interactions where ChatGPT plays the role of a user's boss. classmate.csv: Interactions with ChatGPT acting as the user's classmate. Each turn was coded with user motives of user responses, or the perceived naturalness of ChatGPT responses.
GEM (A General Evaluation Benchmark on Multi-modal Tasks) is a significant benchmark dataset designed to evaluate the performance of cross-modal pre-trained models, including both understanding and generation tasks. Unlike existing datasets such as GLUE, SuperGLUE, XGLUE, and XTREME, which primarily focus on natural language tasks, GEM stands out as a large-scale vision-language benchmark.
The USPTO Backgrounds dataset provides valuable information related to patents and trademarks. Here are some key datasets available from the United States Patent and Trademark Office (USPTO):
MatSeg Dataset for Zero-Shot Material States Segmentation: The dataset contains large-scale synthetic images for training data and highly diverse real-world image benchmarks for testing. Focusing on zero-shot class-agnostic segmentation of materials and their states. This means finding the region of materials states without pre-training on the specific material classes or states. The benchmark contains a wide range of real-world materials and states. For example: wet regions of the surface, scattered dust, minerals of rocks, the sediment of soils, rotten parts of fruits, degraded and corrosive surface regions, food and liquid states, and many others. The focus is on scattered and fragmented materials, as well as soft boundaries partial transition, and partial similarity between regions. It contains both hard segmentation maps and soft and partial similarity annotations for similar but not identical materials.
The dataset was collected from DOTA 2 using OpenDota API via Python. The collection consists of DOTA 2 in-game chat data that were manually categorized into 3 classifications: non-toxic, mild (toxicity), and toxic chats.
This dataset accompanies the paper `Learning the mechanisms of network growth' by the same authors. The dataset contains 6733 networks of size 20,000 each generated in accordance to different combination of three mechanisms: fitness, aging and preferential attachment. The goal is to use machine learning to identify the combination of mechanisms that was used to create the network. The dataset includes static features from the literature and two version of our newly developed dynamic features. net
ConQA is a dataset created using the intersection between VisualGenome and MS-COCO. The goal of this dataset is to provide a new benchmark for text to image retrieval using short and less descriptive queries than the commonly use captions from MS-COCO or Flicker. ConQA consists of 80 queries divided into 50 conceptual and 30 descriptive queries. A descriptive query mentions some of the objects in the image, for instance, people chopping vegetables. While, a conceptual query does not mention objects or only refers to objects in a general context, e.g., working class life.
This dataset of hand drawn images of molecular depictions incorporates 4 sub dataset: