19,997 machine learning datasets
19,997 dataset results
For our experiments, we collected a dataset of procedural knowledge of the LangChain Python library, unseen by many extant LLMs. We selected LangChain as the domain for our dataset because it was published in 2022, which is later than the knowledge cutoff date for many web-scale LLMs, including GPT-3.5, while also having plenty of documentation due to its popularity.
The Concept and Hint-Annotated Math Problems (CHAMP) consists of high school math competition problems, annotated with concepts, or general math facts, and hints, or problem-specific tricks. These annotations allow us to explore the effects of additional information, such as relevant hints, misleading concepts, or related problems.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Annotating data is a time-consuming and costly task, but it is inherently required for supervised machine learning. Active Learning (AL) is an established method that minimizes human labeling effort by iteratively selecting the most informative unlabeled samples for expert annotation, thereby improving the overall classification performance. Even though AL has been known for decades [1], AL is still rarely used in real-world applications. As indicated in the two community web surveys among the NLP community about AL [2], [3], two main reasons continue to hold practitioners back from using AL: first, the complexity of setting AL up, and second, a lack of trust in its effectiveness. We hypothesize that both reasons share the same culprit: the large hyperparameter space of AL. This mostly unexplored hyperparameter space often leads to misleading and irreproducible AL experiment results. In this study, we first compiled a large hyperparameter grid of over 4.6 million hyperparameter combina
A benchmark environment based on the datasets "Adult" and "Names", which allows researchers to test how well their language model can abide by pre-defined access rights rules. Researchers can either directly use the datasets we generated for our ACL 2025 Findings paper or generate their own custom dataset.
An evaluation test bed for assessing the robustness of sentence embedding models against user-informed misinformation edits. A dataset containing perturbed and unperturbed claim pairs used for improving embedding model robustness through knowledge distillation.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
The POLARIS dataset is built from a decade of polarimetric observations (2014–2024) conducted with the SPHERE instrument on the Very Large Telescope (VLT). Specifically, it includes all public polarized light observations obtained using the IRDIS instrument, retrieved from the ESO Science Archive. These raw observations were uniformly preprocessed using a modified version of the IRDAP pipeline to generate high-quality Polarimetric Differential Imaging (PDI) products.
DnR-nonverbal is a dataset for cinematic audio source separation (CASS) based on Divide and Remaster (DnR) dataset.
Aria Scenes is a benchmark dataset for future research on photorealistic reconstruction. The dataset includes 12 .vrs files created in diverse indoor and outdoor environments.
This dataset is a collection of paired wireless signal data and corresponding image ground truth specifically designed for underground object sensing and image reconstruction. It utilizes Channel State Information (CSI) and Received Signal Strength Indicator (RSSI) collected from a low-cost WiFi Wireless Sensor Network (WSN) deployed around a soil volume containing underground potato tubers. The data was acquired in a controlled laboratory environment using an automated system that ensures precise alignment between the wireless measurements and the physical location/shape of the target.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
To evaluate our proposed strategy of asynchronous communication for LLMs, we run games of Mafia with human players, incorporating an LLM-based agent as an additional player, within an asynchronous chat environment.
Annotation Guidlines Three speech and language pathologists, with experience ranging from 2 to 40 years, independently annotated and analyzed the audiovisual samples sourced from the Fluencybank Adults Who Stutter(AWS) dataset. The annotations were created using the ELAN tool.
We introduce ArVoice, a multi-speaker Modern Standard Arabic (MSA) speech corpus with diacritized transcriptions, intended for multi-speaker speech synthesis, and can be useful for other tasks such as speech-based diacritic restoration, voice conversion, and deepfake detection.
the dataset is a monkey doo doo dataset
This is a link to the source code of the Baking-Large domain introduced in the paper.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
WMT 2024 is a collection of datasets used in shared tasks of the Ninth Conference on Machine Translation. The conference builds on a series of annual workshops and conferences on Statistical Machine Translation.