19,997 machine learning datasets
19,997 dataset results
On Sunday, August 29, 2021, Hurricane Ida struck parts of Louisiana and Mississippi with wind gusts reaching up to 172 mph, leaving more than a million customers without electricity, including the entire New Orleans area. During the disaster, Maxar captured high spatial resolution satellite imagery (at 0.4 m/pixel) and was subsequently made publicly available. The original images were segmented into 512*512-pixel patches to maintain spatial context while enabling detailed analysis. From this process, we generated a dataset of 2,135 triplets, each containing a pre-disaster image, a post-disaster image, and a manually annotated damage categorical mask.
first everyday task dataset featuring COT outputs, diverse task designs, detailed re-plan processes, along with SFT and DPO sub-datasets.
This repository presents the dataset used in the PerfCam's original paper. This dataset is to support further research in the area of industrial 3D reconstruction, digital twinning, and predictive maintenance.
This dataset is part of the Finetune-RAG project, which aims to tackle hallucination in retrieval-augmented LLMs. It consists of synthetically curated and processed RAG documents that can be utilised for LLM fine-tuning.
Contains ~60000 HD images of Deformable Linear Objects (DLOs) generated using blender. The dataset contains a variety of industrial-looking backgrounds and contains instance segmentation masks. The main task for this dataset is DLO instance segmentation.
measure the toxicity generated by language models across input severity and harm categories, by creating a new benchmark of open ended prefixes. We sampled 10,376 snippets from web pages across the dimensions and harms as described in the paper - https://arxiv.org/pdf/2505.02009.
pymatgen_code_qa benchmark: qa_benchmark/generated_qa/generation_results_code.json, which consists of 34,621 QA pairs. pymatgen_code_doc benchmark: qa_benchmark/generated_qa/generation_results_doc.json, which consists of 34,604 QA pairs. real-world tool-usage benchmark: src/question_segments, which consists of 49 questions (138 tasks). One subfolder means a question with problem statement, property list and verification code.
SOMPT22 is a multi-object tracking (MOT) benchmark focused on surveillance-style pedestrian tracking.
Genomics Adversarial Attack Sample dataset
The first large-scale dataset for training and evaluating novel-view synthesis from blurred images.
A real-world low-light camera motion blur dataset for evaluating deblurring radiance fields methods.
🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages.
🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages.
🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages.
With the rise of social media, user-generated content has surged, and hate speech has proliferated. Hate speech targets groups or individuals based on race, religion, gender, region, sexual orientation, or physical traits, expressing malice or inciting harm. Recognized as a growing social issue, it affects 9.41 billion Mandarin Chinese speakers (12% of the global population). However, research on Chinese hate speech detection lags, facing two key challenges.
Bc8BioRED is built upon BioRED 2022 with the addition of directionality annotations. The training and development sets from the original 2022 BioRED corpus were combined and reused as the training set, while the test set was used as the development set. Furthermore, the 400 test abstracts from the BioCreative VIII were utilized for evaluation. Bc8BioRED encompasses seven types of entities and eight types of relationships. Each relationship annotation in the Bc8BioRED corpus is categorized by novelty to indicate whether the relationship represents a significant finding or previously known background knowledge. Initially, the BioRED 2022 corpus comprised 600 abstracts for RE system development, with an additional 400 abstracts annotated to enhance coverage of emerging topics. The dataset encompasses directionality annotations (subject/object roles) for each relation pair, resulting in 10,864 directionality annotations.
Ultra-lightweight, multilingual QA eval dataset for rapid testing LLMs.
Verireason-RTL-Coder_7b_reasoning_tb For implementation details, visit our GitHub repository: VeriReason
Verireason-RTL-Coder_7b_reasoning_tb_simple For implementation details, visit our GitHub repository: VeriReason
Node-level tasks.