19,997 machine learning datasets
19,997 dataset results
A 160B bilingual long-text dataset with 3 categories: holistic, aggregated and chaotic long texts.
LADI Overview The Low Altitude Disaster Imagery (LADI) dataset was created to address the relative lack of annotated post-disaster aerial imagery in the computer vision community. Low altitude post-disaster aerial imagery from small planes and UAVs can provide high-resolution imagery to emergency management agencies to help them prioritize response efforts and perform damage assessments. In order to accelerate their workflow, computer vision can be used to automatically identify images that contain features of interest, including infrastructure such as buildings and roads, damage to such infrastructure, and hazards such as floods or debris.
The TCB-DS dataset is a specialized collection of microscopic images focusing on the automatic recognition of cyanobacteria genera. This dataset was meticulously compiled to address the challenges associated with the varying image qualities due to differences in contrast, resolution, size, lighting, and the presence of noise in the original images. It includes 2,591 images with varying dimensions, ranging from a minimum of 11 × 41 pixels to a maximum of 5184 × 3456 pixels.
A version of the WSRD Dataset will be used as a benchmark for the NTIRE24 Challenge on Image Shadow Removal.
TCEC games dataset.
While large language models (LLMs) excel in various natural language tasks in English, their performance in lower-resourced languages like Hebrew, especially for generative tasks such as abstractive summarization, remains unclear. The high morphological richness in Hebrew adds further challenges due to the ambiguity in sentence comprehension and the complexities in meaning construction. In this paper, we address this resource and evaluation gap by introducing HeSum, a novel benchmark specifically designed for abstractive text summarization in Modern Hebrew. HeSum consists of 10,000 article-summary pairs sourced from Hebrew news websites written by professionals. Linguistic analysis confirms HeSum's high abstractness and unique morphological challenges. We show that HeSum presents distinct difficulties for contemporary state-of-the-art LLMs, establishing it as a valuable testbed for generative language technology in Hebrew, and MRLs generative challenges in general.
The "Crime Facts" of "Offenses of Fraudulence" in Judicial Yuan Verdicts Dataset
This repository contains the realFormula dataset presented in the paper MathNet: A Data-Centric Approach for Printed Mathematical Expression Recognition.
We introduce a set of 425 panoramic X-rays with Human annotated Bounding Boxes and Polygons, the 425 images are a subset of UFBA-UESC Dental Dataset. This dataset can be extensively used for detection and segmentation tasks for Dental Panoramic X-rays. Refer to Description for understanding the organisation of annotations and panoramic X-rays. The dataset contains distribution of panoramic X-rays across ten different categories.
This dataset provides a comprehensive resource for detecting and evaluating bias across multiple NLP tasks.
Accompanying supplementary data for the paper. To download this data automatically and use the software, please refer to the details in the README of the linked GitHub repository.
ACCORD CSQA is an extension of the popular CommonsenseQA (CSQA) dataset using ACCORD, a scalable framework for disentangling the commonsense grounding and reasoning abilities of large language models (LLMs) through controlled, multi-hop counterfactuals. ACCORD closes the measurability gap between commonsense and formal reasoning tasks for LLMs. A detailed understanding of LLMs' commonsense reasoning abilities is severely lagging compared to our understanding of their formal reasoning abilities, since commonsense benchmarks are difficult to construct in a manner that is rigorously quantifiable. Specifically, prior commonsense reasoning benchmarks and datasets are limited to one- or two-hop reasoning or include an unknown (i.e., non-measurable) number of reasoning hops and/or distractors. Arbitrary scalability via compositional construction is also typical of formal reasoning tasks but lacking in commonsense reasoning. Finally, most prior commonsense benchmarks either are limited to a si
The Unsplash Dataset is made up of over 350,000+ contributing global photographers and data sourced from hundreds of millions of searches across a nearly unlimited number of uses and contexts. Due to the breadth of intent and semantics contained within the Unsplash dataset, it enables new opportunities for research and learning.
This dataset contains prompts designed to evaluate and challenge the safety mechanisms of generative text-to-image models, with a particular focus on identifying prompts that are likely to produce images containing nudity. Introduced in the 2024 ICML paper Prompting4Debugging: Red-Teaming Text-to-Image Diffusion Models by Finding Problematic Prompts, this dataset is not specific to any single approach or model but is intended to test various mitigating measures against inappropriate content generation in models like Stable Diffusion. This dataset is only for research purposes.
BS-Objaverse 660k Dataset is a set of GPT4-Vision-powered multi-modal captions data. It is constructed to enhance modality alignment and fine-grained visual concept perception for describing detailed information about the shape, texture of Objaverse 3D object.
Dataset composed of two main parts 1. Material characterization of a metal (from lab test)
High-definition Talking Face Dataset (HDTF). The HDTF dataset is collected from youtube website published in recent two years and consists of about 16 hours 720P∼1080P videos. There are over 300 subjects and 10k different sentences in HDTF dataset. Our HDTF dataset has higher video resolution than previous in-thewild datasets and more subjects/sentences than in-the-lab datasets.
Abstract: Through digitization, maintaining and promoting cultural heritage is being strengthened. Concerning this background, this study presents a new Indonesia cultural events dataset and automatic image classification for cultural events. The dataset was developed using the Flickr image platform, and the five cultural events image was collected including the Baliem Festival, Jember Fashion Festival, Nyepi Festival, Pacu Jawi, and Pasola Festival. Further, Convolutional Neural Networks (CNN) was developed for the classification method. A comparison of CNN models (VGG16 and VGG19) using several optimization configurations was performed to get the best model. The results showed that the VGG16 with image augmentation and dropout regularization technique performed best with 94.66% accuracy. This study hoped to support the heritage's digital documentation process and preserve Indonesia's cultural heritage.
Download free fonts in DaFont style from our extensive collection. Find bold, italic, cursive, futuristic fonts, and more. Enhance your projects with unique and stylish typography today!
These are the games used for testing models in the paper Aligning Superhuman AI with Human Behavior Chess as a Model System. The games have been converted from pgn to csv and the per move and per game information has been extracted. If you wish to see the original game files download the December 2019 standard games database release from Lichess. The game_id column matches that used by lichess. The games were selected to have both players of similar Elo Rating and for each divisions of 100 from 1000 to 25000, i.e 1000, 1100, ... 2500 there are 10,000 games. The meaning of each column should be straightforward. The per move information tends to be from the perspective of the active player, although the cp is not, instead the cp_rel columns is the relative centipawn value.