19,997 machine learning datasets
19,997 dataset results
Large multimodal models (LMMs) are processing increasingly longer and richer inputs. Albeit the progress, few public benchmark is available to measure such development. To mitigate this gap, we introduce LongVideoBench, a question-answering benchmark that features video-language interleaved inputs up to an hour long. Our benchmark includes 3,763 varying-length web-collected videos with their subtitles across diverse themes, designed to comprehensively evaluate LMMs on long-term multimodal understanding. To achieve this, we interpret the primary challenge as to accurately retrieve and reason over detailed multimodal information from long inputs. As such, we formulate a novel video question-answering task termed referring reasoning. Specifically, as part of the question, it contains a referring query that references related video contexts, called referred context. The model is then required to reason over relevant video details from the referred context. Following the paradigm of referri
Pentachromatic Cultural Palette Dataset is characterized by unique cultural semantics and values. It is constructed through a carefully - designed multi - step data synthesis process based on PRISM dataset, focusing on the cultural perspectives of different continents (Africa, Asia, Europe, America, and Oceania).
featureinput.npy Reflectance of 5 structures. Each row is 5 reflectance of a target material in 5 structures. Each row has 4000 values, representing 5 pairs of 800 values. For example, in the first 800 values, the first 400 is reflectance measured with target material on structure n, and the second 400 is reflectance measured without target material on structure n.
Whole-body, low-level control/manipulation demonstration dataset for ManiSkill-HAB. Demonstrations are organized by task-subtask-object. All demos use RGBD (128x128) and state. JSON files store metadata (tincluding even labels and success/failure mode), while HDF5 files store demonstration data.
ð BlendNet The dataset contains $12k$ samples. To balance cost savings with data quality and scale, we manually annotated $2k$ samples and used GPT-4o to annotate the remaining $10k$ samples.
ð CADBench CADBench is a comprehensive benchmark to evaluate the ability of LLMs to generate CAD scripts. It contains 500 simulated data samples and 200 data samples collected from online forums.
UAVDB is a high-resolution RGB video dataset meticulously designed for UAV detection tasks across diverse scales and complex backgrounds. Comprising 10,763 training, 2,720 validation, and 4,578 test images (18,061 total) across datasets and camera configurations, it addresses key limitations of existing datasets, such as inaccurate bounding box annotations and limited diversity in environmental contexts, thereby enhancing the reliability and applicability of detection algorithms in real-world scenarios.
IndirectRequests is an LLM-generated dataset of user utterances in a task-oriented dialogue setting where the user does not directly specify their preferred slot value.
The development of the remote sensing fine-grained ship classification field necessitates large-scale realistic fine-grained ship datasets. The FGSC-23 and FGSCR-42 datasets have been pivotal for practical applications, yet they exhibit limitations when assessing advanced classification methods. Specifically, FGSC-23 categorizes data into relatively broad classes, such as lumping all auxiliary ships into one category, which does not align with the granularity required in real-world scenarios. Additionally, the dataset's limited number of categories fails to satisfy the diverse needs of practical applications. FGSCR-42, while providing finer classifications within some broader categories, still lacks comprehensive subcategory coverage. Moreover, the dataset suffers from an imbalance in sample distribution, with six categories containing fewer than ten samples each, which can hinder effective model training.
The MVTec-FS dataset is a refined version of the MVTec AD dataset, designed for few-shot learning research. It contains instance-level annotations of anomaly images and is tailored to support tasks such as:
If you plan to test your method on our road network, you can find road network files in env\map.
Infrared dim-small target detection has gained increasing importance in both military and civilian applications due to its ability to detect thermal radiation, operate effectively at night, passively sense radiation, and offer strong concealment with high resistance to interference. These capabilities make it ideal for systems such as aircraft and bird surveillance, missile guidance, and maritime rescue operations. In these applications, the need for mid- to long-range observations often results in small targets that appear dim and are difficult to detect. This dataset, named SIRST-UAVB, provides infrared images captured in the 3â5 Ξm wavelength range using a mid-wave infrared camera, with a resolution of 640Ã512 pixels and shooting distances ranging from 100 to 800 meters. The dataset predominantly features small targets, which make up 94.3% of the total data and include unmanned aerial vehicles (UAVs) and birds. These targets are presented against complex backgrounds, such as skies,
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
The code and database provided in this repository are related to the paper "DeepNetBeam: A Framework for the Analysis of Functionally Graded Porous Beams", which explores the application of various machine-learning techniques for the analysis of functionally graded porous beams. The three approaches (PINN, DEM, Neural Operator) are implemented to allow flexibility and extension for future use.
This paper describes the first open dataset for full-scale and high-speed autonomous racing. Multi-modal sensor data has been collected from fully autonomous Indy race cars operating at speeds of up to 170 mph (273 kph). Six teams who raced in the Indy Autonomous Challenge have contributed to this dataset. The dataset spans 11 interesting racing scenarios across two race tracks which include solo laps, multi-agent laps, overtaking situations, high-accelerations, banked tracks, obstacle avoidance, pit entry and exit at different speeds. The dataset contains data from 27 racing sessions across the 11 scenarios with over 6.5 hours of sensor data recorded from the track. The data is organized and released in both ROS2 and nuScenes format. We have also developed the ROS2-to-nuScenes conversion library to achieve this. The RACECAR data is unique because of the high-speed environment of autonomous racing. We present several benchmark problems on localization, object detection and tracking (Li
PDS-DPO-9K
DRAM dataset is the first dataset to introduce a fully labeled test set for the task of semantic segmentation of art paintings. The dataset uses a subset of 12 classes used in the PascalVoc12 dataset: Bird, Boat, Bottle, Cat, Chair, Cow, Dog,Horse, Sheep, Person, Potted-Plant, and Background.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
A dataset for training and testing tin various problem types and multi-turn Q&A scenarios, including a training set, test set, and test scripts.
We introduce the AODRaw dataset, which offers 7,785 high-resolution real RAW images with 135,601 annotated instances spanning 62 categories, capturing a broad range of indoor and outdoor scenes under 9 distinct light and weather conditions. AODRaw supports RAW and sRGB object detection.