19,997 machine learning datasets
19,997 dataset results
The dataset contains 10 reference videos and 1467 degraded videos. The videos were transmitted via Microsoft Teams calls in 83 different network conditions and contain various typical videoconferencing impairments. It also includes P.910 Crowd subjective video MOS ratings (see paper for more info).
PAIR-LRT-Human Dataset contains pairs of thermal and RGB images captured using a FLIR Lepton3.5 thermal sensor and a Raspberry Pi camera v2, respectively. The dataset includes a total of 33,228 image pairs captured under different environmental conditions, with one human occupant in a standing position and an upper-body pose. The clothing of the occupant is in one of three colors. The images have a heatmap resolution of 16 × 12 and an RGB resolution of 128 × 96, and the field of view for each image is 71◦ × 57◦. The dataset includes two persons.
For more details see https://huggingface.co/datasets/jpwahle/autoregressive-paraphrase-dataset
we collected a new real-world dataset, called ALPIXVSR, using a ALPIX-Eiger event camera1 . The camera outputs well aligned RGB frames and events. The RGB frames enjoy a resolution of 3264 × 2448 and are generated by a carefully designed image signal processor(ISP) from RAW data with the Quad Bayer pattern , and the events have a resolution with 1632 × 1224.
BEAR (Benchmark on video Action Recognition) is a collection of 18 video datasets grouped into 5 categories (anomaly, gesture, daily, sports, and instructional), which covers a diverse set of real-world applications.
VR-Folding contains garment meshes of 4 categories from CLOTH3D dataset, namely Shirt, Pants, Top and Skirt. For flattening task, there are 5871 videos which contain 585K frames in total. For folding task, there are 3896 videos which contain 204K frames in total. The data for each frame include multi-view RGB-D images, object masks, full garment meshes, and hand poses.
ARKitTrack is a new RGB-D tracking dataset for both static and dynamic scenes captured by consumer-grade LiDAR scanners equipped on Apple's iPhone and iPad. ARKitTrack contains 300 RGBD sequences, 455 targets, and 229.7K video frames in total. This dataset has 123.9K pixel-level target masks along with the bounding box annotations and frame-level attributes.
AI Generated Content (AIGC) refers to any form of content, such as text, images, audio, or video, that is created with the help of artificial intelligence technology. With the flourishing development of deep learning, the efficiency of AIGC generation has increased, and AI-Generated Image (AGI) is becoming more prevalent in areas such as culture, entertainment, education, social media, etc.
The RT-PCR screening tests used and the results of which are reported in SI-DEP made it possible to suspect the presence of the worrisome variant (VOC) Alpha (20I/501Y.V1) and indistinctly from the VOC Beta (20H/501Y. V2) or Gamma (20J/501Y.V3). This screening strategy targeting Alpha, Beta and Gamma VOCs is no longer suited to the increasing diversity of emerging SARS-CoV-2 variants. Since 05/31/2021, the screening strategy has evolved to search for certain mutations of interest that can be found in different variants. It therefore no longer makes it possible to assign the infection to a specific variant but makes it possible to follow the evolution over time and in the territory of the proportion of infections due to a virus carrying these mutations.
The Fraunhofer Portugal AICOS EDoF Dataset was produced within the TAMI project and is composed of images of microscopic fields of view (FOV) of Liquid-based Cervical Cytology (LBC) samples. A total of 15 LBC samples were supplied by the Pathology Services from Hospital Fernando Fonseca and the Portuguese Oncology Institute of Porto. For each LBC sample, a set of images were obtained using a version of µSmartScope [1,2] prototype adapted to the cervical cytology use case [3,4].
A unique dataset comprising multimodal creative and designed documents containing images with corresponding captions paired with music based on around 50mood/themes.
A daily emerging stock market dataset (Chinese CSI 300 dataset) including 300 stocks and 5,088 time steps from the CSMAR database. We construct our stock dataset using a pool of stocks from the CSI 300 index for the last 21 years, from 01/02/2000 to 12/31/2020. Instead of all stocks in the market, we select the stocks that used to belong to the major market index CSI 300, and filter out stocks that have missing price data over the period.
The dataset covers the 2022-23 NBA regular season (2022-10-18 to 2023-01-20) which contains 691 games in 92 game days. There are 582 active players among the 30 teams. Besides 7 basic statistics, we collected 3 tracking statistics, and 3 advanced statistics. We use tracking statistics to more accurately reflect players' movements on the court, and advanced statistics to more properly represent a player's effectiveness and contribution to the game. Together, these two types of data give us a better understanding of factors that are not visible on the scoreboard.
pm2.5 time series data
We construct a large-scale conducting motion dataset, named ConductorMotion100, by deploying pose estimation on conductor view videos of concert performance recordings collected from online video platforms. The construction of ConductorMotion100 removes the need for expensive motion-capture equipment and makes full use of massive online video resources. As a result, the scale of ConductorMotion100 has reached an unprecedented length of 100 hours.
MuCeD, a dataset that is carefully curated and validated by expert pathologists from the All India Institute of Medical Science (AIIMS), Delhi, India. The H&E-stained histopathology images of the human duodenum in MuCeD are captured through an Olympus BX50 microscope at 20x zoom using a DP26 camera with each image being 1920x2148 in dimension. The dataset has 55 images, with bounding boxes for 2,090 IELs and 6,518 ENs annotated using the LabelMe software and are further validated by multiple pathologists. These cells are selected from the epithelial area -- a region of interest that has been explicitly segmented by experts. The epithelial area denotes the area of continuous villi and is used for cell detection, whereas rest of the area is masked out. Further, each image is sliced into 9 subimages and each subimage is re-scaled to 640x640, before it is given as input to object detection models. We divide 55 images into five folds of 11 images each and report 5-fold crossvalidation num
The StatCan Dialogue Dataset: Retrieving Data Tables through Conversations with Genuine Intents
The Archive Query Log (AQL) is a previously unused, comprehensive query log collected at the Internet Archive over the last 25 years. Its first version includes 356 million queries, 166 million search result pages, and 1.7 billion search results across 550 search providers. Although many query logs have been studied in the literature, the search providers that own them generally do not publish their logs to protect user privacy and vital business data. The AQL is the first publicly available query log that combines size, scope, and diversity, enabling research on new retrieval models and search engine analyses. Provided in a privacy-preserving manner, it promotes open research as well as more transparency and accountability in the search industry.
These are the test and training data used for experiments presented in BioNLP 2017.
This dataset contains 9 different seafood types collected from a supermarket in Izmir, Turkey for a university-industry collaboration project at Izmir University of Economics, and this work was published in ASYU 2020. The dataset includes gilt head bream, red sea bream, sea bass, red mullet, horse mackerel, black sea sprat, striped red mullet, trout, shrimp image samples.