19,997 machine learning datasets
19,997 dataset results
Synthetic dataset intended for benchmarking disentanglement frameworks.
EgoMon Gaze & Video Dataset is an Egocentric (first person) Dataset that consists of 7 videos of 30 minutes, more or less, each one of them. - 7 videos with the gaze information plotted on them. - The same videos (without the gaze information plotted on them). - A total of 13428 images, more or less, that corresponds to each frame per second of all these videos. - 7 text files with the gaze data extracted from each video.
VinDr-PCXR is an open, large-scale pediatric chest X-ray dataset for interpretation of common thoracic diseases in children. The dataset contains 9,125 CXR scans retrospectively collected from a major pediatric hospital in Vietnam between 2020 and 2021. Each scan was manually annotated by a pediatric radiologist who has more than ten years of experience. The dataset was labeled for the presence of 36 critical findings and 15 diseases. It aims to aid research in the detection of multiple findings and diseases.
To study the data-scarcity mitigation for learning-based visual localization methods via sim-to-real transfer, we curate and now present the CrossLoc benchmark datasets—a multimodal aerial sim-to-real data available for flights above nature and urban terrains. Unlike the previous computer vision datasets focusing on localization in a single domain (mostly real RGB images), the provided benchmark datasets include various multimodal synthetic cues paired to all real photos. Complementary to the paired real and synthetic data, we offer rich synthetic data that efficiently fills the flight envelope volume in the vicinity of the real data.
This data set provides fine-granular statistics on trading traffic generated by six global exchanges over the course of two days in February 2019 for a set of representative feeds and recorded by the systems of vwd Vereinigte Wirtschaftsdienste GmbH (now known as Infront Financial Technology GmbH).
This dataset accompanies the linked SerialTrack paper and provides test case data (2D/3D, varying particle density) across a range of synthetic and experimental imaging modalities. Included test cases can be used for further code development, validation of and comparisons for existing particle tracking codes, and/or evaluating and learning to use our SerialTrack code on known data.
This is a dataset for multi-document summarization in Portuguese, what means that it has examples of multiple documents (input) related to human-written summaries (output). In particular, it has entries of multiple related texts from Brazilian websites about a subject, and the summary is the Portuguese Wikipedia lead section on the same subject (lead: the first section, i.e., summary, of any Wipedia article). Input texts were extracted from BrWac corpus, and the output from Brazilian Wikipedia dumps page.
Hello Watt collects power usage data at a resolution of 30 minutes. To develop and test our disaggregation methods we consider a subsample consisting of power consumption of 5k households with off-peak pricing contracts for one month. In addition to the type of their water heating, some users also provide such metadata as the home surface area, and the number of inhabitants.
This is a set of debiased Natural Language Inference (NLI) datasets produced by the paper Generating Data to Mitigate Spurious Correlations in Natural Language Inference Datasets. The datasets are constructed by augmenting SNLI or MNLI with data samples that are generated to mitigate the spurious correlations in the original datasets. Please visit this repository for more details.
The dataset has been designed to represent true web videos in the wild, with good visual quality and diverse content characteristics, and will serve as evaluation basis for the Video Browser Showdown 2019-2021 and TREC Video Retrieval (TRECVID) Ad-Hoc Video Search tasks 2019-2021. The dataset comes with a shot segmentation (around 1 million shots) for which we analyze content specifics and statistics. Our research shows that the content of V3C1 is very diverse, has no predominant characteristics and provides a low self-similarity. Thus it is very well suited for video retrieval evaluations as well as for participants of TRECVID AVS or the VBS.
Different from the setting of domain adaptation which uses all labeled source and unlabeled target domain examples for training, domain examples should be divided into two disjoint parts: training and test. UDE-Office-Home is built from Office-Home, so the performance of domain-adapted or domain-expanded models on the source and target domain can be evaluated.
Different from the setting of domain adaptation which uses all labeled source and unlabeled target domain examples for training, domain examples should be divided into two disjoint parts: training and test. UDE-DomainNet is built from DomainNet, so the performance of domain-adapted or domain-expanded models on the source and target domain can be evaluated.
TEM image dataset containing four nanowire morphologies of bio-derived protein nanowires and synthetic peptide nanowires.
Forty prismatic lithium-ion pouch cells were built at the University of Michigan Battery Laboratory. The cells have a nominal capacity of 2.36Ah and comprise a NCM111 cathode and graphite anode. Cells were formed using two different formation protocols: "fast formation" and "baseline formation". After formation, cells were put under cycle life testing at room temperature and 45degC. Cells were cycled until the discharge capacities dropped below 50% of the initial capacities. Data was collected by the cycler equipment (Maccor) during both the formation process as well as during the cycling test. Data was processed in the Voltaiq software and subsequently exported as .csv files.
The dataset has been designed to represent true web videos in the wild, with good visual quality and diverse content characteristics, The test video collection for TRECVID-AVS2019-TRECVID-AVS2021, which contains 1,082,649 web video clips, with even more diverse content, no predominant characteristics and low self-similarity.
Onchocerciasis is causing blindness in over half a million people in the world today. Drug development for the disease is crippled as there is no way of measuring effectiveness of the drug without an invasive procedure. Drug efficacy measurement through assessment of viability of onchocerca worms requires the patients to undergo nodulectomy which is invasive, expensive, time-consuming, skill-dependent, infrastructure dependent and lengthy process.
Data collection was conducted by asking some adults from social media and some students from an elementary school to participate in our experiment. Table.1 shows the number of data gathered for recognizing each color. Due to the fact that two words are used for black in Persian, the number of black samples is more. In addition, because the color recognition is a RAN task, a sequence of data has been gathered. Table.2 depicts the number of sequence data for colors. For the meaningless words, 12 voices have been gathered on average for each word (there are 40 meaningless words in this task).
Dataset for our CVPR paper: "ISNAS-DIP: Image-Specific Neural Architecture Search for Deep Image Prior".
Leaves from genetically unique Juglans regia plants were scanned using X-ray micro-computed tomography (microCT) on the X-ray μCT beamline (8.3.2) at the Advanced Light Source (ALS) in Lawrence Berkeley National Laboratory (LBNL), Berkeley, CA USA).
The link includes both our OPDSynth and OPDReal dataset. For OPDSynth, we select objects with openable parts from an existing dataset of articulated 3D models PartNet-Mobility. For OPDReal, we reconstruct 3D polygonal meshes for articulated objects in real indoor environments and annotate their parts and articulation information.