19,997 machine learning datasets
19,997 dataset results
This tree was generated as indicated above in the methods. The model chosen by the algorithm was LG+F+R8.
This tree was generated as indicated above in the text for file S3 using model LG+F+G.
This tree was generated as indicated above in the methods. The model chosen by the algorithm was LG+R8.
Validity and Novelty are determined in a comparative setting between two conclusions at a time. For Validity and Novelty possible labels are "Conclusion 1 is better", "tie" and "Conclusion 2 is better", for Validity and Novelty respectively.
A dataset for image editing containing >450k samples of:
RGB Arabic Alphabet Sign Language (AASL) dataset
The linked repository holds data from a controlled single-obstacle avoidance experiment recorded in a motion laboratory. If you are using this data, please cite the following sources for the data
In order to avoid the training process of the model being influenced by a portion of HOI classes with a very small number of instances, we remove some of the HOI classes containing a very small number of instances and HOI classes with no interaction from the training \textbf{S}et for the \textbf{D}ouble \textbf{C}hallenge. Finally, there are total 321 HOI classes, 74 object classes and 93 action classes. The training and testing set contain 37,155 and 9,666 images, respectively.
The HC3 (Human ChatGPT Comparison Corpus) dataset consists of nearly 40K questions and their corresponding human/ChatGPT answers. The motivation for this dataset was to study ChatGPT's answers in contrast to human's answers. The questions range from a wide variety of domains, including open-domain, financial, medical, legal, and psychological areas.
StEduCov, a dataset annotated for stances toward online education during the COVID-19 pandemic. StEduCov has 17,097 tweets gathered over 15 months, from March 2020 to May 2021, using Twitter API. The tweets are manually annotated into agree, disagree or neutral classes. We used a set of relevant hashtags and keywords. Specifically, we utilised a combination of hashtags, such as '#COVID 19' or '#Coronavirus' with keywords, such as 'education', 'online learning', 'distance learning' and 'remote learning'. To ensure high annotation quality, three different annotators annotated each tweet and at least one of the reviewers from three judges revised it. They were guided by some instructions, such as that in the case of disagree class, there should be a clear negative statement about online education or its impact. Also, if the tweet is negative but refers to other people (e.g. 'my children hate online learning').
We provide separate training, development and test data. The training data is available right away. The development and test data will be released in several stages, starting with a release of the development sources only.
GovDocs is a corpus of nearly 1 million documents that are freely available for research and may be, to the best of the authors' knowledge, freely redistributed. These documents were obtained by performing searches for words randomly chosen from the Unix dictionary, numbers randomly chosen between 1 and 1 million, and randomized combinations of the two, for documents of specified file types that resided on web servers in the .gov domain using the Yahoo an Google search engines. The documents are representative of a diverse sample of real-world files of various formats produced by a variety of tools, including any malware that may be present in the files. Therefore, the corpus has been used in digital forensics, malware analysis, computer vision, and natural language processing research.
The 50-ha plot at Barro Colorado Island was initially demarcated and fully censused in 1982, and has been fully censused 7 times since, every 5 years from 1985 through 2015. Every measurement of every stem over 8 censuses is included in this archive. Most users will need only the 8 R Analytical Tables in the format tree, which come here zipped together into a single archive (bci.tree.zip), plus the single R Species Table.
A set of 248 search queries annotated with the correct diagnosis. The diagnosis is referenced with a Concept Unique Identifier (CUI). In a retrieval setting, the task consists of retrieving an article from the FindZebra corpus with a CUI that matches the query CUI.
MTTN is a large scale derived and synthesized dataset built with on real prompts and indexed with popular image-text datasets like MS-COCO, Flickr, etc. MTTN consists of over 2.4M sentences that are divided over 5 stages creating a combination amounting to over 12M pairs, along with a vocab size of consisting more than 300 thousands unique words that creates an abundance of variations.
UICaption is a dataset of 114k UI images paired with descriptions of their functionality. It is designed for the tasks of UI action entailment, instruction-based UI image retrieval, grounding referring expressions, and UI entity recognition.
PushWorld is an environment with simplistic physics that requires manipulation planning with both movable obstacles and tools. It contains more than 200 PushWorld puzzles in PDDL and in an OpenAI Gym environment.
Reader Emotion News 20k Dataset
The English Headline Treebank (EHT) is an English headline treebank of 1,055 manually annotated and adjudicated universal dependency (UD) syntactic dependency trees to encourage research in improving NLP pipelines for English headlines.
ConsInv is a stereo RGB + IMU dataset designed for Dynamic SLAM testing and contains two subsets: