19,997 machine learning datasets
19,997 dataset results
Top-notch Flutter App Development Company
RepoQA is a benchmark that aims to exercise the long-context code understanding ability of Language Learning Models (LLMs)². It supports repositories from 5 programming languages: Python, C++, TypeScript, Rust, and Java².
This is an image splicing dataset including different types of preprocessing and postprocessing techniques. Foreground objects are taken from HRSOD and background images are taken from BG20k datasets. 95000 train and 5000 test images are provided.
Welcome to "Reliable Air Ambulance Services in Hyderabad: Your Guide to Emergency Medical Transport." This forum is dedicated to providing comprehensive information, support, and resources about air ambulance services in Hyderabad. Whether you're looking for details on emergency medical transport, comparing service providers, understanding the costs involved, or seeking firsthand experiences and reviews, this is the place for you. Join our community to ask questions, share insights, and stay informed about the latest developments in air ambulance services in Hyderabad. Your health and safety are our top priorities
This dataset contains images and annotations for scene text detection and recognition. It is made up of two parts: (1) 1,175 images manually labeled with a total of 59,588 text instances at the line and word levels; and (2) 929 signboard images collected from the VinText, Total-Text, and ICDAR15 datasets. Each text instance in the first part of our dataset has a quadrilateral bounding box and a ground truth character sequence associated with it. In the second part, images are selected if they contain signboards. This portion of the dataset comprises 20,261 text instances at word levels. This brings the total text instances in our final dataset up to 79,814. Following the ICDAR15 standard, we annotated each image with all of the text instances, polygons, and content that were present. Manual annotations were done on each and every image.
NeoRL-2 includes new task scenarios that better reflect real-world task properties and includes traditional control methods as the data-collecting method. In summary, our contributions are as follows:
The RefoMB dataset is part of a project called RLAIF-V, which stands for "Aligning MLLMs through Open-Source AI Feedback for Super GPT-4V Trustworthiness." It's an open-source multimodal preference dataset that contains more than 30,000 high-quality comparison pairs. The dataset is designed to reduce hallucination in different Multimodal Large Language Models (MLLMs) and improve their trustworthiness by providing high-quality feedback data and an online feedback learning algorithm.
Cancer genomics and precision oncology: The TCGA Research Network started in 2005 has profiled and analyzed a large number of human tumors to discover molecular aberrations at the DNA, RNA, protein, and epigenetic levels and thereby provided reliable diagnostic and prognostic biomarkers for different cancer types since then.The presence of mutated genes is strongly correlated with cancer incidence, very specific causative genes or a small set of genes for most cancers have not been confirmed after decades of genomic studies. Nobel laureate James D. Watson opined at Cancer World 2013: "We can go ahead and sequence every piece of DNA that has ever existed, but I don't think we'll find the Achilles heel of cancer. Importantly,, it is not only necessary to associate genetic mutations with different cancers but also to work on the mechanism of action of mutagens by focusing on enzymes which could invariably mediate oncogenic transformations. For example, overexpression of the ribonucleotide
Cancer genomics and precision oncology: The TCGA Research Network started in 2005 has profiled and analyzed a large number of human tumors to discover molecular aberrations at the DNA, RNA, protein, and epigenetic levels and thereby provided reliable diagnostic and prognostic biomarkers for different cancer types since then.The presence of mutated genes is strongly correlated with cancer incidence, very specific causative genes or a small set of genes for most cancers have not been confirmed after decades of genomic studies. Nobel laureate James D. Watson opined at Cancer World 2013: "We can go ahead and sequence every piece of DNA that has ever existed, but I don't think we'll find the Achilles heel of cancer. Importantly,, it is not only necessary to associate genetic mutations with different cancers but also to work on the mechanism of action of mutagens by focusing on enzymes which could invariably mediate oncogenic transformations. For example, overexpression of the ribonucleotide
TAL-SCQ5K-EN/TAL-SCQ5K-CN are high-quality mathematical competition datasets in English and Chinese language created by TAL Education Group, each consisting of 5K questions(3K training and 2K testing). The questions are in the form of multiple-choice and cover mathematical topics at the primary, junior high, and high school levels. In addition, detailed solution steps are provided to facilitate CoT training and all the mathematical expressions in the questions have been presented as standard text-mode Latex.
MathEval is a benchmark dedicated to a comprehensive evaluation of the mathematical capabilities of large models. It encompasses over 20 evaluation datasets across various mathematical domains, with over 30,000 math problems. The goal is to thoroughly evaluate the performance of large models in tackling problems spanning a wide range of difficulty levels and diverse mathematical subfields (i.e. arithmetic, elementary mathematics, middle and high school competition topics, advanced mathematical, etc.). It serves as a trustworthy reference for cross-model comparisons of mathematical abilities among large models at the current stage and guides how to further enhance the mathematical capabilities of these models in the future.
AGGA (Academic Guidelines for Generative AIs) is a dataset of 80 academic guidelines for the usage of generative AIs and large language models in academia, selected systematically and collected from official university websites across six continents. Comprising 181,225 words, the dataset supports natural language processing tasks such as language modeling, sentiment and semantic analysis, model synthesis, classification, and topic labeling. It can also serve as a benchmark for ambiguity detection and requirements categorization. This resource aims to facilitate research on AI governance in educational contexts, promoting a deeper understanding of the integration of AI technologies in academia.
GenAI-Bench benchmark consists of 1,600 challenging real-world text prompts sourced from professional designers. Compared to benchmarks such as PartiPrompt and T2I-CompBench, GenAI-Bench captures a wider range of aspects in the compositional text-to-visual generation, ranging from basic (scene, attribute, relation) to advanced (counting, comparison, differentiation, logic). GenAI-Bench benchmark also collects human alignment ratings (1-to-5 Likert scales) on images and videos generated by ten leading models, such as Stable Diffusion, DALL-E 3, Midjourney v6, Pika v1, and Gen2.
Multimodal Large Language Models (MLLMs) have shown significant promise in various applications, leading to broad interest from researchers and practitioners alike. However, a comprehensive evaluation of their long-context capabilities remains underexplored. To address these gaps, we introduce the MultiModal Needle-in-a-haystack (MMNeedle) benchmark, specifically designed to assess the long-context capabilities of MLLMs. Besides multi-image input, we employ image stitching to further increase the input context length, and develop a protocol to automatically generate labels for sub-image level retrieval. Essentially, MMNeedle evaluates MLLMs by stress-testing their capability to locate a target sub-image (needle) within a set of images (haystack) based on textual instructions and descriptions of image contents. This setup necessitates an advanced understanding of extensive visual contexts and effective information retrieval within long-context image inputs. With this benchmark, we evalu
This resource contains training and test data for detecting "intent" sentences in email messages. This data comes from the Enron email corpus. Each labeled example is one sentence from an email. We define "intent" here to correspond primarily to the categories "request" and "propose" in the paper:
The DevOps-Eval is an industrial-first evaluation benchmark specifically designed for Large Language Models (LLMs) in the DevOps/AIOps domain¹. It was released by Ant Group in collaboration with Peking University³.
CodeFuseEval is a Code Generation benchmark that combines the multi-tasking scenarios of CodeFuse Model with the benchmarks of HumanEval-x and MBPP. This benchmark is designed to evaluate the performance of models in various multi-tasking tasks, including code completion, code generation from natural language, test case generation, cross-language code translation, and code generation from Chinese commands, among others.
The SKEMPI database contains data on the changes in thermodynamic parameters and kinetic rate constants upon mutation, for protein-protein interactions for which a structure of the complex has been solved and is available in the Protein Databank.
This dataset is flood data in the city of Parepare, South Sulawesi Province, which contains video data collected from social media Instagram. This dataset was created to develop deep learning methods for recognizing floods and surrounding objects, specializing in semantic segmentation methods. This dataset consists of three folders, namely raw video data collected from Instagram, image data resulting from splitting the video into several images, and annotation data containing images that have been color-labeled according to their objects. There are 6 object classifications based on color labels, namely: floods (blue light), buildings (red), plants (green), people (sage), vehicles (orange), and sky (dark blue). This dataset has data in image (JPEG/PNG) and video (MP4) formats. This dataset is suitable for object recognition tasks with the semantic segmentation method. In addition, because this dataset contains original data in the form of videos and images, it can be developed for other
IndirectRequests is an LLM-generated dataset of user utterances in a task-oriented dialogue setting where the user does not directly specify their preferred slot value.