TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

SAIL 2017 (Sentiment Analysis for Indian Languages)

India is a linguistic area with one of the longest histories of contact, influence, use, teaching and learning of English-in-diaspora in the world (Kachru and Nelson, 2006). Thus, a huge number of Indians active on the internet are able in English communication to some degree. India also enjoys huge diversity in language. Apart from Hindi, it has several regional languages that are the primary tongue of people native to the region. This is to the extent that social media including Facebook, WhatsApp, Twitter, etc. contain more than one language, and such phenomena are called code-mixing and code-switching. On the other side, the evolution of sentiments from such social media texts have also created many new opportunities for information access and language technology, but also many new challenges, making it one of the prime present-day research areas. Sentiment analysis in code-mixed data has several real-life applications in opinion mining from social media campaign to feedback analys

1 papers3 benchmarksTexts

Kepler Exoplanet Search Results

Context The Kepler Space Observatory is a NASA-build satellite that was launched in 2009. The telescope is dedicated to searching for exoplanets in star systems besides our own, with the ultimate goal of possibly finding other habitable planets besides our own. The original mission ended in 2013 due to mechanical failures, but the telescope has nevertheless been functional since 2014 on a "K2" extended mission.

1 papers1 benchmarksTabular

SpanEX

Reasoning over spans of tokens from different parts of the input is essential for natural language understanding (NLU) tasks such as fact-checking (FC), machine reading comprehension (MRC) or natural language inference (NLI). We introduce SpanEx, a multi-annotator dataset of human-annotated span interaction explanations for two NLU tasks: NLI and FC.

1 papers0 benchmarksTexts

RadioGalaxyNET

A multimodal dataset of radio galaxies and their corresponding infrared hosts.

1 papers0 benchmarks

Neural Field Arena - Classification

Neural fields (NeFs) have recently emerged as a versatile method for modeling signals of various modalities, including images, shapes, and scenes. Subsequently, many works have explored the use of NeFs as representations for downstream tasks, e.g. classifying an image based on the parameters of a NeF that has been fit to it. However, the impact of the NeF hyperparameters on their quality as downstream representation is scarcely understood and remains largely unexplored. This is partly caused by the large amount of time required to fit datasets of neural fields.

1 papers0 benchmarks3D, Images

Nigeria 2020 cropland dataset

Hand-labelled dataset of crop and non-crop labels distributed throughout Nigeria with respective hd5f data arrays.

1 papers0 benchmarksTime series

PseudoMD-1M

Pre-training dataset used in paper "From Artificially Real to Real: Leveraging Pseudo Data from Large Language Models for Low-Resource Molecule Discovery" (AAAI 2024)

1 papers0 benchmarks

SQL injection dataset

This is a cleaned version of the dataset introduced in Kaggle by user SAJID576. The original dataset's link is https://www.kaggle.com/datasets/sajid576/sql-injection-dataset

1 papers0 benchmarks

GOD (Generic Object Decoding)

The Generic Object Decoding (GOD) Dataset is a specialized resource developed for fMRI-based decoding. It aggregates fMRI data gathered through the presentation of images from 200 representative object categories, originating from the 2011 fall release of ImageNet. The training session incorporated 1,200 images (8 per category from 150 distinct object categories). In contrast, the test session included 50 images (one from each of the 50 object categories). It is noteworthy that the categories in the test session were unique from those in the training session and were introduced in a randomized sequence across runs. On five subjects the fMRI scanning was conducted.

1 papers1 benchmarksImages, Medical, fMRI

CCGR (Cross-Covariate Gait Recognition)

CCGR (Cross-Covariate Gait Recognition), the first gait dataset for studying cross-covariate challenges, contains 970 subjects, approximately 1.6 million sequences, 53 covariates, and 33 views.

1 papers0 benchmarks

AVOIDDS: A dataset for vision-based aircraft detection

Aircraft collision avoidance systems rely on sensor information to detect and track intruding aircraft so that they may issue proper collision avoidance advisories. While typical surveillance sensors for manned aircraft include transponders and onboard radar, autonomous aircraft will require additional sensors both for redundancy and to replace the visual acquisition typically performed by the pilot. As a result, the community has proposed detecting other aircraft using vision-based sensors such as cameras. These sensors require the development of techniques to process images of the environment to detect intruding aircraft. To boost this development, this artifact provides a dataset of 72,000 labeled images of intruder aircraft with various lighting conditions, weather conditions, relative geometries, and geographic locations. For more information on the structure of this dataset as well as benchmark models and a full simulator, see https://github.com/sisl/VisionBasedAircraftDAA.

1 papers0 benchmarks

GEval for KGRC-RDF-star

This repository is an extension of GEval. This repository contains a (software) evaluation framework to perform evaluation and comparison on RDF-star graph embedding techniques. The gold standard datasets for evaluation were created from KGRC-RDF-star. Please see here.

1 papers0 benchmarksGraphs

Turbulence

$\textbf{Turbulence}$ is a new benchmark for systematically evaluating the correctness and robustness of instruction-tuned large language models (LLMs) for code generation. Turbulence consists of a large set of natural language question templates, each of which is a programming problem, parameterised so that it can be asked in many different forms. Each question template has an associated test oracle that judges whether a code solution returned by an LLM is correct. Thus, from a single question template, it is possible to ask an LLM a $\textit{neighbourhood}$ of very similar programming questions, and assess the correctness of the result returned for each question. This new benchmark systematically and automatically identifies cases where LLMs are able to solve some problems in a neighbourhood but do not manage to generalise to solve the whole neighbourhood. Therefore, this method is effective at highlighting robustness issues.

1 papers1 benchmarks

2D-ATOMS

Official dataset for Towards A Holistic Landscape of Situated Theory of Mind in Large Language Models. Ziqiao Ma, Jacob Sansom, Run Peng, Joyce Chai. EMNLP Findings, 2023.

1 papers0 benchmarksTexts

LTI LangID Corpus

The LTI LangID Corpus is a dataset used for language identification (LangID) tasks. It contains text data in various languages. The dataset has had multiple releases, with the first release containing 781 "core" languages and 1091 languages overall.

1 papers0 benchmarks

Wikimedia

The Wikimedia dataset refers to a collection of data related to Wikimedia projects, which include Wikipedia, Wikivoyage, Wiktionary, Wikisource, and others. These datasets are publicly available and can be used for various purposes such as research, backup, and offline use.

1 papers0 benchmarks

MegaAcceptability

The MegaAcceptability dataset is a collection of ordinal acceptability judgments for clause-embedding verbs of English in various surface-syntactic frames and matrix tenses. Here are some key details: - It includes judgments for 1,007 clause-embedding verbs of English in 50 surface-syntactic frames and three matrix tenses. - The dataset combines the MegaAcceptability version 1.0 and data collected for 25,000 additional verb-frame pairs on Amazon’s Mechanical Turk using Ibex on Mechanical Turk. - The dataset has been used to address questions in linguistic theory.

1 papers0 benchmarks

BikeDNA BIG: Denmark analysis

See https://zenodo.org/records/8340383

1 papers0 benchmarks

BikeDNA Denmark Analysis

See https://zenodo.org/records/10185500

1 papers0 benchmarks

Amazon Polarity

The Amazon Polarity dataset is a set of reviews from Amazon. The dataset is constructed by taking review scores 1 and 2 as negative (class 1), and 4 and 5 as positive (class 2). Reviews with a score of 3 are ignored. The dataset spans a period of 18 years, including approximately 35 million reviews up to March 2013. Each class in the dataset has 1,800,000 training samples and 200,000 testing samples. The dataset includes product and user information, ratings, and a plaintext review.

1 papers0 benchmarks
PreviousPage 483 of 1000Next