AudioSet
AudioVideosCC BY 4.0Introduced 2017-01-01
Audioset is an audio event dataset, which consists of over 2M human-annotated 10-second video clips. These clips are collected from YouTube, therefore many of which are in poor-quality and contain multiple sound-sources. A hierarchical ontology of 632 event classes is employed to annotate these data, which means that the same sound could be annotated as different labels. For example, the sound of barking is annotated as Animal, Pets, and Dog. All the videos are split into Evaluation/Balanced-Train/Unbalanced-Train set.
Source: Curriculum Audiovisual Learning
Benchmarks
Audio Classification/Test mAPAudio Classification/AUCAudio Classification/d-primeAudio Source Separation/SDRAudio Source Separation/SARAudio Source Separation/SIRAudio Source Separation/SDRiAudio Source Separation/SI-SDRiAudio Tagging/mean average precisionClassification/Test mAPClassification/AUCClassification/d-primeClassification/Average mAPMulti-modal Classification/Average mAPTarget Sound Extraction/SDRiTarget Sound Extraction/SI-SDRi