M3AV

Multimodal, Multigenre, and Multipurpose Audio-Visual

Introduced 2024-03-21

The M3AV (Multimodal, Multigenre, and Multipurpose Audio-Visual) is a novel dataset proposed for academic lectures¹. It contains almost 367 hours of videos from five sources covering topics in computer science, mathematics, and medical and biology¹⁴.

These videos carry rich multimodal information including speech, the facial and body movements of the speakers, as well as the texts and pictures in the slides and possibly even the papers¹. The dataset has high-quality human annotations of the slide text and spoken words, particularly high-valued name entities¹.

The M3AV dataset can be used for multiple audio-visual recognition and understanding tasks¹. Along with the dataset, three benchmark tasks are proposed that reflect the perception and understanding of the multimodal information in those videos. These tasks include automatic speech recognition (ASR) with a particular focus on contextual ASR (CASR), spontaneous text-to-speech (TTS) synthesis, and slide and script generation (SSG)¹.

This dataset effectively represents the typical scenario met in academic presentations where new terminologies constantly appear and are crucial to the overall understanding¹. It's a challenging dataset for evaluations performed on contextual speech recognition, speech synthesis, and slide and script generation tasks¹.

(1) M3AV: A Multimodal, Multigenre, and Multipurpose Audio-Visual Academic .... https://arxiv.org/html/2403.14168v2. (2) M3^3AV: A Multimodal, Multigenre, and Multipurpose Audio-Visual .... https://arxiv.org/abs/2403.14168. (3) Mac Benchmarks - Geekbench. https://browser.geekbench.com/mac-benchmarks. (4) M3AV: A Multimodal, Multigenre, and Multipurpose Audio-Visual Academic .... https://jack-zc8.github.io/M3AV-dataset-page/.