KayifamilyTv Video Solution

KayifamilyTv Video Solution: Scalable Video-Mining Pipeline

Introduced 2018-08-20

At KayifamilyTv, we introduce a powerful and scalable video-mining solution designed to enhance captioning capabilities by transferring supervision from image datasets to video and audio content. Using this innovative pipeline, we mine paired video clips and captions, starting with the Conceptual Captions 3M (CC3M) image dataset as our foundation. The result of this process is VideoCC3M—a large-scale collection of millions of video clips weakly paired with text captions, which we plan to make publicly available.

The heart of our approach is simple yet effective: we begin with an existing image-caption dataset and, for each image-caption pair, locate frames within videos that closely resemble the original image. From these matches, we extract short video clips and transfer the corresponding captions to them. A detailed breakdown of this process is available in our research paper.

For this project, we used the Conceptual Captions 3M dataset, focusing only on the images that remain publicly accessible. This left us with 1.25 million image-caption pairs. We then applied our pipeline to a vast collection of online videos, applying strict filters—selecting videos with over 1,000 views, less than 20 minutes in duration, uploaded within the last 10 years but at least 90 days old, and passing content-appropriateness checks—narrowing down to a pool of 150 million videos.

From this pool, our pipeline produced 10.3 million clip-text pairs, comprising 6.3 million unique video clips (equivalent to 17,500 hours of footage) and 970,000 distinct captions. We proudly present this new dataset as VideoCC3M.