Partially Relevant Video Retrieval

Jianfeng Dong, Xianke Chen, Minsong Zhang, Xun Yang, ShuJie Chen, Xirong Li, Xun Wang

2022-08-26Video Retrieval Partially Relevant Video Retrieval Multiple Instance Learning Text to Video Retrieval Video Captioning Moment Retrieval Retrieval Video Corpus Moment Retrieval

Paper PDF Code(official)

Abstract

Current methods for text-to-video retrieval (T2VR) are trained and tested on video-captioning oriented datasets such as MSVD, MSR-VTT and VATEX. A key property of these datasets is that videos are assumed to be temporally pre-trimmed with short duration, whilst the provided captions well describe the gist of the video content. Consequently, for a given paired video and caption, the video is supposed to be fully relevant to the caption. In reality, however, as queries are not known a priori, pre-trimmed video clips may not contain sufficient content to fully meet the query. This suggests a gap between the literature and the real world. To fill the gap, we propose in this paper a novel T2VR subtask termed Partially Relevant Video Retrieval (PRVR). An untrimmed video is considered to be partially relevant w.r.t. a given textual query if it contains a moment relevant to the query. PRVR aims to retrieve such partially relevant videos from a large collection of untrimmed videos. PRVR differs from single video moment retrieval and video corpus moment retrieval, as the latter two are to retrieve moments rather than untrimmed videos. We formulate PRVR as a multiple instance learning (MIL) problem, where a video is simultaneously viewed as a bag of video clips and a bag of video frames. Clips and frames represent video content at different time scales. We propose a Multi-Scale Similarity Learning (MS-SL) network that jointly learns clip-scale and frame-scale similarities for PRVR. Extensive experiments on three datasets (TVR, ActivityNet Captions, and Charades-STA) demonstrate the viability of the proposed method. We also show that our method can be used for improving video corpus moment retrieval.

Results

Task	Dataset	Metric	Value	Model
Text to Video Retrieval	TVR	Recall@Sum	172.3	ms-sl
Text to Video Retrieval	ActivityNet Captions	Recall@Sum	140.1	ms-sl
Text to Video Retrieval	Charades-STA	Recall@Sum	68.4	ms-sl
10-shot image generation	TVR	Recall@Sum	172.3	ms-sl
10-shot image generation	ActivityNet Captions	Recall@Sum	140.1	ms-sl
10-shot image generation	Charades-STA	Recall@Sum	68.4	ms-sl

Partially Relevant Video Retrieval

Abstract

Results

Related Papers

Partially Relevant Video Retrieval

Abstract

Results

Related Papers