Egocentric Video-Language Pretraining

Kevin Qinghong Lin, Alex Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Zhongcong Xu, Difei Gao, RongCheng Tu, Wenzhe Zhao, Weijie Kong, Chengfei Cai, Hongfa Wang, Dima Damen, Bernard Ghanem, Wei Liu, Mike Zheng Shou

2022-06-03Question Answering Video-Text Retrieval Text Retrieval Moment Queries Multi-Instance Retrieval Video Summarization Contrastive Learning Temporal Localization Object State Change Classification Action Recognition Retrieval Natural Language Queries

Paper PDF Code Code(official)

Abstract

Video-Language Pretraining (VLP), which aims to learn transferable representation to advance a wide range of video-text downstream tasks, has recently received increasing attention. Best performing works rely on large-scale, 3rd-person video-text datasets, such as HowTo100M. In this work, we exploit the recently released Ego4D dataset to pioneer Egocentric VLP along three directions. (i) We create EgoClip, a 1st-person video-text pretraining dataset comprising 3.8M clip-text pairs well-chosen from Ego4D, covering a large variety of human daily activities. (ii) We propose a novel pretraining objective, dubbed EgoNCE, which adapts video-text contrastive learning to the egocentric domain by mining egocentric-aware positive and negative samples. (iii) We introduce EgoMCQ, a development benchmark that is close to EgoClip and hence can support effective validation and fast exploration of our design decisions in EgoClip and EgoNCE. Furthermore, we demonstrate strong performance on five egocentric downstream tasks across three datasets: video-text retrieval on EPIC-KITCHENS-100; action recognition on Charades-Ego; natural language query, moment query, and object state change classification on Ego4D challenge benchmarks. The dataset and code are available at https://github.com/showlab/EgoVLP.

Results

Task	Dataset	Metric	Value	Model
Video	Query-Focused Video Summarization Dataset	F1 (avg)	49.72	EgoVLP
Question Answering	EgoTaskQA	Direct	42.51	EgoVLP
Activity Recognition	Charades-Ego	mAP	32.1	EgoVLP
Video Summarization	Query-Focused Video Summarization Dataset	F1 (avg)	49.72	EgoVLP
Action Recognition	Charades-Ego	mAP	32.1	EgoVLP
Natural Language Queries	Ego4D	R@1 IoU=0.3	10.46	EgoVLP
Natural Language Queries	Ego4D	R@1 IoU=0.5	6.24	EgoVLP
Natural Language Queries	Ego4D	R@1 Mean(0.3 and 0.5)	8.35	EgoVLP
Natural Language Queries	Ego4D	R@5 IoU=0.3	16.76	EgoVLP
Natural Language Queries	Ego4D	R@5 IoU=0.5	11.29	EgoVLP

Egocentric Video-Language Pretraining

Abstract

Results

Related Papers

Egocentric Video-Language Pretraining

Abstract

Results

Related Papers