Learning to Anticipate Egocentric Actions by Imagination

Yu Wu, Linchao Zhu, Xiaohan Wang, Yi Yang, Fei Wu

2021-01-13Action Anticipation Autonomous Driving Contrastive Learning

Abstract

Anticipating actions before they are executed is crucial for a wide range of practical applications, including autonomous driving and robotics. In this paper, we study the egocentric action anticipation task, which predicts future action seconds before it is performed for egocentric videos. Previous approaches focus on summarizing the observed content and directly predicting future action based on past observations. We believe it would benefit the action anticipation if we could mine some cues to compensate for the missing information of the unobserved frames. We then propose to decompose the action anticipation into a series of future feature predictions. We imagine how the visual feature changes in the near future and then predicts future action labels based on these imagined representations. Differently, our ImagineRNN is optimized in a contrastive learning way instead of feature regression. We utilize a proxy task to train the ImagineRNN, i.e., selecting the correct future states from distractors. We further improve ImagineRNN by residual anticipation, i.e., changing its target to predicting the feature difference of adjacent frames instead of the frame content. This promotes the network to focus on our target, i.e., the future action, as the difference between adjacent frame features is more important for forecasting the future. Extensive experiments on two large-scale egocentric action datasets validate the effectiveness of our method. Our method significantly outperforms previous methods on both the seen test set and the unseen test set of the EPIC Kitchens Action Anticipation Challenge.

Results

Task	Dataset	Metric	Value	Model
Activity Recognition	EPIC-KITCHENS-55 (Seen test set (S1))	Top 1 Accuracy - Act.	14.66	ImagineRNN
Activity Recognition	EPIC-KITCHENS-55 (Seen test set (S1))	Top 1 Accuracy - Noun	22.79	ImagineRNN
Activity Recognition	EPIC-KITCHENS-55 (Seen test set (S1))	Top 1 Accuracy - Verb	35.44	ImagineRNN
Activity Recognition	EPIC-KITCHENS-55 (Seen test set (S1))	Top 5 Accuracy - Act.	34.98	ImagineRNN
Activity Recognition	EPIC-KITCHENS-55 (Seen test set (S1))	Top 5 Accuracy - Noun	52.09	ImagineRNN
Activity Recognition	EPIC-KITCHENS-55 (Seen test set (S1))	Top 5 Accuracy - Verb	79.72	ImagineRNN
Activity Recognition	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 1 Accuracy - Act.	9.25	ImagineRNN
Activity Recognition	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 1 Accuracy - Noun	15.5	ImagineRNN
Activity Recognition	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 1 Accuracy - Verb	29.33	ImagineRNN
Activity Recognition	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 5 Accuracy - Act.	22.19	ImagineRNN
Activity Recognition	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 5 Accuracy - Noun	35.78	ImagineRNN
Activity Recognition	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 5 Accuracy - Verb	70.67	ImagineRNN
Action Recognition	EPIC-KITCHENS-55 (Seen test set (S1))	Top 1 Accuracy - Act.	14.66	ImagineRNN
Action Recognition	EPIC-KITCHENS-55 (Seen test set (S1))	Top 1 Accuracy - Noun	22.79	ImagineRNN
Action Recognition	EPIC-KITCHENS-55 (Seen test set (S1))	Top 1 Accuracy - Verb	35.44	ImagineRNN
Action Recognition	EPIC-KITCHENS-55 (Seen test set (S1))	Top 5 Accuracy - Act.	34.98	ImagineRNN
Action Recognition	EPIC-KITCHENS-55 (Seen test set (S1))	Top 5 Accuracy - Noun	52.09	ImagineRNN
Action Recognition	EPIC-KITCHENS-55 (Seen test set (S1))	Top 5 Accuracy - Verb	79.72	ImagineRNN
Action Recognition	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 1 Accuracy - Act.	9.25	ImagineRNN
Action Recognition	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 1 Accuracy - Noun	15.5	ImagineRNN
Action Recognition	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 1 Accuracy - Verb	29.33	ImagineRNN
Action Recognition	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 5 Accuracy - Act.	22.19	ImagineRNN
Action Recognition	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 5 Accuracy - Noun	35.78	ImagineRNN
Action Recognition	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 5 Accuracy - Verb	70.67	ImagineRNN
Action Anticipation	EPIC-KITCHENS-55 (Seen test set (S1))	Top 1 Accuracy - Act.	14.66	ImagineRNN
Action Anticipation	EPIC-KITCHENS-55 (Seen test set (S1))	Top 1 Accuracy - Noun	22.79	ImagineRNN
Action Anticipation	EPIC-KITCHENS-55 (Seen test set (S1))	Top 1 Accuracy - Verb	35.44	ImagineRNN
Action Anticipation	EPIC-KITCHENS-55 (Seen test set (S1))	Top 5 Accuracy - Act.	34.98	ImagineRNN
Action Anticipation	EPIC-KITCHENS-55 (Seen test set (S1))	Top 5 Accuracy - Noun	52.09	ImagineRNN
Action Anticipation	EPIC-KITCHENS-55 (Seen test set (S1))	Top 5 Accuracy - Verb	79.72	ImagineRNN
Action Anticipation	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 1 Accuracy - Act.	9.25	ImagineRNN
Action Anticipation	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 1 Accuracy - Noun	15.5	ImagineRNN
Action Anticipation	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 1 Accuracy - Verb	29.33	ImagineRNN
Action Anticipation	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 5 Accuracy - Act.	22.19	ImagineRNN
Action Anticipation	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 5 Accuracy - Noun	35.78	ImagineRNN
Action Anticipation	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 5 Accuracy - Verb	70.67	ImagineRNN
2D Human Pose Estimation	EPIC-KITCHENS-55 (Seen test set (S1))	Top 1 Accuracy - Act.	14.66	ImagineRNN
2D Human Pose Estimation	EPIC-KITCHENS-55 (Seen test set (S1))	Top 1 Accuracy - Noun	22.79	ImagineRNN
2D Human Pose Estimation	EPIC-KITCHENS-55 (Seen test set (S1))	Top 1 Accuracy - Verb	35.44	ImagineRNN
2D Human Pose Estimation	EPIC-KITCHENS-55 (Seen test set (S1))	Top 5 Accuracy - Act.	34.98	ImagineRNN
2D Human Pose Estimation	EPIC-KITCHENS-55 (Seen test set (S1))	Top 5 Accuracy - Noun	52.09	ImagineRNN
2D Human Pose Estimation	EPIC-KITCHENS-55 (Seen test set (S1))	Top 5 Accuracy - Verb	79.72	ImagineRNN
2D Human Pose Estimation	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 1 Accuracy - Act.	9.25	ImagineRNN
2D Human Pose Estimation	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 1 Accuracy - Noun	15.5	ImagineRNN
2D Human Pose Estimation	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 1 Accuracy - Verb	29.33	ImagineRNN
2D Human Pose Estimation	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 5 Accuracy - Act.	22.19	ImagineRNN
2D Human Pose Estimation	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 5 Accuracy - Noun	35.78	ImagineRNN
2D Human Pose Estimation	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 5 Accuracy - Verb	70.67	ImagineRNN
Action Recognition In Videos	EPIC-KITCHENS-55 (Seen test set (S1))	Top 1 Accuracy - Act.	14.66	ImagineRNN
Action Recognition In Videos	EPIC-KITCHENS-55 (Seen test set (S1))	Top 1 Accuracy - Noun	22.79	ImagineRNN
Action Recognition In Videos	EPIC-KITCHENS-55 (Seen test set (S1))	Top 1 Accuracy - Verb	35.44	ImagineRNN
Action Recognition In Videos	EPIC-KITCHENS-55 (Seen test set (S1))	Top 5 Accuracy - Act.	34.98	ImagineRNN
Action Recognition In Videos	EPIC-KITCHENS-55 (Seen test set (S1))	Top 5 Accuracy - Noun	52.09	ImagineRNN
Action Recognition In Videos	EPIC-KITCHENS-55 (Seen test set (S1))	Top 5 Accuracy - Verb	79.72	ImagineRNN
Action Recognition In Videos	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 1 Accuracy - Act.	9.25	ImagineRNN
Action Recognition In Videos	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 1 Accuracy - Noun	15.5	ImagineRNN
Action Recognition In Videos	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 1 Accuracy - Verb	29.33	ImagineRNN
Action Recognition In Videos	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 5 Accuracy - Act.	22.19	ImagineRNN
Action Recognition In Videos	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 5 Accuracy - Noun	35.78	ImagineRNN
Action Recognition In Videos	EPIC-KITCHENS-55 (Unseen test set (S2)	Top 5 Accuracy - Verb	70.67	ImagineRNN

Learning to Anticipate Egocentric Actions by Imagination

Abstract

Results

Related Papers

Learning to Anticipate Egocentric Actions by Imagination

Abstract

Results

Related Papers