Temporally Coherent Embeddings for Self-Supervised Video Representation Learning

Joshua Knights, Ben Harwood, Daniel Ward, Anthony Vanderkop, Olivia Mackenzie-Ross, Peyman Moghadam

2020-03-21Representation Learning Metric Learning Self-Supervised Learning Action Recognition Temporal Action Localization Self-Supervised Action Recognition

Paper PDF Code(official)

Abstract

This paper presents TCE: Temporally Coherent Embeddings for self-supervised video representation learning. The proposed method exploits inherent structure of unlabeled video data to explicitly enforce temporal coherency in the embedding space, rather than indirectly learning it through ranking or predictive proxy tasks. In the same way that high-level visual information in the world changes smoothly, we believe that nearby frames in learned representations will benefit from demonstrating similar properties. Using this assumption, we train our TCE model to encode videos such that adjacent frames exist close to each other and videos are separated from one another. Using TCE we learn robust representations from large quantities of unlabeled video data. We thoroughly analyse and evaluate our self-supervised learned TCE models on a downstream task of video action recognition using multiple challenging benchmarks (Kinetics400, UCF101, HMDB51). With a simple but effective 2D-CNN backbone and only RGB stream inputs, TCE pre-trained representations outperform all previous selfsupervised 2D-CNN and 3D-CNN pre-trained on UCF101. The code and pre-trained models for this paper can be downloaded at: https://github.com/csiro-robotics/TCE

Results

Task	Dataset	Metric	Value	Model
Activity Recognition	UCF101	3-fold Accuracy	71.2	TCE (ResNet-50)
Activity Recognition	UCF101	3-fold Accuracy	68.8	TCE (ResNet-18, Split 1)
Activity Recognition	UCF101	3-fold Accuracy	68.2	TCE (ResNet18, Split 1)
Activity Recognition	HMDB51	Top-1 Accuracy	36.6	TCE (ResNet-50)
Activity Recognition	HMDB51	Top-1 Accuracy	34.2	TCE (ResNet-18)
Action Recognition	UCF101	3-fold Accuracy	71.2	TCE (ResNet-50)
Action Recognition	UCF101	3-fold Accuracy	68.8	TCE (ResNet-18, Split 1)
Action Recognition	UCF101	3-fold Accuracy	68.2	TCE (ResNet18, Split 1)
Action Recognition	HMDB51	Top-1 Accuracy	36.6	TCE (ResNet-50)
Action Recognition	HMDB51	Top-1 Accuracy	34.2	TCE (ResNet-18)

Temporally Coherent Embeddings for Self-Supervised Video Representation Learning

Abstract

Results

Related Papers

Temporally Coherent Embeddings for Self-Supervised Video Representation Learning

Abstract

Results

Related Papers