Weakly-supervised Temporal Action Localization by Uncertainty Modeling

Pilhyeon Lee, Jinglu Wang, Yan Lu, Hyeran Byun

2020-06-12Weakly Supervised Action Localization Action Classification Action Localization Multiple Instance Learning Weakly-supervised Temporal Action Localization Out-of-Distribution Detection Temporal Action Localization

Paper PDF Code(official)Code

Abstract

Weakly-supervised temporal action localization aims to learn detecting temporal intervals of action classes with only video-level labels. To this end, it is crucial to separate frames of action classes from the background frames (i.e., frames not belonging to any action classes). In this paper, we present a new perspective on background frames where they are modeled as out-of-distribution samples regarding their inconsistency. Then, background frames can be detected by estimating the probability of each frame being out-of-distribution, known as uncertainty, but it is infeasible to directly learn uncertainty without frame-level labels. To realize the uncertainty learning in the weakly-supervised setting, we leverage the multiple instance learning formulation. Moreover, we further introduce a background entropy loss to better discriminate background frames by encouraging their in-distribution (action) probabilities to be uniformly distributed over all action classes. Experimental results show that our uncertainty modeling is effective at alleviating the interference of background frames and brings a large performance gain without bells and whistles. We demonstrate that our model significantly outperforms state-of-the-art methods on the benchmarks, THUMOS'14 and ActivityNet (1.2 & 1.3). Our code is available at https://github.com/Pilhyeon/WTAL-Uncertainty-Modeling.

Results

Task	Dataset	Metric	Value	Model
Video	THUMOS 2014	mAP@0.1:0.5	51.6	Lee et al.
Video	THUMOS 2014	mAP@0.1:0.7	41.9	Lee et al.
Video	THUMOS 2014	mAP@0.5	33.7	Lee et al.
Video	THUMOS14	avg-mAP (0.1-0.5)	51.6	Lee et al.
Video	THUMOS14	avg-mAP (0.1:0.7)	41.9	Lee et al.
Video	THUMOS14	avg-mAP (0.3-0.7)	32.9	Lee et al.
Video	THUMOS’14	mAP@0.5	33.7	Lee et al.
Video	ActivityNet-1.3	mAP@0.5	37	Lee et al.
Video	ActivityNet-1.3	mAP@0.5:0.95	23.7	Lee et al.
Video	ActivityNet-1.2	Mean mAP	25.9	Lee et al.
Video	ActivityNet-1.2	mAP@0.5	41.2	Lee et al.
Temporal Action Localization	THUMOS 2014	mAP@0.1:0.5	51.6	Lee et al.
Temporal Action Localization	THUMOS 2014	mAP@0.1:0.7	41.9	Lee et al.
Temporal Action Localization	THUMOS 2014	mAP@0.5	33.7	Lee et al.
Temporal Action Localization	THUMOS14	avg-mAP (0.1-0.5)	51.6	Lee et al.
Temporal Action Localization	THUMOS14	avg-mAP (0.1:0.7)	41.9	Lee et al.
Temporal Action Localization	THUMOS14	avg-mAP (0.3-0.7)	32.9	Lee et al.
Temporal Action Localization	THUMOS’14	mAP@0.5	33.7	Lee et al.
Temporal Action Localization	ActivityNet-1.3	mAP@0.5	37	Lee et al.
Temporal Action Localization	ActivityNet-1.3	mAP@0.5:0.95	23.7	Lee et al.
Temporal Action Localization	ActivityNet-1.2	Mean mAP	25.9	Lee et al.
Temporal Action Localization	ActivityNet-1.2	mAP@0.5	41.2	Lee et al.
Zero-Shot Learning	THUMOS 2014	mAP@0.1:0.5	51.6	Lee et al.
Zero-Shot Learning	THUMOS 2014	mAP@0.1:0.7	41.9	Lee et al.
Zero-Shot Learning	THUMOS 2014	mAP@0.5	33.7	Lee et al.
Zero-Shot Learning	THUMOS14	avg-mAP (0.1-0.5)	51.6	Lee et al.
Zero-Shot Learning	THUMOS14	avg-mAP (0.1:0.7)	41.9	Lee et al.
Zero-Shot Learning	THUMOS14	avg-mAP (0.3-0.7)	32.9	Lee et al.
Zero-Shot Learning	THUMOS’14	mAP@0.5	33.7	Lee et al.
Zero-Shot Learning	ActivityNet-1.3	mAP@0.5	37	Lee et al.
Zero-Shot Learning	ActivityNet-1.3	mAP@0.5:0.95	23.7	Lee et al.
Zero-Shot Learning	ActivityNet-1.2	Mean mAP	25.9	Lee et al.
Zero-Shot Learning	ActivityNet-1.2	mAP@0.5	41.2	Lee et al.
Action Localization	THUMOS 2014	mAP@0.1:0.5	51.6	Lee et al.
Action Localization	THUMOS 2014	mAP@0.1:0.7	41.9	Lee et al.
Action Localization	THUMOS 2014	mAP@0.5	33.7	Lee et al.
Action Localization	THUMOS14	avg-mAP (0.1-0.5)	51.6	Lee et al.
Action Localization	THUMOS14	avg-mAP (0.1:0.7)	41.9	Lee et al.
Action Localization	THUMOS14	avg-mAP (0.3-0.7)	32.9	Lee et al.
Action Localization	THUMOS’14	mAP@0.5	33.7	Lee et al.
Action Localization	ActivityNet-1.3	mAP@0.5	37	Lee et al.
Action Localization	ActivityNet-1.3	mAP@0.5:0.95	23.7	Lee et al.
Action Localization	ActivityNet-1.2	Mean mAP	25.9	Lee et al.
Action Localization	ActivityNet-1.2	mAP@0.5	41.2	Lee et al.
Weakly Supervised Action Localization	THUMOS 2014	mAP@0.1:0.5	51.6	Lee et al.
Weakly Supervised Action Localization	THUMOS 2014	mAP@0.1:0.7	41.9	Lee et al.
Weakly Supervised Action Localization	THUMOS 2014	mAP@0.5	33.7	Lee et al.
Weakly Supervised Action Localization	THUMOS14	avg-mAP (0.1-0.5)	51.6	Lee et al.
Weakly Supervised Action Localization	THUMOS14	avg-mAP (0.1:0.7)	41.9	Lee et al.
Weakly Supervised Action Localization	THUMOS14	avg-mAP (0.3-0.7)	32.9	Lee et al.
Weakly Supervised Action Localization	THUMOS’14	mAP@0.5	33.7	Lee et al.
Weakly Supervised Action Localization	ActivityNet-1.3	mAP@0.5	37	Lee et al.
Weakly Supervised Action Localization	ActivityNet-1.3	mAP@0.5:0.95	23.7	Lee et al.
Weakly Supervised Action Localization	ActivityNet-1.2	Mean mAP	25.9	Lee et al.
Weakly Supervised Action Localization	ActivityNet-1.2	mAP@0.5	41.2	Lee et al.

Weakly-supervised Temporal Action Localization by Uncertainty Modeling

Abstract

Results

Related Papers

Weakly-supervised Temporal Action Localization by Uncertainty Modeling

Abstract

Results

Related Papers