Spatiotemporal CNN for Video Object Segmentation

Kai Xu, Longyin Wen, Guorong Li, Liefeng Bo, Qingming Huang

2019-04-04CVPR 2019 6Visual Object Tracking Semi-Supervised Video Object Segmentation Segmentation Semantic Segmentation Video Segmentation Video Object Segmentation Video Semantic Segmentation

Paper PDF Code(official)

Abstract

In this paper, we present a unified, end-to-end trainable spatiotemporal CNN model for VOS, which consists of two branches, i.e., the temporal coherence branch and the spatial segmentation branch. Specifically, the temporal coherence branch pretrained in an adversarial fashion from unlabeled video data, is designed to capture the dynamic appearance and motion cues of video sequences to guide object segmentation. The spatial segmentation branch focuses on segmenting objects accurately based on the learned appearance and motion cues. To obtain accurate segmentation results, we design a coarse-to-fine process to sequentially apply a designed attention module on multi-scale feature maps, and concatenate them to produce the final prediction. In this way, the spatial segmentation branch is enforced to gradually concentrate on object regions. These two branches are jointly fine-tuned on video segmentation sequences in an end-to-end manner. Several experiments are carried out on three challenging datasets (i.e., DAVIS-2016, DAVIS-2017 and Youtube-Object) to show that our method achieves favorable performance against the state-of-the-arts. Code is available at https://github.com/longyin880815/STCNN.

Results

Task	Dataset	Metric	Value	Model
Video	DAVIS 2017 (val)	F-measure (Mean)	64.6	Spatiotemporal CNN
Video	DAVIS 2017 (val)	J&F	61.65	Spatiotemporal CNN
Video	DAVIS 2017 (val)	Jaccard (Mean)	58.7	Spatiotemporal CNN
Video	DAVIS 2016	F-measure (Mean)	83.8	Spatiotemporal CNN
Video	DAVIS 2016	J&F	83.8	Spatiotemporal CNN
Video	DAVIS 2016	Jaccard (Mean)	83.8	Spatiotemporal CNN
Video	YouTube	mIoU	0.796	Spatiotemporal CNN
Video	DAVIS (no YouTube-VOS training)	D16 val (F)	83.8	STCNN
Video	DAVIS (no YouTube-VOS training)	D16 val (G)	83.8	STCNN
Video	DAVIS (no YouTube-VOS training)	D16 val (J)	83.8	STCNN
Video	DAVIS (no YouTube-VOS training)	D17 val (F)	64.6	STCNN
Video	DAVIS (no YouTube-VOS training)	D17 val (G)	61.7	STCNN
Video	DAVIS (no YouTube-VOS training)	D17 val (J)	58.7	STCNN
Video	DAVIS (no YouTube-VOS training)	FPS	0.26	STCNN
Video Object Segmentation	DAVIS 2017 (val)	F-measure (Mean)	64.6	Spatiotemporal CNN
Video Object Segmentation	DAVIS 2017 (val)	J&F	61.65	Spatiotemporal CNN
Video Object Segmentation	DAVIS 2017 (val)	Jaccard (Mean)	58.7	Spatiotemporal CNN
Video Object Segmentation	DAVIS 2016	F-measure (Mean)	83.8	Spatiotemporal CNN
Video Object Segmentation	DAVIS 2016	J&F	83.8	Spatiotemporal CNN
Video Object Segmentation	DAVIS 2016	Jaccard (Mean)	83.8	Spatiotemporal CNN
Video Object Segmentation	YouTube	mIoU	0.796	Spatiotemporal CNN
Video Object Segmentation	DAVIS (no YouTube-VOS training)	D16 val (F)	83.8	STCNN
Video Object Segmentation	DAVIS (no YouTube-VOS training)	D16 val (G)	83.8	STCNN
Video Object Segmentation	DAVIS (no YouTube-VOS training)	D16 val (J)	83.8	STCNN
Video Object Segmentation	DAVIS (no YouTube-VOS training)	D17 val (F)	64.6	STCNN
Video Object Segmentation	DAVIS (no YouTube-VOS training)	D17 val (G)	61.7	STCNN
Video Object Segmentation	DAVIS (no YouTube-VOS training)	D17 val (J)	58.7	STCNN
Video Object Segmentation	DAVIS (no YouTube-VOS training)	FPS	0.26	STCNN
Semi-Supervised Video Object Segmentation	DAVIS 2017 (val)	F-measure (Mean)	64.6	Spatiotemporal CNN
Semi-Supervised Video Object Segmentation	DAVIS 2017 (val)	J&F	61.65	Spatiotemporal CNN
Semi-Supervised Video Object Segmentation	DAVIS 2017 (val)	Jaccard (Mean)	58.7	Spatiotemporal CNN
Semi-Supervised Video Object Segmentation	DAVIS 2016	F-measure (Mean)	83.8	Spatiotemporal CNN
Semi-Supervised Video Object Segmentation	DAVIS 2016	J&F	83.8	Spatiotemporal CNN
Semi-Supervised Video Object Segmentation	DAVIS 2016	Jaccard (Mean)	83.8	Spatiotemporal CNN
Semi-Supervised Video Object Segmentation	YouTube	mIoU	0.796	Spatiotemporal CNN
Semi-Supervised Video Object Segmentation	DAVIS (no YouTube-VOS training)	D16 val (F)	83.8	STCNN
Semi-Supervised Video Object Segmentation	DAVIS (no YouTube-VOS training)	D16 val (G)	83.8	STCNN
Semi-Supervised Video Object Segmentation	DAVIS (no YouTube-VOS training)	D16 val (J)	83.8	STCNN
Semi-Supervised Video Object Segmentation	DAVIS (no YouTube-VOS training)	D17 val (F)	64.6	STCNN
Semi-Supervised Video Object Segmentation	DAVIS (no YouTube-VOS training)	D17 val (G)	61.7	STCNN
Semi-Supervised Video Object Segmentation	DAVIS (no YouTube-VOS training)	D17 val (J)	58.7	STCNN
Semi-Supervised Video Object Segmentation	DAVIS (no YouTube-VOS training)	FPS	0.26	STCNN

Spatiotemporal CNN for Video Object Segmentation

Abstract

Results

Related Papers

Spatiotemporal CNN for Video Object Segmentation

Abstract

Results

Related Papers