Learning Position and Target Consistency for Memory-based Video Object Segmentation

Li Hu, Peng Zhang, Bang Zhang, Pan Pan, Yinghui Xu, Rong Jin

2021-04-09CVPR 2021 1Semi-Supervised Video Object Segmentation One-shot visual object segmentation Segmentation Semantic Segmentation Video Object Segmentation Video Semantic Segmentation

Paper PDF

Abstract

This paper studies the problem of semi-supervised video object segmentation(VOS). Multiple works have shown that memory-based approaches can be effective for video object segmentation. They are mostly based on pixel-level matching, both spatially and temporally. The main shortcoming of memory-based approaches is that they do not take into account the sequential order among frames and do not exploit object-level knowledge from the target. To address this limitation, we propose to Learn position and target Consistency framework for Memory-based video object segmentation, termed as LCM. It applies the memory mechanism to retrieve pixels globally, and meanwhile learns position consistency for more reliable segmentation. The learned location response promotes a better discrimination between target and distractors. Besides, LCM introduces an object-level relationship from the target to maintain target consistency, making LCM more robust to error drifting. Experiments show that our LCM achieves state-of-the-art performance on both DAVIS and Youtube-VOS benchmark. And we rank the 1st in the DAVIS 2020 challenge semi-supervised VOS task.

Results

Task	Dataset	Metric	Value	Model
Video	DAVIS (no YouTube-VOS training)	D17 val (F)	77.2	LCM
Video	DAVIS (no YouTube-VOS training)	D17 val (G)	75.2	LCM
Video	DAVIS (no YouTube-VOS training)	D17 val (J)	73.1	LCM
Video	DAVIS (no YouTube-VOS training)	FPS	8.47	LCM
Video Object Segmentation	DAVIS (no YouTube-VOS training)	D17 val (F)	77.2	LCM
Video Object Segmentation	DAVIS (no YouTube-VOS training)	D17 val (G)	75.2	LCM
Video Object Segmentation	DAVIS (no YouTube-VOS training)	D17 val (J)	73.1	LCM
Video Object Segmentation	DAVIS (no YouTube-VOS training)	FPS	8.47	LCM
Semi-Supervised Video Object Segmentation	DAVIS (no YouTube-VOS training)	D17 val (F)	77.2	LCM
Semi-Supervised Video Object Segmentation	DAVIS (no YouTube-VOS training)	D17 val (G)	75.2	LCM
Semi-Supervised Video Object Segmentation	DAVIS (no YouTube-VOS training)	D17 val (J)	73.1	LCM
Semi-Supervised Video Object Segmentation	DAVIS (no YouTube-VOS training)	FPS	8.47	LCM

Learning Position and Target Consistency for Memory-based Video Object Segmentation

Abstract

Results

Related Papers

Learning Position and Target Consistency for Memory-based Video Object Segmentation

Abstract

Results

Related Papers