Referring Segmentation in Images and Videos with Cross-Modal Self-Attention Network

Linwei Ye, Mrigank Rochan, Zhi Liu, Xiaoqin Zhang, Yang Wang

2021-02-09Referring Expression Referring Expression Segmentation Segmentation Video Segmentation Video Semantic Segmentation

Paper PDF

Abstract

We consider the problem of referring segmentation in images and videos with natural language. Given an input image (or video) and a referring expression, the goal is to segment the entity referred by the expression in the image or video. In this paper, we propose a cross-modal self-attention (CMSA) module to utilize fine details of individual words and the input image or video, which effectively captures the long-range dependencies between linguistic and visual features. Our model can adaptively focus on informative words in the referring expression and important regions in the visual input. We further propose a gated multi-level fusion (GMLF) module to selectively integrate self-attentive cross-modal features corresponding to different levels of visual features. This module controls the feature fusion of information flow of features at different levels with high-level and low-level semantic information related to different attentive words. Besides, we introduce cross-frame self-attention (CFSA) module to effectively integrate temporal information in consecutive frames which extends our method in the case of referring segmentation in videos. Experiments on benchmark datasets of four referring image datasets and two actor and action video segmentation datasets consistently demonstrate that our proposed approach outperforms existing state-of-the-art methods.

Results

Task	Dataset	Metric	Value	Model
Instance Segmentation	A2D Sentences	IoU mean	0.432	CMSA+CFSA
Instance Segmentation	A2D Sentences	IoU overall	0.618	CMSA+CFSA
Instance Segmentation	A2D Sentences	Precision@0.5	0.487	CMSA+CFSA
Instance Segmentation	A2D Sentences	Precision@0.6	0.431	CMSA+CFSA
Instance Segmentation	A2D Sentences	Precision@0.7	0.358	CMSA+CFSA
Instance Segmentation	A2D Sentences	Precision@0.8	0.231	CMSA+CFSA
Instance Segmentation	A2D Sentences	Precision@0.9	0.052	CMSA+CFSA
Instance Segmentation	J-HMDB	IoU mean	0.581	CMSA+CFSA
Instance Segmentation	J-HMDB	IoU overall	0.628	CMSA+CFSA
Instance Segmentation	J-HMDB	Precision@0.5	0.764	CMSA+CFSA
Instance Segmentation	J-HMDB	Precision@0.6	0.625	CMSA+CFSA
Instance Segmentation	J-HMDB	Precision@0.7	0.389	CMSA+CFSA
Instance Segmentation	J-HMDB	Precision@0.8	0.09	CMSA+CFSA
Instance Segmentation	J-HMDB	Precision@0.9	0.001	CMSA+CFSA
Referring Expression Segmentation	A2D Sentences	IoU mean	0.432	CMSA+CFSA
Referring Expression Segmentation	A2D Sentences	IoU overall	0.618	CMSA+CFSA
Referring Expression Segmentation	A2D Sentences	Precision@0.5	0.487	CMSA+CFSA
Referring Expression Segmentation	A2D Sentences	Precision@0.6	0.431	CMSA+CFSA
Referring Expression Segmentation	A2D Sentences	Precision@0.7	0.358	CMSA+CFSA
Referring Expression Segmentation	A2D Sentences	Precision@0.8	0.231	CMSA+CFSA
Referring Expression Segmentation	A2D Sentences	Precision@0.9	0.052	CMSA+CFSA
Referring Expression Segmentation	J-HMDB	IoU mean	0.581	CMSA+CFSA
Referring Expression Segmentation	J-HMDB	IoU overall	0.628	CMSA+CFSA
Referring Expression Segmentation	J-HMDB	Precision@0.5	0.764	CMSA+CFSA
Referring Expression Segmentation	J-HMDB	Precision@0.6	0.625	CMSA+CFSA
Referring Expression Segmentation	J-HMDB	Precision@0.7	0.389	CMSA+CFSA
Referring Expression Segmentation	J-HMDB	Precision@0.8	0.09	CMSA+CFSA
Referring Expression Segmentation	J-HMDB	Precision@0.9	0.001	CMSA+CFSA

Referring Segmentation in Images and Videos with Cross-Modal Self-Attention Network

Abstract

Results

Related Papers

Referring Segmentation in Images and Videos with Cross-Modal Self-Attention Network

Abstract

Results

Related Papers