TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Papers/Referring Segmentation in Images and Videos with Cross-Mod...

Referring Segmentation in Images and Videos with Cross-Modal Self-Attention Network

Linwei Ye, Mrigank Rochan, Zhi Liu, Xiaoqin Zhang, Yang Wang

2021-02-09Referring ExpressionReferring Expression SegmentationSegmentationVideo SegmentationVideo Semantic Segmentation
PaperPDF

Abstract

We consider the problem of referring segmentation in images and videos with natural language. Given an input image (or video) and a referring expression, the goal is to segment the entity referred by the expression in the image or video. In this paper, we propose a cross-modal self-attention (CMSA) module to utilize fine details of individual words and the input image or video, which effectively captures the long-range dependencies between linguistic and visual features. Our model can adaptively focus on informative words in the referring expression and important regions in the visual input. We further propose a gated multi-level fusion (GMLF) module to selectively integrate self-attentive cross-modal features corresponding to different levels of visual features. This module controls the feature fusion of information flow of features at different levels with high-level and low-level semantic information related to different attentive words. Besides, we introduce cross-frame self-attention (CFSA) module to effectively integrate temporal information in consecutive frames which extends our method in the case of referring segmentation in videos. Experiments on benchmark datasets of four referring image datasets and two actor and action video segmentation datasets consistently demonstrate that our proposed approach outperforms existing state-of-the-art methods.

Results

TaskDatasetMetricValueModel
Instance SegmentationA2D SentencesIoU mean0.432CMSA+CFSA
Instance SegmentationA2D SentencesIoU overall0.618CMSA+CFSA
Instance SegmentationA2D SentencesPrecision@0.50.487CMSA+CFSA
Instance SegmentationA2D SentencesPrecision@0.60.431CMSA+CFSA
Instance SegmentationA2D SentencesPrecision@0.70.358CMSA+CFSA
Instance SegmentationA2D SentencesPrecision@0.80.231CMSA+CFSA
Instance SegmentationA2D SentencesPrecision@0.90.052CMSA+CFSA
Instance SegmentationJ-HMDBIoU mean0.581CMSA+CFSA
Instance SegmentationJ-HMDBIoU overall0.628CMSA+CFSA
Instance SegmentationJ-HMDBPrecision@0.50.764CMSA+CFSA
Instance SegmentationJ-HMDBPrecision@0.60.625CMSA+CFSA
Instance SegmentationJ-HMDBPrecision@0.70.389CMSA+CFSA
Instance SegmentationJ-HMDBPrecision@0.80.09CMSA+CFSA
Instance SegmentationJ-HMDBPrecision@0.90.001CMSA+CFSA
Referring Expression SegmentationA2D SentencesIoU mean0.432CMSA+CFSA
Referring Expression SegmentationA2D SentencesIoU overall0.618CMSA+CFSA
Referring Expression SegmentationA2D SentencesPrecision@0.50.487CMSA+CFSA
Referring Expression SegmentationA2D SentencesPrecision@0.60.431CMSA+CFSA
Referring Expression SegmentationA2D SentencesPrecision@0.70.358CMSA+CFSA
Referring Expression SegmentationA2D SentencesPrecision@0.80.231CMSA+CFSA
Referring Expression SegmentationA2D SentencesPrecision@0.90.052CMSA+CFSA
Referring Expression SegmentationJ-HMDBIoU mean0.581CMSA+CFSA
Referring Expression SegmentationJ-HMDBIoU overall0.628CMSA+CFSA
Referring Expression SegmentationJ-HMDBPrecision@0.50.764CMSA+CFSA
Referring Expression SegmentationJ-HMDBPrecision@0.60.625CMSA+CFSA
Referring Expression SegmentationJ-HMDBPrecision@0.70.389CMSA+CFSA
Referring Expression SegmentationJ-HMDBPrecision@0.80.09CMSA+CFSA
Referring Expression SegmentationJ-HMDBPrecision@0.90.001CMSA+CFSA

Related Papers

SeC: Advancing Complex Video Object Segmentation via Progressive Concept Construction2025-07-21Deep Learning-Based Fetal Lung Segmentation from Diffusion-weighted MRI Images and Lung Maturity Evaluation for Fetal Growth Restriction2025-07-17DiffOSeg: Omni Medical Image Segmentation via Multi-Expert Collaboration Diffusion Model2025-07-17From Variability To Accuracy: Conditional Bernoulli Diffusion Models with Consensus-Driven Correction for Thin Structure Segmentation2025-07-17Unleashing Vision Foundation Models for Coronary Artery Segmentation: Parallel ViT-CNN Encoding and Variational Fusion2025-07-17SCORE: Scene Context Matters in Open-Vocabulary Remote Sensing Instance Segmentation2025-07-17Unified Medical Image Segmentation with State Space Modeling Snake2025-07-17A Privacy-Preserving Semantic-Segmentation Method Using Domain-Adaptation Technique2025-07-17