TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Papers/Towards Robust Video Object Segmentation with Adaptive Obj...

Towards Robust Video Object Segmentation with Adaptive Object Calibration

Xiaohao Xu, Jinglu Wang, Xiang Ming, Yan Lu

2022-07-02Visual Object TrackingSemi-Supervised Video Object SegmentationSegmentationSemantic SegmentationVideo SegmentationVideo Object SegmentationVideo Semantic Segmentation
PaperPDFCode(official)

Abstract

In the booming video era, video segmentation attracts increasing research attention in the multimedia community. Semi-supervised video object segmentation (VOS) aims at segmenting objects in all target frames of a video, given annotated object masks of reference frames. Most existing methods build pixel-wise reference-target correlations and then perform pixel-wise tracking to obtain target masks. Due to neglecting object-level cues, pixel-level approaches make the tracking vulnerable to perturbations, and even indiscriminate among similar objects. Towards robust VOS, the key insight is to calibrate the representation and mask of each specific object to be expressive and discriminative. Accordingly, we propose a new deep network, which can adaptively construct object representations and calibrate object masks to achieve stronger robustness. First, we construct the object representations by applying an adaptive object proxy (AOP) aggregation method, where the proxies represent arbitrary-shaped segments at multi-levels for reference. Then, prototype masks are initially generated from the reference-target correlations based on AOP. Afterwards, such proto-masks are further calibrated through network modulation, conditioning on the object proxy representations. We consolidate this conditional mask calibration process in a progressive manner, where the object representations and proto-masks evolve to be discriminative iteratively. Extensive experiments are conducted on the standard VOS benchmarks, YouTube-VOS-18/19 and DAVIS-17. Our model achieves the state-of-the-art performance among existing published works, and also exhibits superior robustness against perturbations. Our project repo is at https://github.com/JerryX1110/Robust-Video-Object-Segmentation

Results

TaskDatasetMetricValueModel
VideoDAVIS 2017F-Score85.9AOC-MF (val)
VideoDAVIS 2017Jaccard (Mean)81.7AOC-MF (val)
VideoDAVIS 2016F-Score94.7AOC-MF (val)
VideoDAVIS 2016Jaccard (Mean)88.5AOC-MF (val)
Object TrackingYouTube-VOSF-Measure (Seen)87.4AOC-MF
Object TrackingYouTube-VOSF-Measure (Unseen)87.1AOC-MF
Object TrackingYouTube-VOSJaccard (Seen)82.7AOC-MF
Object TrackingYouTube-VOSJaccard (Unseen)78.8AOC-MF
Object TrackingYouTube-VOSO (Average of Measures)84AOC-MF
Object TrackingYouTube-VOSF-Measure (Seen)87.2AOC-Base
Object TrackingYouTube-VOSF-Measure (Unseen)86.3AOC-Base
Object TrackingYouTube-VOSJaccard (Seen)82.6AOC-Base
Object TrackingYouTube-VOSJaccard (Unseen)78.3AOC-Base
Object TrackingYouTube-VOSO (Average of Measures)83.6AOC-Base
Video Object SegmentationDAVIS 2017F-Score85.9AOC-MF (val)
Video Object SegmentationDAVIS 2017Jaccard (Mean)81.7AOC-MF (val)
Video Object SegmentationDAVIS 2016F-Score94.7AOC-MF (val)
Video Object SegmentationDAVIS 2016Jaccard (Mean)88.5AOC-MF (val)
Visual Object TrackingYouTube-VOSF-Measure (Seen)87.4AOC-MF
Visual Object TrackingYouTube-VOSF-Measure (Unseen)87.1AOC-MF
Visual Object TrackingYouTube-VOSJaccard (Seen)82.7AOC-MF
Visual Object TrackingYouTube-VOSJaccard (Unseen)78.8AOC-MF
Visual Object TrackingYouTube-VOSO (Average of Measures)84AOC-MF
Visual Object TrackingYouTube-VOSF-Measure (Seen)87.2AOC-Base
Visual Object TrackingYouTube-VOSF-Measure (Unseen)86.3AOC-Base
Visual Object TrackingYouTube-VOSJaccard (Seen)82.6AOC-Base
Visual Object TrackingYouTube-VOSJaccard (Unseen)78.3AOC-Base
Visual Object TrackingYouTube-VOSO (Average of Measures)83.6AOC-Base

Related Papers

SeC: Advancing Complex Video Object Segmentation via Progressive Concept Construction2025-07-21Deep Learning-Based Fetal Lung Segmentation from Diffusion-weighted MRI Images and Lung Maturity Evaluation for Fetal Growth Restriction2025-07-17DiffOSeg: Omni Medical Image Segmentation via Multi-Expert Collaboration Diffusion Model2025-07-17From Variability To Accuracy: Conditional Bernoulli Diffusion Models with Consensus-Driven Correction for Thin Structure Segmentation2025-07-17Unleashing Vision Foundation Models for Coronary Artery Segmentation: Parallel ViT-CNN Encoding and Variational Fusion2025-07-17SCORE: Scene Context Matters in Open-Vocabulary Remote Sensing Instance Segmentation2025-07-17Unified Medical Image Segmentation with State Space Modeling Snake2025-07-17A Privacy-Preserving Semantic-Segmentation Method Using Domain-Adaptation Technique2025-07-17