Leveraging Hidden Positives for Unsupervised Semantic Segmentation

Hyun Seok Seong, WonJun Moon, SuBeen Lee, Jae-Pil Heo

2023-03-27CVPR 2023 1Unsupervised Semantic Segmentation Semantic Segmentation Contrastive Learning

Abstract

Dramatic demand for manpower to label pixel-level annotations triggered the advent of unsupervised semantic segmentation. Although the recent work employing the vision transformer (ViT) backbone shows exceptional performance, there is still a lack of consideration for task-specific training guidance and local semantic consistency. To tackle these issues, we leverage contrastive learning by excavating hidden positives to learn rich semantic relationships and ensure semantic consistency in local regions. Specifically, we first discover two types of global hidden positives, task-agnostic and task-specific ones for each anchor based on the feature similarities defined by a fixed pre-trained backbone and a segmentation head-in-training, respectively. A gradual increase in the contribution of the latter induces the model to capture task-specific semantic features. In addition, we introduce a gradient propagation strategy to learn semantic consistency between adjacent patches, under the inherent premise that nearby patches are highly likely to possess the same semantics. Specifically, we add the loss propagating to local hidden positives, semantically similar nearby patches, in proportion to the predefined similarity scores. With these training schemes, our proposed method achieves new state-of-the-art (SOTA) results in COCO-stuff, Cityscapes, and Potsdam-3 datasets. Our code is available at: https://github.com/hynnsk/HP.

Results

Task	Dataset	Metric	Value	Model
Semantic Segmentation	Potsdam-3	Accuracy	82.4	HP
Semantic Segmentation	Cityscapes test	Accuracy	80.1	HP
Semantic Segmentation	Cityscapes test	mIoU	18.4	HP
Semantic Segmentation	COCO-Stuff-27	Clustering [Accuracy]	57.2	HP (ViT-S/8)
Semantic Segmentation	COCO-Stuff-27	Clustering [mIoU]	24.6	HP (ViT-S/8)
Semantic Segmentation	COCO-Stuff-27	Linear Classifier [Accuracy]	75.6	HP (ViT-S/8)
Semantic Segmentation	COCO-Stuff-27	Linear Classifier [mIoU]	42.7	HP (ViT-S/8)
Semantic Segmentation	COCO-Stuff-27	Clustering [Accuracy]	54.5	HP (ViT-S/16)
Semantic Segmentation	COCO-Stuff-27	Clustering [mIoU]	24.3	HP (ViT-S/16)
Semantic Segmentation	COCO-Stuff-27	Linear Classifier [Accuracy]	74.1	HP (ViT-S/16)
Semantic Segmentation	COCO-Stuff-27	Linear Classifier [mIoU]	39.1	HP (ViT-S/16)
Unsupervised Semantic Segmentation	Potsdam-3	Accuracy	82.4	HP
Unsupervised Semantic Segmentation	Cityscapes test	Accuracy	80.1	HP
Unsupervised Semantic Segmentation	Cityscapes test	mIoU	18.4	HP
Unsupervised Semantic Segmentation	COCO-Stuff-27	Clustering [Accuracy]	57.2	HP (ViT-S/8)
Unsupervised Semantic Segmentation	COCO-Stuff-27	Clustering [mIoU]	24.6	HP (ViT-S/8)
Unsupervised Semantic Segmentation	COCO-Stuff-27	Linear Classifier [Accuracy]	75.6	HP (ViT-S/8)
Unsupervised Semantic Segmentation	COCO-Stuff-27	Linear Classifier [mIoU]	42.7	HP (ViT-S/8)
Unsupervised Semantic Segmentation	COCO-Stuff-27	Clustering [Accuracy]	54.5	HP (ViT-S/16)
Unsupervised Semantic Segmentation	COCO-Stuff-27	Clustering [mIoU]	24.3	HP (ViT-S/16)
Unsupervised Semantic Segmentation	COCO-Stuff-27	Linear Classifier [Accuracy]	74.1	HP (ViT-S/16)
Unsupervised Semantic Segmentation	COCO-Stuff-27	Linear Classifier [mIoU]	39.1	HP (ViT-S/16)
10-shot image generation	Potsdam-3	Accuracy	82.4	HP
10-shot image generation	Cityscapes test	Accuracy	80.1	HP
10-shot image generation	Cityscapes test	mIoU	18.4	HP
10-shot image generation	COCO-Stuff-27	Clustering [Accuracy]	57.2	HP (ViT-S/8)
10-shot image generation	COCO-Stuff-27	Clustering [mIoU]	24.6	HP (ViT-S/8)
10-shot image generation	COCO-Stuff-27	Linear Classifier [Accuracy]	75.6	HP (ViT-S/8)
10-shot image generation	COCO-Stuff-27	Linear Classifier [mIoU]	42.7	HP (ViT-S/8)
10-shot image generation	COCO-Stuff-27	Clustering [Accuracy]	54.5	HP (ViT-S/16)
10-shot image generation	COCO-Stuff-27	Clustering [mIoU]	24.3	HP (ViT-S/16)
10-shot image generation	COCO-Stuff-27	Linear Classifier [Accuracy]	74.1	HP (ViT-S/16)
10-shot image generation	COCO-Stuff-27	Linear Classifier [mIoU]	39.1	HP (ViT-S/16)

Leveraging Hidden Positives for Unsupervised Semantic Segmentation

Abstract

Results

Related Papers

Leveraging Hidden Positives for Unsupervised Semantic Segmentation

Abstract

Results

Related Papers