Unsupervised Semantic Segmentation by Distilling Feature Correspondences

Mark Hamilton, Zhoutong Zhang, Bharath Hariharan, Noah Snavely, William T. Freeman

2022-03-16ICLR 2022 4Unsupervised Semantic Segmentation Form Semantic Segmentation

Abstract

Unsupervised semantic segmentation aims to discover and localize semantically meaningful categories within image corpora without any form of annotation. To solve this task, algorithms must produce features for every pixel that are both semantically meaningful and compact enough to form distinct clusters. Unlike previous works which achieve this with a single end-to-end framework, we propose to separate feature learning from cluster compactification. Empirically, we show that current unsupervised feature learning frameworks already generate dense features whose correlations are semantically consistent. This observation motivates us to design STEGO ($\textbf{S}$elf-supervised $\textbf{T}$ransformer with $\textbf{E}$nergy-based $\textbf{G}$raph $\textbf{O}$ptimization), a novel framework that distills unsupervised features into high-quality discrete semantic labels. At the core of STEGO is a novel contrastive loss function that encourages features to form compact clusters while preserving their relationships across the corpora. STEGO yields a significant improvement over the prior state of the art, on both the CocoStuff ($\textbf{+14 mIoU}$) and Cityscapes ($\textbf{+9 mIoU}$) semantic segmentation challenges.

Results

Task	Dataset	Metric	Value	Model
Semantic Segmentation	Potsdam-3	Accuracy	77	STEGO
Semantic Segmentation	Cityscapes test	Accuracy	73.2	STEGO
Semantic Segmentation	Cityscapes test	mIoU	21	STEGO
Semantic Segmentation	COCO-Stuff-27	Clustering [Accuracy]	56.9	STEGO (ViT-B/8)
Semantic Segmentation	COCO-Stuff-27	Clustering [mIoU]	28.2	STEGO (ViT-B/8)
Semantic Segmentation	COCO-Stuff-27	Clustering [mIoU]	24.5	STEGO (ViT-S/8)
Semantic Segmentation	COCO-Stuff-27	Linear Classifier [Accuracy]	74.4	STEGO (ViT-S/8)
Semantic Segmentation	COCO-Stuff-27	Linear Classifier [mIoU]	38.3	STEGO (ViT-S/8)
Unsupervised Semantic Segmentation	Potsdam-3	Accuracy	77	STEGO
Unsupervised Semantic Segmentation	Cityscapes test	Accuracy	73.2	STEGO
Unsupervised Semantic Segmentation	Cityscapes test	mIoU	21	STEGO
Unsupervised Semantic Segmentation	COCO-Stuff-27	Clustering [Accuracy]	56.9	STEGO (ViT-B/8)
Unsupervised Semantic Segmentation	COCO-Stuff-27	Clustering [mIoU]	28.2	STEGO (ViT-B/8)
Unsupervised Semantic Segmentation	COCO-Stuff-27	Clustering [mIoU]	24.5	STEGO (ViT-S/8)
Unsupervised Semantic Segmentation	COCO-Stuff-27	Linear Classifier [Accuracy]	74.4	STEGO (ViT-S/8)
Unsupervised Semantic Segmentation	COCO-Stuff-27	Linear Classifier [mIoU]	38.3	STEGO (ViT-S/8)
10-shot image generation	Potsdam-3	Accuracy	77	STEGO
10-shot image generation	Cityscapes test	Accuracy	73.2	STEGO
10-shot image generation	Cityscapes test	mIoU	21	STEGO
10-shot image generation	COCO-Stuff-27	Clustering [Accuracy]	56.9	STEGO (ViT-B/8)
10-shot image generation	COCO-Stuff-27	Clustering [mIoU]	28.2	STEGO (ViT-B/8)
10-shot image generation	COCO-Stuff-27	Clustering [mIoU]	24.5	STEGO (ViT-S/8)
10-shot image generation	COCO-Stuff-27	Linear Classifier [Accuracy]	74.4	STEGO (ViT-S/8)
10-shot image generation	COCO-Stuff-27	Linear Classifier [mIoU]	38.3	STEGO (ViT-S/8)

Unsupervised Semantic Segmentation by Distilling Feature Correspondences

Abstract

Results

Related Papers

Unsupervised Semantic Segmentation by Distilling Feature Correspondences

Abstract

Results

Related Papers