VPNeXt -- Rethinking Dense Decoding for Plain Vision Transformer

Xikai Tang, Ye Huang, Guangqiang Yin, Lixin Duan

2025-02-23Semantic Segmentation

Abstract

We present VPNeXt, a new and simple model for the Plain Vision Transformer (ViT). Unlike the many related studies that share the same homogeneous paradigms, VPNeXt offers a fresh perspective on dense representation based on ViT. In more detail, the proposed VPNeXt addressed two concerns about the existing paradigm: (1) Is it necessary to use a complex Transformer Mask Decoder architecture to obtain good representations? (2) Does the Plain ViT really need to depend on the mock pyramid feature for upsampling? For (1), we investigated the potential underlying reasons that contributed to the effectiveness of the Transformer Decoder and introduced the Visual Context Replay (VCR) to achieve similar effects efficiently. For (2), we introduced the ViTUp module. This module fully utilizes the previously overlooked ViT real pyramid feature to achieve better upsampling results compared to the earlier mock pyramid feature. This represents the first instance of such functionality in the field of semantic segmentation for Plain ViT. We performed ablation studies on related modules to verify their effectiveness gradually. We conducted relevant comparative experiments and visualizations to show that VPNeXt achieved state-of-the-art performance with a simple and effective design. Moreover, the proposed VPNeXt significantly exceeded the long-established mIoU wall/barrier of the VOC2012 dataset, setting a new state-of-the-art by a large margin, which also stands as the largest improvement since 2015.

Results

Task	Dataset	Metric	Value	Model
Semantic Segmentation	Cityscapes val	mIoU	84.4	VPNeXt
Semantic Segmentation	PASCAL Context	mIoU	71.1	VPNeXt
Semantic Segmentation	COCO-Stuff test	mIoU	53.7	VPNeXt
10-shot image generation	Cityscapes val	mIoU	84.4	VPNeXt
10-shot image generation	PASCAL Context	mIoU	71.1	VPNeXt
10-shot image generation	COCO-Stuff test	mIoU	53.7	VPNeXt

VPNeXt -- Rethinking Dense Decoding for Plain Vision Transformer

Abstract

Results

Related Papers

VPNeXt -- Rethinking Dense Decoding for Plain Vision Transformer

Abstract

Results

Related Papers