Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions

Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, Ling Shao

2021-02-24ICCV 2021 10Image Classification Semantic Segmentation Instance Segmentation object-detection Object Detection

Paper PDF Code Code Code Code Code Code Code Code Code(official)Code Code

Abstract

Although using convolutional neural networks (CNNs) as backbones achieves great successes in computer vision, this work investigates a simple backbone network useful for many dense prediction tasks without convolutions. Unlike the recently-proposed Transformer model (e.g., ViT) that is specially designed for image classification, we propose Pyramid Vision Transformer~(PVT), which overcomes the difficulties of porting Transformer to various dense prediction tasks. PVT has several merits compared to prior arts. (1) Different from ViT that typically has low-resolution outputs and high computational and memory cost, PVT can be not only trained on dense partitions of the image to achieve high output resolution, which is important for dense predictions but also using a progressive shrinking pyramid to reduce computations of large feature maps. (2) PVT inherits the advantages from both CNN and Transformer, making it a unified backbone in various vision tasks without convolutions by simply replacing CNN backbones. (3) We validate PVT by conducting extensive experiments, showing that it boosts the performance of many downstream tasks, e.g., object detection, semantic, and instance segmentation. For example, with a comparable number of parameters, RetinaNet+PVT achieves 40.4 AP on the COCO dataset, surpassing RetinNet+ResNet50 (36.3 AP) by 4.1 absolute AP. We hope PVT could serve as an alternative and useful backbone for pixel-level predictions and facilitate future researches. Code is available at https://github.com/whai362/PVT.

Results

Task	Dataset	Metric	Value	Model
Object Detection	COCO minival	AP50	63.6	PVT-Large (RetinaNet 3x,MS)
Object Detection	COCO minival	AP75	46.1	PVT-Large (RetinaNet 3x,MS)
Object Detection	COCO minival	APL	59.5	PVT-Large (RetinaNet 3x,MS)
Object Detection	COCO minival	APM	46	PVT-Large (RetinaNet 3x,MS)
Object Detection	COCO minival	APS	26.1	PVT-Large (RetinaNet 3x,MS)
Object Detection	COCO minival	box AP	43.4	PVT-Large (RetinaNet 3x,MS)
Object Detection	COCO minival	AP50	63.7	PVT-Large (RetinaNet 1x)
Object Detection	COCO minival	AP75	45.4	PVT-Large (RetinaNet 1x)
Object Detection	COCO minival	APL	58.4	PVT-Large (RetinaNet 1x)
Object Detection	COCO minival	APM	46	PVT-Large (RetinaNet 1x)
Object Detection	COCO minival	APS	25.8	PVT-Large (RetinaNet 1x)
Object Detection	COCO minival	box AP	42.6	PVT-Large (RetinaNet 1x)
3D	COCO minival	AP50	63.6	PVT-Large (RetinaNet 3x,MS)
3D	COCO minival	AP75	46.1	PVT-Large (RetinaNet 3x,MS)
3D	COCO minival	APL	59.5	PVT-Large (RetinaNet 3x,MS)
3D	COCO minival	APM	46	PVT-Large (RetinaNet 3x,MS)
3D	COCO minival	APS	26.1	PVT-Large (RetinaNet 3x,MS)
3D	COCO minival	box AP	43.4	PVT-Large (RetinaNet 3x,MS)
3D	COCO minival	AP50	63.7	PVT-Large (RetinaNet 1x)
3D	COCO minival	AP75	45.4	PVT-Large (RetinaNet 1x)
3D	COCO minival	APL	58.4	PVT-Large (RetinaNet 1x)
3D	COCO minival	APM	46	PVT-Large (RetinaNet 1x)
3D	COCO minival	APS	25.8	PVT-Large (RetinaNet 1x)
3D	COCO minival	box AP	42.6	PVT-Large (RetinaNet 1x)
2D Classification	COCO minival	AP50	63.6	PVT-Large (RetinaNet 3x,MS)
2D Classification	COCO minival	AP75	46.1	PVT-Large (RetinaNet 3x,MS)
2D Classification	COCO minival	APL	59.5	PVT-Large (RetinaNet 3x,MS)
2D Classification	COCO minival	APM	46	PVT-Large (RetinaNet 3x,MS)
2D Classification	COCO minival	APS	26.1	PVT-Large (RetinaNet 3x,MS)
2D Classification	COCO minival	box AP	43.4	PVT-Large (RetinaNet 3x,MS)
2D Classification	COCO minival	AP50	63.7	PVT-Large (RetinaNet 1x)
2D Classification	COCO minival	AP75	45.4	PVT-Large (RetinaNet 1x)
2D Classification	COCO minival	APL	58.4	PVT-Large (RetinaNet 1x)
2D Classification	COCO minival	APM	46	PVT-Large (RetinaNet 1x)
2D Classification	COCO minival	APS	25.8	PVT-Large (RetinaNet 1x)
2D Classification	COCO minival	box AP	42.6	PVT-Large (RetinaNet 1x)
2D Object Detection	COCO minival	AP50	63.6	PVT-Large (RetinaNet 3x,MS)
2D Object Detection	COCO minival	AP75	46.1	PVT-Large (RetinaNet 3x,MS)
2D Object Detection	COCO minival	APL	59.5	PVT-Large (RetinaNet 3x,MS)
2D Object Detection	COCO minival	APM	46	PVT-Large (RetinaNet 3x,MS)
2D Object Detection	COCO minival	APS	26.1	PVT-Large (RetinaNet 3x,MS)
2D Object Detection	COCO minival	box AP	43.4	PVT-Large (RetinaNet 3x,MS)
2D Object Detection	COCO minival	AP50	63.7	PVT-Large (RetinaNet 1x)
2D Object Detection	COCO minival	AP75	45.4	PVT-Large (RetinaNet 1x)
2D Object Detection	COCO minival	APL	58.4	PVT-Large (RetinaNet 1x)
2D Object Detection	COCO minival	APM	46	PVT-Large (RetinaNet 1x)
2D Object Detection	COCO minival	APS	25.8	PVT-Large (RetinaNet 1x)
2D Object Detection	COCO minival	box AP	42.6	PVT-Large (RetinaNet 1x)
16k	COCO minival	AP50	63.6	PVT-Large (RetinaNet 3x,MS)
16k	COCO minival	AP75	46.1	PVT-Large (RetinaNet 3x,MS)
16k	COCO minival	APL	59.5	PVT-Large (RetinaNet 3x,MS)
16k	COCO minival	APM	46	PVT-Large (RetinaNet 3x,MS)
16k	COCO minival	APS	26.1	PVT-Large (RetinaNet 3x,MS)
16k	COCO minival	box AP	43.4	PVT-Large (RetinaNet 3x,MS)
16k	COCO minival	AP50	63.7	PVT-Large (RetinaNet 1x)
16k	COCO minival	AP75	45.4	PVT-Large (RetinaNet 1x)
16k	COCO minival	APL	58.4	PVT-Large (RetinaNet 1x)
16k	COCO minival	APM	46	PVT-Large (RetinaNet 1x)
16k	COCO minival	APS	25.8	PVT-Large (RetinaNet 1x)
16k	COCO minival	box AP	42.6	PVT-Large (RetinaNet 1x)

Abstract

Results

Task	Dataset	Metric	Value	Model
Object Detection	COCO minival	AP50	63.6	PVT-Large (RetinaNet 3x,MS)
Object Detection	COCO minival	AP75	46.1	PVT-Large (RetinaNet 3x,MS)
Object Detection	COCO minival	APL	59.5	PVT-Large (RetinaNet 3x,MS)
Object Detection	COCO minival	APM	46	PVT-Large (RetinaNet 3x,MS)
Object Detection	COCO minival	APS	26.1	PVT-Large (RetinaNet 3x,MS)
Object Detection	COCO minival	box AP	43.4	PVT-Large (RetinaNet 3x,MS)
Object Detection	COCO minival	AP50	63.7	PVT-Large (RetinaNet 1x)
Object Detection	COCO minival	AP75	45.4	PVT-Large (RetinaNet 1x)
Object Detection	COCO minival	APL	58.4	PVT-Large (RetinaNet 1x)
Object Detection	COCO minival	APM	46	PVT-Large (RetinaNet 1x)
Object Detection	COCO minival	APS	25.8	PVT-Large (RetinaNet 1x)
Object Detection	COCO minival	box AP	42.6	PVT-Large (RetinaNet 1x)
3D	COCO minival	AP50	63.6	PVT-Large (RetinaNet 3x,MS)
3D	COCO minival	AP75	46.1	PVT-Large (RetinaNet 3x,MS)
3D	COCO minival	APL	59.5	PVT-Large (RetinaNet 3x,MS)
3D	COCO minival	APM	46	PVT-Large (RetinaNet 3x,MS)
3D	COCO minival	APS	26.1	PVT-Large (RetinaNet 3x,MS)
3D	COCO minival	box AP	43.4	PVT-Large (RetinaNet 3x,MS)
3D	COCO minival	AP50	63.7	PVT-Large (RetinaNet 1x)
3D	COCO minival	AP75	45.4	PVT-Large (RetinaNet 1x)
3D	COCO minival	APL	58.4	PVT-Large (RetinaNet 1x)
3D	COCO minival	APM	46	PVT-Large (RetinaNet 1x)
3D	COCO minival	APS	25.8	PVT-Large (RetinaNet 1x)
3D	COCO minival	box AP	42.6	PVT-Large (RetinaNet 1x)
2D Classification	COCO minival	AP50	63.6	PVT-Large (RetinaNet 3x,MS)
2D Classification	COCO minival	AP75	46.1	PVT-Large (RetinaNet 3x,MS)
2D Classification	COCO minival	APL	59.5	PVT-Large (RetinaNet 3x,MS)
2D Classification	COCO minival	APM	46	PVT-Large (RetinaNet 3x,MS)
2D Classification	COCO minival	APS	26.1	PVT-Large (RetinaNet 3x,MS)
2D Classification	COCO minival	box AP	43.4	PVT-Large (RetinaNet 3x,MS)
2D Classification	COCO minival	AP50	63.7	PVT-Large (RetinaNet 1x)
2D Classification	COCO minival	AP75	45.4	PVT-Large (RetinaNet 1x)
2D Classification	COCO minival	APL	58.4	PVT-Large (RetinaNet 1x)
2D Classification	COCO minival	APM	46	PVT-Large (RetinaNet 1x)
2D Classification	COCO minival	APS	25.8	PVT-Large (RetinaNet 1x)
2D Classification	COCO minival	box AP	42.6	PVT-Large (RetinaNet 1x)
2D Object Detection	COCO minival	AP50	63.6	PVT-Large (RetinaNet 3x,MS)
2D Object Detection	COCO minival	AP75	46.1	PVT-Large (RetinaNet 3x,MS)
2D Object Detection	COCO minival	APL	59.5	PVT-Large (RetinaNet 3x,MS)
2D Object Detection	COCO minival	APM	46	PVT-Large (RetinaNet 3x,MS)
2D Object Detection	COCO minival	APS	26.1	PVT-Large (RetinaNet 3x,MS)
2D Object Detection	COCO minival	box AP	43.4	PVT-Large (RetinaNet 3x,MS)
2D Object Detection	COCO minival	AP50	63.7	PVT-Large (RetinaNet 1x)
2D Object Detection	COCO minival	AP75	45.4	PVT-Large (RetinaNet 1x)
2D Object Detection	COCO minival	APL	58.4	PVT-Large (RetinaNet 1x)
2D Object Detection	COCO minival	APM	46	PVT-Large (RetinaNet 1x)
2D Object Detection	COCO minival	APS	25.8	PVT-Large (RetinaNet 1x)
2D Object Detection	COCO minival	box AP	42.6	PVT-Large (RetinaNet 1x)
16k	COCO minival	AP50	63.6	PVT-Large (RetinaNet 3x,MS)
16k	COCO minival	AP75	46.1	PVT-Large (RetinaNet 3x,MS)
16k	COCO minival	APL	59.5	PVT-Large (RetinaNet 3x,MS)
16k	COCO minival	APM	46	PVT-Large (RetinaNet 3x,MS)
16k	COCO minival	APS	26.1	PVT-Large (RetinaNet 3x,MS)
16k	COCO minival	box AP	43.4	PVT-Large (RetinaNet 3x,MS)
16k	COCO minival	AP50	63.7	PVT-Large (RetinaNet 1x)
16k	COCO minival	AP75	45.4	PVT-Large (RetinaNet 1x)
16k	COCO minival	APL	58.4	PVT-Large (RetinaNet 1x)
16k	COCO minival	APM	46	PVT-Large (RetinaNet 1x)
16k	COCO minival	APS	25.8	PVT-Large (RetinaNet 1x)
16k	COCO minival	box AP	42.6	PVT-Large (RetinaNet 1x)

Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions

Abstract

Results

Related Papers

Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions

Abstract

Results

Related Papers