Deep Dual-resolution Networks for Real-time and Accurate Semantic Segmentation of Road Scenes

Yuanduo Hong, Huihui Pan, Weichao Sun, Yisong Jia

2021-01-15Autonomous Vehicles Scene Parsing All-day Semantic Segmentation Real-Time Semantic Segmentation Semantic Segmentation

Paper PDF Code Code Code Code(official)Code Code Code Code

Abstract

Semantic segmentation is a key technology for autonomous vehicles to understand the surrounding scenes. The appealing performances of contemporary models usually come at the expense of heavy computations and lengthy inference time, which is intolerable for self-driving. Using light-weight architectures (encoder-decoder or two-pathway) or reasoning on low-resolution images, recent methods realize very fast scene parsing, even running at more than 100 FPS on a single 1080Ti GPU. However, there is still a significant gap in performance between these real-time methods and the models based on dilation backbones. To tackle this problem, we proposed a family of efficient backbones specially designed for real-time semantic segmentation. The proposed deep dual-resolution networks (DDRNets) are composed of two deep branches between which multiple bilateral fusions are performed. Additionally, we design a new contextual information extractor named Deep Aggregation Pyramid Pooling Module (DAPPM) to enlarge effective receptive fields and fuse multi-scale context based on low-resolution feature maps. Our method achieves a new state-of-the-art trade-off between accuracy and speed on both Cityscapes and CamVid dataset. In particular, on a single 2080Ti GPU, DDRNet-23-slim yields 77.4% mIoU at 102 FPS on Cityscapes test set and 74.7% mIoU at 230 FPS on CamVid test set. With widely used test augmentation, our method is superior to most state-of-the-art models and requires much less computation. Codes and trained models are available online.

Results

Task	Dataset	Metric	Value	Model
Semantic Segmentation	Cityscapes test	Time (ms)	9.8	DDRNet-23-slim
Semantic Segmentation	CamVid	Time (ms)	10.6	DDRNet-23(Cityscapes-Pretrained)
Semantic Segmentation	CamVid	mIoU	80.6	DDRNet-23(Cityscapes-Pretrained)
Semantic Segmentation	CamVid	Time (ms)	4.3	DDRNet-23-slim
Semantic Segmentation	CamVid	mIoU	74.7	DDRNet-23-slim
Semantic Segmentation	Cityscapes val	Frame (fps)	37.1	DDRNet23
Semantic Segmentation	Cityscapes val	mIoU	79.4	DDRNet23
Semantic Segmentation	Cityscapes val	Frame (fps)	101.6	DDRNet23-slim
Semantic Segmentation	Cityscapes val	mIoU	77.4	DDRNet23-slim
2D Semantic Segmentation	All-day CityScapes	mIoU	68.6	DDR-Net
10-shot image generation	Cityscapes test	Time (ms)	9.8	DDRNet-23-slim
10-shot image generation	CamVid	Time (ms)	10.6	DDRNet-23(Cityscapes-Pretrained)
10-shot image generation	CamVid	mIoU	80.6	DDRNet-23(Cityscapes-Pretrained)
10-shot image generation	CamVid	Time (ms)	4.3	DDRNet-23-slim
10-shot image generation	CamVid	mIoU	74.7	DDRNet-23-slim
10-shot image generation	Cityscapes val	Frame (fps)	37.1	DDRNet23
10-shot image generation	Cityscapes val	mIoU	79.4	DDRNet23
10-shot image generation	Cityscapes val	Frame (fps)	101.6	DDRNet23-slim
10-shot image generation	Cityscapes val	mIoU	77.4	DDRNet23-slim

Deep Dual-resolution Networks for Real-time and Accurate Semantic Segmentation of Road Scenes

Abstract

Results

Related Papers

Deep Dual-resolution Networks for Real-time and Accurate Semantic Segmentation of Road Scenes

Abstract

Results

Related Papers