SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation

Vijay Badrinarayanan, Alex Kendall, Roberto Cipolla

2015-11-02Thermal Image Segmentation Crowd Counting Real-Time Semantic Segmentation Scene Segmentation Scene Understanding Segmentation Lesion Segmentation Semantic Segmentation Medical Image Segmentation General Classification Image Segmentation

Abstract

We present a novel and practical deep fully convolutional neural network architecture for semantic pixel-wise segmentation termed SegNet. This core trainable segmentation engine consists of an encoder network, a corresponding decoder network followed by a pixel-wise classification layer. The architecture of the encoder network is topologically identical to the 13 convolutional layers in the VGG16 network. The role of the decoder network is to map the low resolution encoder feature maps to full input resolution feature maps for pixel-wise classification. The novelty of SegNet lies is in the manner in which the decoder upsamples its lower resolution input feature map(s). Specifically, the decoder uses pooling indices computed in the max-pooling step of the corresponding encoder to perform non-linear upsampling. This eliminates the need for learning to upsample. The upsampled maps are sparse and are then convolved with trainable filters to produce dense feature maps. We compare our proposed architecture with the widely adopted FCN and also with the well known DeepLab-LargeFOV, DeconvNet architectures. This comparison reveals the memory versus accuracy trade-off involved in achieving good segmentation performance. SegNet was primarily motivated by scene understanding applications. Hence, it is designed to be efficient both in terms of memory and computational time during inference. It is also significantly smaller in the number of trainable parameters than other competing architectures. We also performed a controlled benchmark of SegNet and other architectures on both road scenes and SUN RGB-D indoor scene segmentation tasks. We show that SegNet provides good performance with competitive inference time and more efficient inference memory-wise as compared to other architectures. We also provide a Caffe implementation of SegNet and a web demo at http://mi.eng.cam.ac.uk/projects/segnet/.

Results

Task	Dataset	Metric	Value	Model
Medical Image Segmentation	RITE	Dice	52.23	SegNet
Medical Image Segmentation	RITE	Jaccard Index	39.14	SegNet
Medical Image Segmentation	Anatomical Tracings of Lesions After Stroke (ATLAS)	Dice	0.2767	SegNet
Medical Image Segmentation	Anatomical Tracings of Lesions After Stroke (ATLAS)	IoU	0.1911	SegNet
Medical Image Segmentation	Anatomical Tracings of Lesions After Stroke (ATLAS)	Precision	0.3938	SegNet
Medical Image Segmentation	Anatomical Tracings of Lesions After Stroke (ATLAS)	Recall	0.2532	SegNet
Crowds	UCF-QNRF	MAE	270	Encoder-Decoder
Semantic Segmentation	TLCGIS	IoU	77.8	SegNet
Semantic Segmentation	SkyScapes-Dense	Mean IoU	23.14	SegNet
Semantic Segmentation	ADE20K	Validation mIoU	21.64	SegNet
Semantic Segmentation	SUN-RGBD	Mean IoU	31.84	SegNet
Semantic Segmentation	MFN Dataset	mIOU	42.3	SegNet
Semantic Segmentation	CamVid	Frame (fps)	4.6	SegNet
Semantic Segmentation	CamVid	Time (ms)	217	SegNet
Scene Segmentation	SUN-RGBD	Mean IoU	31.84	SegNet
Scene Segmentation	MFN Dataset	mIOU	42.3	SegNet
2D Object Detection	MFN Dataset	mIOU	42.3	SegNet
10-shot image generation	TLCGIS	IoU	77.8	SegNet
10-shot image generation	SkyScapes-Dense	Mean IoU	23.14	SegNet
10-shot image generation	ADE20K	Validation mIoU	21.64	SegNet
10-shot image generation	SUN-RGBD	Mean IoU	31.84	SegNet
10-shot image generation	MFN Dataset	mIOU	42.3	SegNet
10-shot image generation	CamVid	Frame (fps)	4.6	SegNet
10-shot image generation	CamVid	Time (ms)	217	SegNet

SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation

Abstract

Results

Related Papers

SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation

Abstract

Results

Related Papers