Aggregated Residual Transformations for Deep Neural Networks

Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, Kaiming He

2016-11-16CVPR 2017 7Image Classification Domain Generalization General Classification

Abstract

We present a simple, highly modularized network architecture for image classification. Our network is constructed by repeating a building block that aggregates a set of transformations with the same topology. Our simple design results in a homogeneous, multi-branch architecture that has only a few hyper-parameters to set. This strategy exposes a new dimension, which we call "cardinality" (the size of the set of transformations), as an essential factor in addition to the dimensions of depth and width. On the ImageNet-1K dataset, we empirically show that even under the restricted condition of maintaining complexity, increasing cardinality is able to improve classification accuracy. Moreover, increasing cardinality is more effective than going deeper or wider when we increase the capacity. Our models, named ResNeXt, are the foundations of our entry to the ILSVRC 2016 classification task in which we secured 2nd place. We further investigate ResNeXt on an ImageNet-5K set and the COCO detection set, also showing better results than its ResNet counterpart. The code and models are publicly available online.

Results

Task	Dataset	Metric	Value	Model
Domain Adaptation	VizWiz-Classification	Accuracy - All Images	51.7	ResNeXt-101 32x16d
Domain Adaptation	VizWiz-Classification	Accuracy - Clean Images	54.8	ResNeXt-101 32x16d
Domain Adaptation	VizWiz-Classification	Accuracy - Corrupted Images	48.1	ResNeXt-101 32x16d
Image Classification	GasHisSDB	Accuracy	98.59	ResNeXt-50-32x4d
Image Classification	GasHisSDB	F1-Score	99.25	ResNeXt-50-32x4d
Image Classification	GasHisSDB	Precision	99.94	ResNeXt-50-32x4d
Image Classification	ImageNet	GFLOPs	31.5	ResNeXt-101 64x4
Image Classification	ImageNet	Top 5 Accuracy	94.7	ResNeXt-101 64x4
Domain Generalization	VizWiz-Classification	Accuracy - All Images	51.7	ResNeXt-101 32x16d
Domain Generalization	VizWiz-Classification	Accuracy - Clean Images	54.8	ResNeXt-101 32x16d
Domain Generalization	VizWiz-Classification	Accuracy - Corrupted Images	48.1	ResNeXt-101 32x16d

Aggregated Residual Transformations for Deep Neural Networks

Abstract

Results

Related Papers

Aggregated Residual Transformations for Deep Neural Networks

Abstract

Results

Related Papers