Vision Models Are More Robust And Fair When Pretrained On Uncurated Images Without Supervision

Priya Goyal, Quentin Duval, Isaac Seessel, Mathilde Caron, Ishan Misra, Levent Sagun, Armand Joulin, Piotr Bojanowski

2022-02-16Fairness Self-Supervised Image Classification Image Classification Action Classification Self-Supervised Learning Traffic Sign Recognition Domain Generalization Copy Detection Word Embeddings Action Recognition Meme Classification Out-of-Distribution Generalization Fine-Grained Image Classification Semi-Supervised Image Classification

Paper PDF Code(official)

Abstract

Discriminative self-supervised learning allows training models on any random group of internet images, and possibly recover salient information that helps differentiate between the images. Applied to ImageNet, this leads to object centric features that perform on par with supervised features on most object-centric downstream tasks. In this work, we question if using this ability, we can learn any salient and more representative information present in diverse unbounded set of images from across the globe. To do so, we train models on billions of random images without any data pre-processing or prior assumptions about what we want the model to learn. We scale our model size to dense 10 billion parameters to avoid underfitting on a large data size. We extensively study and validate our model performance on over 50 benchmarks including fairness, robustness to distribution shift, geographical diversity, fine grained recognition, image copy detection and many image classification datasets. The resulting model, not only captures well semantic information, it also captures information about artistic style and learns salient information such as geolocations and multilingual word embeddings based on visual content only. More importantly, we discover that such model is more robust, more fair, less harmful and less biased than supervised models or models trained on object centric datasets such as ImageNet.

Results

Task	Dataset	Metric	Value	Model
Domain Adaptation	ImageNet-R	Top-1 Error Rate	43.9	SEER (RegNet10B)
Domain Adaptation	ImageNet-A	Top-1 accuracy %	52.7	SEER (RegNet10B)
Domain Adaptation	ImageNet-Sketch	Top-1 accuracy	45.6	SEER (RegNet10B)
Video	Kinetics-700	Top-1 Accuracy	51.9	SEER (RegNet10B)
Image Classification	KITTI-Dist	Top 1 Accuracy	78.34	SEER (RegNet10B)
Image Classification	Places205	Top 1 Accuracy	69	SEER (RegNet10B - finetuned - 384px)
Image Classification	ImageNet V2	Top 1 Accuracy	76.2	SEER (RegNet10B)
Image Classification	DTD	Accuracy	80.5	SEER (RegNet10B - linear eval)
Image Classification	CLEVR/Count	Top 1 Accuracy	89.28	SEER (RegNet10B)
Image Classification	CLEVR/Count	Top 1 Accuracy	87.98	SEER (RegNetY-128GF)
Image Classification	ObjectNet	Top-1 Accuracy	60.2	SEER (RegNet10B)
Image Classification	RESISC45	Top 1 Accuracy	95.61	SEER (RegNet10B)
Image Classification	RESISC45	Top 1 Accuracy	94.73	SwAV (ResNet50-w5)
Image Classification	RESISC45	Top 1 Accuracy	93.97	DINO (DeiT-B/16)
Image Classification	RESISC45	Top 1 Accuracy	93.35	MoCo-v3 (ViT-B/16)
Image Classification	RESISC45	Top 1 Accuracy	92.7	CLIP (ViT-B/16)
Image Classification	RESISC45	Top 1 Accuracy	92.48	DeiT-B/16
Image Classification	RESISC45	Top 1 Accuracy	89.77	SimCLR-v2 (ResNet152-w3 + SK)
Image Classification	RESISC45	Top 1 Accuracy	88.56	ResNet50 (ImageNet-supervised)
Image Classification	RESISC45	Top 1 Accuracy	85.4	MoCo-v2 (ResNet50)
Image Classification	CIFAR-10	Percentage correct	90	SEER (RegNet10B)
Image Classification	Flowers-102	Accuracy	96.3	SEER (RegNet10B)
Image Classification	CIFAR-100	Percentage correct	81.53	SEER (RegNet10B)
Image Classification	MNIST	Accuracy	99.42	SEER (RegNet10B)
Image Classification	MNIST	Percentage error	0.58	SEER (RegNet10B)
Image Classification	CLEVR/Dist	Top 1 Accuracy	74.98	SEER (RegNet10B)
Image Classification	CLEVR/Dist	Top 1 Accuracy	72.67	SEER (RegNetY-128GF)
Image Classification	STL-10	Percentage correct	97.3	SEER (RegNet10B)
Image Classification	Food-101	Accuracy (%)	90.3	SEER (RegNet10B - linear eval)
Image Classification	EuroSAT	Accuracy (%)	97.5	SEER (RegNet10B - linear eval)
Image Classification	SVHN	Percentage error	13.6	SEER (RegNet10B)
Image Classification	Caltech-101	Accuracy	91	SEER (RegNet10B - linear eval)
Image Classification	SUN397	Accuracy	80	SEER (RegNet10B - linear eval)
Fine-Grained Image Classification	Caltech-101	Accuracy	91	SEER (RegNet10B - linear eval)
Fine-Grained Image Classification	SUN397	Accuracy	80	SEER (RegNet10B - linear eval)
Meme Classification	Hateful Memes	ROC-AUC	0.734	SEER (RegNet10B)
Domain Generalization	ImageNet-R	Top-1 Error Rate	43.9	SEER (RegNet10B)
Domain Generalization	ImageNet-A	Top-1 accuracy %	52.7	SEER (RegNet10B)
Domain Generalization	ImageNet-Sketch	Top-1 accuracy	45.6	SEER (RegNet10B)

Vision Models Are More Robust And Fair When Pretrained On Uncurated Images Without Supervision

Abstract

Results

Related Papers

Vision Models Are More Robust And Fair When Pretrained On Uncurated Images Without Supervision

Abstract

Results

Related Papers