PIGEON: Predicting Image Geolocations

Lukas Haas, Michal Skreta, Silas Alberti, Chelsea Finn

2023-07-11CVPR 2024 1Photo geolocation estimation

Abstract

Planet-scale image geolocalization remains a challenging problem due to the diversity of images originating from anywhere in the world. Although approaches based on vision transformers have made significant progress in geolocalization accuracy, success in prior literature is constrained to narrow distributions of images of landmarks, and performance has not generalized to unseen places. We present a new geolocalization system that combines semantic geocell creation, multi-task contrastive pretraining, and a novel loss function. Additionally, our work is the first to perform retrieval over location clusters for guess refinements. We train two models for evaluations on street-level data and general-purpose image geolocalization; the first model, PIGEON, is trained on data from the game of Geoguessr and is capable of placing over 40% of its guesses within 25 kilometers of the target location globally. We also develop a bot and deploy PIGEON in a blind experiment against humans, ranking in the top 0.01% of players. We further challenge one of the world's foremost professional Geoguessr players to a series of six matches with millions of viewers, winning all six games. Our second model, PIGEOTTO, differs in that it is trained on a dataset of images from Flickr and Wikipedia, achieving state-of-the-art results on a wide range of image geolocalization benchmarks, outperforming the previous SOTA by up to 7.7 percentage points on the city accuracy level and up to 38.8 percentage points on the country level. Our findings suggest that PIGEOTTO is the first image geolocalization model that effectively generalizes to unseen places and that our approach can pave the way for highly accurate, planet-scale image geolocalization systems. Our code is available on GitHub.

Results

Task	Dataset	Metric	Value	Model
Image Classification	Im2GPS3k	City level (25 km)	36.7	PIGEOTTO
Image Classification	Im2GPS3k	Continent level (2500 km)	85.3	PIGEOTTO
Image Classification	Im2GPS3k	Country level (750 km)	72.4	PIGEOTTO
Image Classification	Im2GPS3k	Median Error (km)	147.3	PIGEOTTO
Image Classification	Im2GPS3k	Region level (200 km)	53.8	PIGEOTTO
Image Classification	Im2GPS3k	Street level (1 km)	11.3	PIGEOTTO
Image Classification	Im2GPS	City level (25 km)	40.9	PIGEOTTO
Image Classification	Im2GPS	Continent level (2500 km)	91.1	PIGEOTTO
Image Classification	Im2GPS	Country level (750 km)	82.3	PIGEOTTO
Image Classification	Im2GPS	Median Error (km)	70.5	PIGEOTTO
Image Classification	Im2GPS	Region level (200 km)	63.3	PIGEOTTO
Image Classification	Im2GPS	Street level (1 km)	14.8	PIGEOTTO
Image Classification	YFCC26k	City level (25 km)	25.8	PIGEOTTO
Image Classification	YFCC26k	Continent level (2500 km)	79	PIGEOTTO
Image Classification	YFCC26k	Country level (750 km)	63.2	PIGEOTTO
Image Classification	YFCC26k	Median Error (km)	333.3	PIGEOTTO
Image Classification	YFCC26k	Region level (200 km)	42.7	PIGEOTTO
Image Classification	YFCC26k	Street level (1 km)	10.5	PIGEOTTO
Image Classification	GWS15k	City level (25 km)	9.2	PIGEOTTO
Image Classification	GWS15k	Continent level (2500 km)	85.1	PIGEOTTO
Image Classification	GWS15k	Country level (750 km)	65.7	PIGEOTTO
Image Classification	GWS15k	Median Error (km)	415.4	PIGEOTTO
Image Classification	GWS15k	Region level (200 km)	31.2	PIGEOTTO
Image Classification	GWS15k	Street level (1 km)	0.7	PIGEOTTO
Image Classification	YFCC4k	City (25 km)	23.7	PIGEOTTO
Image Classification	YFCC4k	Continent (2500 km)	77.7	PIGEOTTO
Image Classification	YFCC4k	Country (750 km)	62.2	PIGEOTTO
Image Classification	YFCC4k	Median Error (km)	383	PIGEOTTO
Image Classification	YFCC4k	Region (200 km)	40.6	PIGEOTTO
Image Classification	YFCC4k	Street (1 km)	10.4	PIGEOTTO
4K 60Fps	Im2GPS3k	City level (25 km)	36.7	PIGEOTTO
4K 60Fps	Im2GPS3k	Continent level (2500 km)	85.3	PIGEOTTO
4K 60Fps	Im2GPS3k	Country level (750 km)	72.4	PIGEOTTO
4K 60Fps	Im2GPS3k	Median Error (km)	147.3	PIGEOTTO
4K 60Fps	Im2GPS3k	Region level (200 km)	53.8	PIGEOTTO
4K 60Fps	Im2GPS3k	Street level (1 km)	11.3	PIGEOTTO
4K 60Fps	Im2GPS	City level (25 km)	40.9	PIGEOTTO
4K 60Fps	Im2GPS	Continent level (2500 km)	91.1	PIGEOTTO
4K 60Fps	Im2GPS	Country level (750 km)	82.3	PIGEOTTO
4K 60Fps	Im2GPS	Median Error (km)	70.5	PIGEOTTO
4K 60Fps	Im2GPS	Region level (200 km)	63.3	PIGEOTTO
4K 60Fps	Im2GPS	Street level (1 km)	14.8	PIGEOTTO
4K 60Fps	YFCC26k	City level (25 km)	25.8	PIGEOTTO
4K 60Fps	YFCC26k	Continent level (2500 km)	79	PIGEOTTO
4K 60Fps	YFCC26k	Country level (750 km)	63.2	PIGEOTTO
4K 60Fps	YFCC26k	Median Error (km)	333.3	PIGEOTTO
4K 60Fps	YFCC26k	Region level (200 km)	42.7	PIGEOTTO
4K 60Fps	YFCC26k	Street level (1 km)	10.5	PIGEOTTO
4K 60Fps	GWS15k	City level (25 km)	9.2	PIGEOTTO
4K 60Fps	GWS15k	Continent level (2500 km)	85.1	PIGEOTTO
4K 60Fps	GWS15k	Country level (750 km)	65.7	PIGEOTTO
4K 60Fps	GWS15k	Median Error (km)	415.4	PIGEOTTO
4K 60Fps	GWS15k	Region level (200 km)	31.2	PIGEOTTO
4K 60Fps	GWS15k	Street level (1 km)	0.7	PIGEOTTO
4K 60Fps	YFCC4k	City (25 km)	23.7	PIGEOTTO
4K 60Fps	YFCC4k	Continent (2500 km)	77.7	PIGEOTTO
4K 60Fps	YFCC4k	Country (750 km)	62.2	PIGEOTTO
4K 60Fps	YFCC4k	Median Error (km)	383	PIGEOTTO
4K 60Fps	YFCC4k	Region (200 km)	40.6	PIGEOTTO
4K 60Fps	YFCC4k	Street (1 km)	10.4	PIGEOTTO

Abstract

Results

Task	Dataset	Metric	Value	Model
Image Classification	Im2GPS3k	City level (25 km)	36.7	PIGEOTTO
Image Classification	Im2GPS3k	Continent level (2500 km)	85.3	PIGEOTTO
Image Classification	Im2GPS3k	Country level (750 km)	72.4	PIGEOTTO
Image Classification	Im2GPS3k	Median Error (km)	147.3	PIGEOTTO
Image Classification	Im2GPS3k	Region level (200 km)	53.8	PIGEOTTO
Image Classification	Im2GPS3k	Street level (1 km)	11.3	PIGEOTTO
Image Classification	Im2GPS	City level (25 km)	40.9	PIGEOTTO
Image Classification	Im2GPS	Continent level (2500 km)	91.1	PIGEOTTO
Image Classification	Im2GPS	Country level (750 km)	82.3	PIGEOTTO
Image Classification	Im2GPS	Median Error (km)	70.5	PIGEOTTO
Image Classification	Im2GPS	Region level (200 km)	63.3	PIGEOTTO
Image Classification	Im2GPS	Street level (1 km)	14.8	PIGEOTTO
Image Classification	YFCC26k	City level (25 km)	25.8	PIGEOTTO
Image Classification	YFCC26k	Continent level (2500 km)	79	PIGEOTTO
Image Classification	YFCC26k	Country level (750 km)	63.2	PIGEOTTO
Image Classification	YFCC26k	Median Error (km)	333.3	PIGEOTTO
Image Classification	YFCC26k	Region level (200 km)	42.7	PIGEOTTO
Image Classification	YFCC26k	Street level (1 km)	10.5	PIGEOTTO
Image Classification	GWS15k	City level (25 km)	9.2	PIGEOTTO
Image Classification	GWS15k	Continent level (2500 km)	85.1	PIGEOTTO
Image Classification	GWS15k	Country level (750 km)	65.7	PIGEOTTO
Image Classification	GWS15k	Median Error (km)	415.4	PIGEOTTO
Image Classification	GWS15k	Region level (200 km)	31.2	PIGEOTTO
Image Classification	GWS15k	Street level (1 km)	0.7	PIGEOTTO
Image Classification	YFCC4k	City (25 km)	23.7	PIGEOTTO
Image Classification	YFCC4k	Continent (2500 km)	77.7	PIGEOTTO
Image Classification	YFCC4k	Country (750 km)	62.2	PIGEOTTO
Image Classification	YFCC4k	Median Error (km)	383	PIGEOTTO
Image Classification	YFCC4k	Region (200 km)	40.6	PIGEOTTO
Image Classification	YFCC4k	Street (1 km)	10.4	PIGEOTTO
4K 60Fps	Im2GPS3k	City level (25 km)	36.7	PIGEOTTO
4K 60Fps	Im2GPS3k	Continent level (2500 km)	85.3	PIGEOTTO
4K 60Fps	Im2GPS3k	Country level (750 km)	72.4	PIGEOTTO
4K 60Fps	Im2GPS3k	Median Error (km)	147.3	PIGEOTTO
4K 60Fps	Im2GPS3k	Region level (200 km)	53.8	PIGEOTTO
4K 60Fps	Im2GPS3k	Street level (1 km)	11.3	PIGEOTTO
4K 60Fps	Im2GPS	City level (25 km)	40.9	PIGEOTTO
4K 60Fps	Im2GPS	Continent level (2500 km)	91.1	PIGEOTTO
4K 60Fps	Im2GPS	Country level (750 km)	82.3	PIGEOTTO
4K 60Fps	Im2GPS	Median Error (km)	70.5	PIGEOTTO
4K 60Fps	Im2GPS	Region level (200 km)	63.3	PIGEOTTO
4K 60Fps	Im2GPS	Street level (1 km)	14.8	PIGEOTTO
4K 60Fps	YFCC26k	City level (25 km)	25.8	PIGEOTTO
4K 60Fps	YFCC26k	Continent level (2500 km)	79	PIGEOTTO
4K 60Fps	YFCC26k	Country level (750 km)	63.2	PIGEOTTO
4K 60Fps	YFCC26k	Median Error (km)	333.3	PIGEOTTO
4K 60Fps	YFCC26k	Region level (200 km)	42.7	PIGEOTTO
4K 60Fps	YFCC26k	Street level (1 km)	10.5	PIGEOTTO
4K 60Fps	GWS15k	City level (25 km)	9.2	PIGEOTTO
4K 60Fps	GWS15k	Continent level (2500 km)	85.1	PIGEOTTO
4K 60Fps	GWS15k	Country level (750 km)	65.7	PIGEOTTO
4K 60Fps	GWS15k	Median Error (km)	415.4	PIGEOTTO
4K 60Fps	GWS15k	Region level (200 km)	31.2	PIGEOTTO
4K 60Fps	GWS15k	Street level (1 km)	0.7	PIGEOTTO
4K 60Fps	YFCC4k	City (25 km)	23.7	PIGEOTTO
4K 60Fps	YFCC4k	Continent (2500 km)	77.7	PIGEOTTO
4K 60Fps	YFCC4k	Country (750 km)	62.2	PIGEOTTO
4K 60Fps	YFCC4k	Median Error (km)	383	PIGEOTTO
4K 60Fps	YFCC4k	Region (200 km)	40.6	PIGEOTTO
4K 60Fps	YFCC4k	Street (1 km)	10.4	PIGEOTTO

PIGEON: Predicting Image Geolocations

Abstract

Results

Related Papers

PIGEON: Predicting Image Geolocations

Abstract

Results

Related Papers