Text2Pos: Text-to-Point-Cloud Cross-Modal Localization

Manuel Kolmet, Qunjie Zhou, Aljosa Osep, Laura Leal-Taixe

2022-03-28CVPR 2022 1Visual Place Recognition

Abstract

Natural language-based communication with mobile devices and home appliances is becoming increasingly popular and has the potential to become natural for communicating with mobile robots in the future. Towards this goal, we investigate cross-modal text-to-point-cloud localization that will allow us to specify, for example, a vehicle pick-up or goods delivery location. In particular, we propose Text2Pos, a cross-modal localization module that learns to align textual descriptions with localization cues in a coarse- to-fine manner. Given a point cloud of the environment, Text2Pos locates a position that is specified via a natural language-based description of the immediate surroundings. To train Text2Pos and study its performance, we construct KITTI360Pose, the first dataset for this task based on the recently introduced KITTI360 dataset. Our experiments show that we can localize 65% of textual queries within 15m distance to query locations for top-10 retrieved locations. This is a starting point that we hope will spark future developments towards language-based navigation.

Results

Task	Dataset	Metric	Value	Model
Visual Place Recognition	KITTI360pose	Localization Recall@1	0.14	Text2Pos

Related Papers

Visual Place Recognition for Large-Scale UAV Applications2025-07-20 Query-Based Adaptive Aggregation for Multi-Dataset Joint Training Toward Universal Visual Place Recognition2025-07-04 Adversarial Attacks and Detection in Visual Place Recognition for Safer Robot Navigation2025-06-19 Astra: Toward General-Purpose Mobile Robots via Hierarchical Multimodal Learning2025-06-06 HypeVPR: Exploring Hyperbolic Space for Perspective to Equirectangular Visual Place Recognition2025-06-05 TAT-VPR: Ternary Adaptive Transformer for Dynamic and Efficient Visual Place Recognition2025-05-22 Place Recognition: A Comprehensive Review, Current Challenges and Future Directions2025-05-20 MMS-VPR: Multimodal Street-Level Visual Place Recognition Dataset and Benchmark2025-05-18