Cross-modulated Attention Transformer for RGBT Tracking

Yun Xiao, jiacong Zhao, Andong Lu, Chenglong Li, Yin Lin, Bing Yin, Cong Liu

2024-08-05Rgb-T Tracking

Abstract

Existing Transformer-based RGBT trackers achieve remarkable performance benefits by leveraging self-attention to extract uni-modal features and cross-attention to enhance multi-modal feature interaction and template-search correlation computation. Nevertheless, the independent search-template correlation calculations ignore the consistency between branches, which can result in ambiguous and inappropriate correlation weights. It not only limits the intra-modal feature representation, but also harms the robustness of cross-attention for multi-modal feature interaction and search-template correlation computation. To address these issues, we propose a novel approach called Cross-modulated Attention Transformer (CAFormer), which performs intra-modality self-correlation, inter-modality feature interaction, and search-template correlation computation in a unified attention model, for RGBT tracking. In particular, we first independently generate correlation maps for each modality and feed them into the designed Correlation Modulated Enhancement module, modulating inaccurate correlation weights by seeking the consensus between modalities. Such kind of design unifies self-attention and cross-attention schemes, which not only alleviates inaccurate attention weight computation in self-attention but also eliminates redundant computation introduced by extra cross-attention scheme. In addition, we propose a collaborative token elimination strategy to further improve tracking inference efficiency and accuracy. Extensive experiments on five public RGBT tracking benchmarks show the outstanding performance of the proposed CAFormer against state-of-the-art methods.

Results

Task	Dataset	Metric	Value	Model
Visual Tracking	LasHeR	Precision	70	CAFormer
Visual Tracking	LasHeR	Success	55.6	CAFormer
Visual Tracking	GTOT	Precision	91.8	CAFormer
Visual Tracking	GTOT	Success	76.9	CAFormer
Visual Tracking	RGBT234	Precision	88.3	CAFormer
Visual Tracking	RGBT234	Success	66.4	CAFormer
Visual Tracking	RGBT210	Precision	85.6	CAFormer
Visual Tracking	RGBT210	Success	63.2	CAFormer

Abstract

Results

Task	Dataset	Metric	Value	Model
Visual Tracking	LasHeR	Precision	70	CAFormer
Visual Tracking	LasHeR	Success	55.6	CAFormer
Visual Tracking	GTOT	Precision	91.8	CAFormer
Visual Tracking	GTOT	Success	76.9	CAFormer
Visual Tracking	RGBT234	Precision	88.3	CAFormer
Visual Tracking	RGBT234	Success	66.4	CAFormer
Visual Tracking	RGBT210	Precision	85.6	CAFormer
Visual Tracking	RGBT210	Success	63.2	CAFormer

Cross-modulated Attention Transformer for RGBT Tracking

Abstract

Results

Related Papers

Cross-modulated Attention Transformer for RGBT Tracking

Abstract

Results

Related Papers