Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training

Gen Li, Nan Duan, Yuejian Fang, Ming Gong, Daxin Jiang, Ming Zhou

2019-08-16Image-text Retrieval Image-text matching Text Retrieval Masked Language Modeling Image-to-Text Retrieval Retrieval Visual Commonsense Reasoning Language Modelling

Paper PDF

Abstract

We propose Unicoder-VL, a universal encoder that aims to learn joint representations of vision and language in a pre-training manner. Borrow ideas from cross-lingual pre-trained models, such as XLM and Unicoder, both visual and linguistic contents are fed into a multi-layer Transformer for the cross-modal pre-training, where three pre-trained tasks are employed, including Masked Language Modeling (MLM), Masked Object Classification (MOC) and Visual-linguistic Matching (VLM). The first two tasks learn context-aware representations for input tokens based on linguistic and visual contents jointly. The last task tries to predict whether an image and a text describe each other. After pretraining on large-scale image-caption pairs, we transfer Unicoder-VL to caption-based image-text retrieval and visual commonsense reasoning, with just one additional output layer. We achieve state-of-the-art or comparable results on both two tasks and show the powerful ability of the cross-modal pre-training.

Results

Task	Dataset	Metric	Value	Model
Image Retrieval with Multi-Modal Query	CommercialAdsDataset	ADD(S) AUC	83.16	Unicoder-VL
Cross-Modal Information Retrieval	CommercialAdsDataset	ADD(S) AUC	83.16	Unicoder-VL
Cross-Modal Retrieval	CommercialAdsDataset	ADD(S) AUC	83.16	Unicoder-VL
Image-to-Text Retrieval	COCO (Common Objects in Context)	Recall@10	97.2	Unicoder-VL

Related Papers

Visual-Language Model Knowledge Distillation Method for Image Quality Assessment2025-07-21 From Roots to Rewards: Dynamic Tree Reasoning with RL2025-07-17 HapticCap: A Multimodal Dataset and Task for Understanding User Experience of Vibration Haptic Signals2025-07-17 A Survey of Context Engineering for Large Language Models2025-07-17 MCoT-RE: Multi-Faceted Chain-of-Thought and Re-Ranking for Training-Free Zero-Shot Composed Image Retrieval2025-07-17 Making Language Model a Hierarchical Classifier and Generator2025-07-17 VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning2025-07-17 The Generative Energy Arena (GEA): Incorporating Energy Awareness in Large Language Model (LLM) Human Evaluations2025-07-17