Compact Trilinear Interaction for Visual Question Answering

Tuong Do, Thanh-Toan Do, Huy Tran, Erman Tjiputra, Quang D. Tran

2019-09-26ICCV 2019 10Question Answering Benchmarking Knowledge Distillation Visual Question Answering (VQA)Visual Question Answering

Paper PDF Code(official)

Abstract

In Visual Question Answering (VQA), answers have a great correlation with question meaning and visual contents. Thus, to selectively utilize image, question and answer information, we propose a novel trilinear interaction model which simultaneously learns high level associations between these three inputs. In addition, to overcome the interaction complexity, we introduce a multimodal tensor-based PARALIND decomposition which efficiently parameterizes trilinear interaction between the three inputs. Moreover, knowledge distillation is first time applied in Free-form Opened-ended VQA. It is not only for reducing the computational cost and required memory but also for transferring knowledge from trilinear interaction model to bilinear interaction model. The extensive experiments on benchmarking datasets TDIUC, VQA-2.0, and Visual7W show that the proposed compact trilinear interaction model achieves state-of-the-art results when using a single model on all three datasets.

Results

Task	Dataset	Metric	Value	Model
Visual Question Answering (VQA)	TDIUC	Accuracy	87	BAN2-CTI
Visual Question Answering (VQA)	Visual7W	Percentage correct	72.3	CTI (with Boxes)
Visual Question Answering (VQA)	VQA v2 test-dev	Accuracy	67.4	BAN2-CTI

Related Papers

Visual-Language Model Knowledge Distillation Method for Image Quality Assessment2025-07-21 Visual Place Recognition for Large-Scale UAV Applications2025-07-20 From Roots to Rewards: Dynamic Tree Reasoning with RL2025-07-17 Enter the Mind Palace: Reasoning and Planning for Long-term Active Embodied Question Answering2025-07-17 Vision-and-Language Training Helps Deploy Taxonomic Knowledge but Does Not Fundamentally Alter It2025-07-17 City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning2025-07-17 Training Transformers with Enforced Lipschitz Constants2025-07-17 Disentangling coincident cell events using deep transfer learning and compressive sensing2025-07-17