iCAN: Instance-Centric Attention Network for Human-Object Interaction Detection

Chen Gao, Yuliang Zou, Jia-Bin Huang

2018-08-30Human-Object Interaction Detection

Abstract

Recent years have witnessed rapid progress in detecting and recognizing individual object instances. To understand the situation in a scene, however, computers need to recognize how humans interact with surrounding objects. In this paper, we tackle the challenging task of detecting human-object interactions (HOI). Our core idea is that the appearance of a person or an object instance contains informative cues on which relevant parts of an image to attend to for facilitating interaction prediction. To exploit these cues, we propose an instance-centric attention module that learns to dynamically highlight regions in an image conditioned on the appearance of each instance. Such an attention-based network allows us to selectively aggregate features relevant for recognizing HOIs. We validate the efficacy of the proposed network on the Verb in COCO and HICO-DET datasets and show that our approach compares favorably with the state-of-the-arts.

Results

Task	Dataset	Metric	Value	Model
Human-Object Interaction Detection	V-COCO	AP(S1)	44.7	iCAN
Human-Object Interaction Detection	Ambiguious-HOI	mAP	8.14	iCAN
Human-Object Interaction Detection	HICO-DET	mAP	14.84	iCAN

Related Papers

RoHOI: Robustness Benchmark for Human-Object Interaction Detection2025-07-12 Bilateral Collaboration with Large Vision-Language Models for Open Vocabulary Human-Object Interaction Detection2025-07-09 VolumetricSMPL: A Neural Volumetric Body Model for Efficient Interactions, Contacts, and Collisions2025-06-29 HOIverse: A Synthetic Scene Graph Dataset With Human Object Interactions2025-06-24 On the Robustness of Human-Object Interaction Detection against Distribution Shift2025-06-22 Egocentric Human-Object Interaction Detection: A New Benchmark and Method2025-06-17 InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions2025-06-11 HunyuanVideo-HOMA: Generic Human-Object Interaction in Multimodal Driven Human Animation2025-06-10