InterMask: 3D Human Interaction Generation via Collaborative Masked Modelling

Muhammad Gohar Javed, Chuan Guo, Li Cheng, Xingyu Li

2024-10-13Motion Synthesis

Abstract

Generating realistic 3D human-human interactions from textual descriptions remains a challenging task. Existing approaches, typically based on diffusion models, often generate unnatural and unrealistic results. In this work, we introduce InterMask, a novel framework for generating human interactions using collaborative masked modeling in discrete space. InterMask first employs a VQ-VAE to transform each motion sequence into a 2D discrete motion token map. Unlike traditional 1D VQ token maps, it better preserves fine-grained spatio-temporal details and promotes spatial awareness within each token. Building on this representation, InterMask utilizes a generative masked modeling framework to collaboratively model the tokens of two interacting individuals. This is achieved by employing a transformer architecture specifically designed to capture complex spatio-temporal interdependencies. During training, it randomly masks the motion tokens of both individuals and learns to predict them. In inference, starting from fully masked sequences, it progressively fills in the tokens for both individuals. With its enhanced motion representation, dedicated architecture, and effective learning strategy, InterMask achieves state-of-the-art results, producing high-fidelity and diverse human interactions. It outperforms previous methods, achieving an FID of $5.154$ (vs $5.535$ for in2IN) on the InterHuman dataset and $0.399$ (vs $5.207$ for InterGen) on the InterX dataset. Additionally, InterMask seamlessly supports reaction generation without the need for model redesign or fine-tuning.

Results

Task	Dataset	Metric	Value	Model
Pose Tracking	Inter-X	FID	0.399	InterMask
Pose Tracking	Inter-X	MMDist	3.705	InterMask
Pose Tracking	Inter-X	MModality	2.261	InterMask
Pose Tracking	Inter-X	R-Precision Top3	0.705	InterMask
Pose Tracking	InterHuman	FID	5.154	InterMask
Pose Tracking	InterHuman	MMDist	3.79	InterMask
Pose Tracking	InterHuman	MModality	1.737	InterMask
Pose Tracking	InterHuman	R-Precision Top3	0.683	InterMask
Motion Synthesis	Inter-X	FID	0.399	InterMask
Motion Synthesis	Inter-X	MMDist	3.705	InterMask
Motion Synthesis	Inter-X	MModality	2.261	InterMask
Motion Synthesis	Inter-X	R-Precision Top3	0.705	InterMask
Motion Synthesis	InterHuman	FID	5.154	InterMask
Motion Synthesis	InterHuman	MMDist	3.79	InterMask
Motion Synthesis	InterHuman	MModality	1.737	InterMask
Motion Synthesis	InterHuman	R-Precision Top3	0.683	InterMask
10-shot image generation	Inter-X	FID	0.399	InterMask
10-shot image generation	Inter-X	MMDist	3.705	InterMask
10-shot image generation	Inter-X	MModality	2.261	InterMask
10-shot image generation	Inter-X	R-Precision Top3	0.705	InterMask
10-shot image generation	InterHuman	FID	5.154	InterMask
10-shot image generation	InterHuman	MMDist	3.79	InterMask
10-shot image generation	InterHuman	MModality	1.737	InterMask
10-shot image generation	InterHuman	R-Precision Top3	0.683	InterMask
3D Human Pose Tracking	Inter-X	FID	0.399	InterMask
3D Human Pose Tracking	Inter-X	MMDist	3.705	InterMask
3D Human Pose Tracking	Inter-X	MModality	2.261	InterMask
3D Human Pose Tracking	Inter-X	R-Precision Top3	0.705	InterMask
3D Human Pose Tracking	InterHuman	FID	5.154	InterMask
3D Human Pose Tracking	InterHuman	MMDist	3.79	InterMask
3D Human Pose Tracking	InterHuman	MModality	1.737	InterMask
3D Human Pose Tracking	InterHuman	R-Precision Top3	0.683	InterMask

Abstract

Results

Task	Dataset	Metric	Value	Model
Pose Tracking	Inter-X	FID	0.399	InterMask
Pose Tracking	Inter-X	MMDist	3.705	InterMask
Pose Tracking	Inter-X	MModality	2.261	InterMask
Pose Tracking	Inter-X	R-Precision Top3	0.705	InterMask
Pose Tracking	InterHuman	FID	5.154	InterMask
Pose Tracking	InterHuman	MMDist	3.79	InterMask
Pose Tracking	InterHuman	MModality	1.737	InterMask
Pose Tracking	InterHuman	R-Precision Top3	0.683	InterMask
Motion Synthesis	Inter-X	FID	0.399	InterMask
Motion Synthesis	Inter-X	MMDist	3.705	InterMask
Motion Synthesis	Inter-X	MModality	2.261	InterMask
Motion Synthesis	Inter-X	R-Precision Top3	0.705	InterMask
Motion Synthesis	InterHuman	FID	5.154	InterMask
Motion Synthesis	InterHuman	MMDist	3.79	InterMask
Motion Synthesis	InterHuman	MModality	1.737	InterMask
Motion Synthesis	InterHuman	R-Precision Top3	0.683	InterMask
10-shot image generation	Inter-X	FID	0.399	InterMask
10-shot image generation	Inter-X	MMDist	3.705	InterMask
10-shot image generation	Inter-X	MModality	2.261	InterMask
10-shot image generation	Inter-X	R-Precision Top3	0.705	InterMask
10-shot image generation	InterHuman	FID	5.154	InterMask
10-shot image generation	InterHuman	MMDist	3.79	InterMask
10-shot image generation	InterHuman	MModality	1.737	InterMask
10-shot image generation	InterHuman	R-Precision Top3	0.683	InterMask
3D Human Pose Tracking	Inter-X	FID	0.399	InterMask
3D Human Pose Tracking	Inter-X	MMDist	3.705	InterMask
3D Human Pose Tracking	Inter-X	MModality	2.261	InterMask
3D Human Pose Tracking	Inter-X	R-Precision Top3	0.705	InterMask
3D Human Pose Tracking	InterHuman	FID	5.154	InterMask
3D Human Pose Tracking	InterHuman	MMDist	3.79	InterMask
3D Human Pose Tracking	InterHuman	MModality	1.737	InterMask
3D Human Pose Tracking	InterHuman	R-Precision Top3	0.683	InterMask

InterMask: 3D Human Interaction Generation via Collaborative Masked Modelling

Abstract

Results

Related Papers

InterMask: 3D Human Interaction Generation via Collaborative Masked Modelling

Abstract

Results

Related Papers