XFormer: Fast and Accurate Monocular 3D Body Capture

Lihui Qian, Xintong Han, Faqiang Wang, Hongyu Liu, Haoye Dong, Zhiwen Li, Huawei Wei, Zhe Lin, Cheng-Bin Jin

2023-05-183D Human Pose Estimation

Abstract

We present XFormer, a novel human mesh and motion capture method that achieves real-time performance on consumer CPUs given only monocular images as input. The proposed network architecture contains two branches: a keypoint branch that estimates 3D human mesh vertices given 2D keypoints, and an image branch that makes predictions directly from the RGB image features. At the core of our method is a cross-modal transformer block that allows information to flow across these two branches by modeling the attention between 2D keypoint coordinates and image spatial features. Our architecture is smartly designed, which enables us to train on various types of datasets including images with 2D/3D annotations, images with 3D pseudo labels, and motion capture datasets that do not have associated images. This effectively improves the accuracy and generalization ability of our system. Built on a lightweight backbone (MobileNetV3), our method runs blazing fast (over 30fps on a single CPU core) and still yields competitive accuracy. Furthermore, with an HRNet backbone, XFormer delivers state-of-the-art performance on Huamn3.6 and 3DPW datasets.

Results

Task	Dataset	Metric	Value	Model
3D Human Pose Estimation	MPI-INF-3DHP	MPJPE	109.8	XFormer (HRNet)
3D Human Pose Estimation	MPI-INF-3DHP	PA-MPJPE	64.5	XFormer (HRNet)
3D Human Pose Estimation	3DPW	MPJPE	75	XFormer (HRNet)
3D Human Pose Estimation	3DPW	MPVPE	87.1	XFormer (HRNet)
3D Human Pose Estimation	3DPW	PA-MPJPE	45.7	XFormer (HRNet)
Pose Estimation	MPI-INF-3DHP	MPJPE	109.8	XFormer (HRNet)
Pose Estimation	MPI-INF-3DHP	PA-MPJPE	64.5	XFormer (HRNet)
Pose Estimation	3DPW	MPJPE	75	XFormer (HRNet)
Pose Estimation	3DPW	MPVPE	87.1	XFormer (HRNet)
Pose Estimation	3DPW	PA-MPJPE	45.7	XFormer (HRNet)
3D	MPI-INF-3DHP	MPJPE	109.8	XFormer (HRNet)
3D	MPI-INF-3DHP	PA-MPJPE	64.5	XFormer (HRNet)
3D	3DPW	MPJPE	75	XFormer (HRNet)
3D	3DPW	MPVPE	87.1	XFormer (HRNet)
3D	3DPW	PA-MPJPE	45.7	XFormer (HRNet)
1 Image, 2*2 Stitchi	MPI-INF-3DHP	MPJPE	109.8	XFormer (HRNet)
1 Image, 2*2 Stitchi	MPI-INF-3DHP	PA-MPJPE	64.5	XFormer (HRNet)
1 Image, 2*2 Stitchi	3DPW	MPJPE	75	XFormer (HRNet)
1 Image, 2*2 Stitchi	3DPW	MPVPE	87.1	XFormer (HRNet)
1 Image, 2*2 Stitchi	3DPW	PA-MPJPE	45.7	XFormer (HRNet)

XFormer: Fast and Accurate Monocular 3D Body Capture

Abstract

Results

Related Papers

XFormer: Fast and Accurate Monocular 3D Body Capture

Abstract

Results

Related Papers