Skeleton-Based Action Recognition with Spatial Reasoning and Temporal Stack Learning

Chenyang Si, Ya Jing, Wei Wang, Liang Wang, Tieniu Tan

2018-05-07ECCV 2018 9Spatial Reasoning Skeleton Based Action Recognition Human-Object Interaction Detection Action Recognition Temporal Action Localization

Paper PDF

Abstract

Skeleton-based action recognition has made great progress recently, but many problems still remain unsolved. For example, most of the previous methods model the representations of skeleton sequences without abundant spatial structure information and detailed temporal dynamics features. In this paper, we propose a novel model with spatial reasoning and temporal stack learning (SR-TSL) for skeleton based action recognition, which consists of a spatial reasoning network (SRN) and a temporal stack learning network (TSLN). The SRN can capture the high-level spatial structural information within each frame by a residual graph neural network, while the TSLN can model the detailed temporal dynamics of skeleton sequences by a composition of multiple skip-clip LSTMs. During training, we propose a clip-based incremental loss to optimize the model. We perform extensive experiments on the SYSU 3D Human-Object Interaction dataset and NTU RGB+D dataset and verify the effectiveness of each network of our model. The comparison results illustrate that our approach achieves much better results than state-of-the-art methods.

Results

Task	Dataset	Metric	Value	Model
Video	NTU RGB+D	Accuracy (CS)	84.8	SR-TSL
Video	NTU RGB+D	Accuracy (CV)	92.4	SR-TSL
Temporal Action Localization	NTU RGB+D	Accuracy (CS)	84.8	SR-TSL
Temporal Action Localization	NTU RGB+D	Accuracy (CV)	92.4	SR-TSL
Zero-Shot Learning	NTU RGB+D	Accuracy (CS)	84.8	SR-TSL
Zero-Shot Learning	NTU RGB+D	Accuracy (CV)	92.4	SR-TSL
Activity Recognition	NTU RGB+D	Accuracy (CS)	84.8	SR-TSL
Activity Recognition	NTU RGB+D	Accuracy (CV)	92.4	SR-TSL
Action Localization	NTU RGB+D	Accuracy (CS)	84.8	SR-TSL
Action Localization	NTU RGB+D	Accuracy (CV)	92.4	SR-TSL
Action Detection	NTU RGB+D	Accuracy (CS)	84.8	SR-TSL
Action Detection	NTU RGB+D	Accuracy (CV)	92.4	SR-TSL
3D Action Recognition	NTU RGB+D	Accuracy (CS)	84.8	SR-TSL
3D Action Recognition	NTU RGB+D	Accuracy (CV)	92.4	SR-TSL
Action Recognition	NTU RGB+D	Accuracy (CS)	84.8	SR-TSL
Action Recognition	NTU RGB+D	Accuracy (CV)	92.4	SR-TSL

Skeleton-Based Action Recognition with Spatial Reasoning and Temporal Stack Learning

Abstract

Results

Related Papers

Skeleton-Based Action Recognition with Spatial Reasoning and Temporal Stack Learning

Abstract

Results

Related Papers