Action recognition with spatial-temporal discriminative filter banks

Brais Martinez, Davide Modolo, Yuanjun Xiong, Joseph Tighe

2019-08-20ICCV 2019 10Action Classification Action Recognition

Abstract

Action recognition has seen a dramatic performance improvement in the last few years. Most of the current state-of-the-art literature either aims at improving performance through changes to the backbone CNN network, or they explore different trade-offs between computational efficiency and performance, again through altering the backbone network. However, almost all of these works maintain the same last layers of the network, which simply consist of a global average pooling followed by a fully connected layer. In this work we focus on how to improve the representation capacity of the network, but rather than altering the backbone, we focus on improving the last layers of the network, where changes have low impact in terms of computational cost. In particular, we show that current architectures have poor sensitivity to finer details and we exploit recent advances in the fine-grained recognition literature to improve our model in this aspect. With the proposed approach, we obtain state-of-the-art performance on Kinetics-400 and Something-Something-V1, the two major large-scale action recognition benchmarks.

Results

Task	Dataset	Metric	Value	Model
Video	Kinetics-400	Acc@1	78.8	GB + DF + LB (ResNet 152, ImageNet pretrained)
Activity Recognition	Something-Something V1	Top 1 Accuracy	53.4	GB + DF + LB (ResNet152, ImageNet pretrained)
Action Recognition	Something-Something V1	Top 1 Accuracy	53.4	GB + DF + LB (ResNet152, ImageNet pretrained)

Related Papers

A Real-Time System for Egocentric Hand-Object Interaction Detection in Industrial Domains2025-07-17 Zero-shot Skeleton-based Action Recognition with Prototype-guided Feature Alignment2025-07-01 EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception2025-06-26 Feature Hallucination for Self-supervised Action Recognition2025-06-25 CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition2025-06-25 Including Semantic Information via Word Embeddings for Skeleton-based Action Recognition2025-06-23 Adapting Vision-Language Models for Evaluating World Models2025-06-22 Active Multimodal Distillation for Few-shot Action Recognition2025-06-16