Bridge-Prompt: Towards Ordinal Action Understanding in Instructional Videos

Muheng Li, Lei Chen, Yueqi Duan, Zhilan Hu, Jianjiang Feng, Jie zhou, Jiwen Lu

2022-03-26CVPR 2022 1Action Segmentation Human Activity Recognition Activity Recognition Action Understanding

Abstract

Action recognition models have shown a promising capability to classify human actions in short video clips. In a real scenario, multiple correlated human actions commonly occur in particular orders, forming semantically meaningful human activities. Conventional action recognition approaches focus on analyzing single actions. However, they fail to fully reason about the contextual relations between adjacent actions, which provide potential temporal logic for understanding long videos. In this paper, we propose a prompt-based framework, Bridge-Prompt (Br-Prompt), to model the semantics across adjacent actions, so that it simultaneously exploits both out-of-context and contextual information from a series of ordinal actions in instructional videos. More specifically, we reformulate the individual action labels as integrated text prompts for supervision, which bridge the gap between individual action semantics. The generated text prompts are paired with corresponding video clips, and together co-train the text encoder and the video encoder via a contrastive approach. The learned vision encoder has a stronger capability for ordinal-action-related downstream tasks, e.g. action segmentation and human activity recognition. We evaluate the performances of our approach on several video datasets: Georgia Tech Egocentric Activities (GTEA), 50Salads, and the Breakfast dataset. Br-Prompt achieves state-of-the-art on multiple benchmarks. Code is available at https://github.com/ttlmh/Bridge-Prompt

Results

Task	Dataset	Metric	Value	Model
Action Localization	50 Salads	Acc	88.1	Br-Prompt+ASFormer
Action Localization	50 Salads	Edit	83.8	Br-Prompt+ASFormer
Action Localization	50 Salads	F1@10%	89.2	Br-Prompt+ASFormer
Action Localization	50 Salads	F1@25%	87.8	Br-Prompt+ASFormer
Action Localization	50 Salads	F1@50%	81.3	Br-Prompt+ASFormer
Action Localization	GTEA	Acc	81.2	Br-Prompt+ASFormer
Action Localization	GTEA	Edit	91.6	Br-Prompt+ASFormer
Action Localization	GTEA	F1@10%	94.1	Br-Prompt+ASFormer
Action Localization	GTEA	F1@25%	92	Br-Prompt+ASFormer
Action Localization	GTEA	F1@50%	83	Br-Prompt+ASFormer
Action Segmentation	50 Salads	Acc	88.1	Br-Prompt+ASFormer
Action Segmentation	50 Salads	Edit	83.8	Br-Prompt+ASFormer
Action Segmentation	50 Salads	F1@10%	89.2	Br-Prompt+ASFormer
Action Segmentation	50 Salads	F1@25%	87.8	Br-Prompt+ASFormer
Action Segmentation	50 Salads	F1@50%	81.3	Br-Prompt+ASFormer
Action Segmentation	GTEA	Acc	81.2	Br-Prompt+ASFormer
Action Segmentation	GTEA	Edit	91.6	Br-Prompt+ASFormer
Action Segmentation	GTEA	F1@10%	94.1	Br-Prompt+ASFormer
Action Segmentation	GTEA	F1@25%	92	Br-Prompt+ASFormer
Action Segmentation	GTEA	F1@50%	83	Br-Prompt+ASFormer

Bridge-Prompt: Towards Ordinal Action Understanding in Instructional Videos

Abstract

Results

Related Papers

Bridge-Prompt: Towards Ordinal Action Understanding in Instructional Videos

Abstract

Results

Related Papers