End-to-end Dense Video Captioning as Sequence Generation

Wanrong Zhu, Bo Pang, Ashish V. Thapliyal, William Yang Wang, Radu Soricut

2022-04-18COLING 2022 10Descriptive Video Captioning Dense Video Captioning

Abstract

Dense video captioning aims to identify the events of interest in an input video, and generate descriptive captions for each event. Previous approaches usually follow a two-stage generative process, which first proposes a segment for each event, then renders a caption for each identified segment. Recent advances in large-scale sequence generation pretraining have seen great success in unifying task formulation for a great variety of tasks, but so far, more complex tasks such as dense video captioning are not able to fully utilize this powerful paradigm. In this work, we show how to model the two subtasks of dense video captioning jointly as one sequence generation task, and simultaneously predict the events and the corresponding descriptions. Experiments on YouCook2 and ViTT show encouraging results and indicate the feasibility of training complex tasks such as end-to-end dense video captioning integrated into large-scale pretrained models.

Results

Task	Dataset	Metric	Value	Model
Video Captioning	ViTT	CIDEr	25	E2ESG
Video Captioning	ViTT	METEOR	8.1	E2ESG
Dense Video Captioning	ViTT	CIDEr	25	E2ESG
Dense Video Captioning	ViTT	METEOR	8.1	E2ESG

Related Papers

DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization2025-07-17 Assay2Mol: large language model-based drug design using BioAssay context2025-07-16 Describe Anything Model for Visual Question Answering on Text-rich Images2025-07-16 UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks2025-07-15 FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation2025-07-09 Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor2025-07-04 Prompt Disentanglement via Language Guidance and Representation Alignment for Domain Generalization2025-07-03 Dataset Distillation via Vision-Language Category Prototype2025-06-30