TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Papers/Recurrent Vision Transformers for Object Detection with Ev...

Recurrent Vision Transformers for Object Detection with Event Cameras

Mathias Gehrig, Davide Scaramuzza

2022-12-11CVPR 2023 1Event-based visionobject-detectionObject Detection
PaperPDFCode(official)

Abstract

We present Recurrent Vision Transformers (RVTs), a novel backbone for object detection with event cameras. Event cameras provide visual information with sub-millisecond latency at a high-dynamic range and with strong robustness against motion blur. These unique properties offer great potential for low-latency object detection and tracking in time-critical scenarios. Prior work in event-based vision has achieved outstanding detection performance but at the cost of substantial inference time, typically beyond 40 milliseconds. By revisiting the high-level design of recurrent vision backbones, we reduce inference time by a factor of 6 while retaining similar performance. To achieve this, we explore a multi-stage design that utilizes three key concepts in each stage: First, a convolutional prior that can be regarded as a conditional positional embedding. Second, local and dilated global self-attention for spatial feature interaction. Third, recurrent temporal feature aggregation to minimize latency while retaining temporal information. RVTs can be trained from scratch to reach state-of-the-art performance on event-based object detection - achieving an mAP of 47.2% on the Gen1 automotive dataset. At the same time, RVTs offer fast inference (<12 ms on a T4 GPU) and favorable parameter efficiency (5 times fewer than prior art). Our study brings new insights into effective design choices that can be fruitful for research beyond event-based vision.

Results

TaskDatasetMetricValueModel
Object DetectionGEN1 DetectionParams18.5RVT-B
Object DetectionGEN1 DetectionmAP47.2RVT-B
Object DetectionGEN1 DetectionParams9.9RVT-S
Object DetectionGEN1 DetectionmAP46.5RVT-S
Object DetectionGEN1 DetectionParams4.4RVT-T
Object DetectionGEN1 DetectionmAP44.1RVT-T
3DGEN1 DetectionParams18.5RVT-B
3DGEN1 DetectionmAP47.2RVT-B
3DGEN1 DetectionParams9.9RVT-S
3DGEN1 DetectionmAP46.5RVT-S
3DGEN1 DetectionParams4.4RVT-T
3DGEN1 DetectionmAP44.1RVT-T
2D ClassificationGEN1 DetectionParams18.5RVT-B
2D ClassificationGEN1 DetectionmAP47.2RVT-B
2D ClassificationGEN1 DetectionParams9.9RVT-S
2D ClassificationGEN1 DetectionmAP46.5RVT-S
2D ClassificationGEN1 DetectionParams4.4RVT-T
2D ClassificationGEN1 DetectionmAP44.1RVT-T
2D Object DetectionGEN1 DetectionParams18.5RVT-B
2D Object DetectionGEN1 DetectionmAP47.2RVT-B
2D Object DetectionGEN1 DetectionParams9.9RVT-S
2D Object DetectionGEN1 DetectionmAP46.5RVT-S
2D Object DetectionGEN1 DetectionParams4.4RVT-T
2D Object DetectionGEN1 DetectionmAP44.1RVT-T
16kGEN1 DetectionParams18.5RVT-B
16kGEN1 DetectionmAP47.2RVT-B
16kGEN1 DetectionParams9.9RVT-S
16kGEN1 DetectionmAP46.5RVT-S
16kGEN1 DetectionParams4.4RVT-T
16kGEN1 DetectionmAP44.1RVT-T

Related Papers

A Real-Time System for Egocentric Hand-Object Interaction Detection in Industrial Domains2025-07-17RS-TinyNet: Stage-wise Feature Fusion Network for Detecting Tiny Objects in Remote Sensing Images2025-07-17Decoupled PROB: Decoupled Query Initialization Tasks and Objectness-Class Learning for Open World Object Detection2025-07-17Dual LiDAR-Based Traffic Movement Count Estimation at a Signalized Intersection: Deployment, Data Collection, and Preliminary Analysis2025-07-17Vision-based Perception for Autonomous Vehicles in Obstacle Avoidance Scenarios2025-07-16Tomato Multi-Angle Multi-Pose Dataset for Fine-Grained Phenotyping2025-07-15ECORE: Energy-Conscious Optimized Routing for Deep Learning Models at the Edge2025-07-08Beyond One Shot, Beyond One Perspective: Cross-View and Long-Horizon Distillation for Better LiDAR Representations2025-07-07