TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Papers/Fine-Grained Action Detection with RGB and Pose Informatio...

Fine-Grained Action Detection with RGB and Pose Information using Two Stream Convolutional Networks

Leonard Hacker, Finn Bartels, Pierre-Etienne Martin

2023-02-06Action DetectionAction ClassificationFine-Grained Action DetectionStroke Classification
PaperPDFCode(official)

Abstract

As participants of the MediaEval 2022 Sport Task, we propose a two-stream network approach for the classification and detection of table tennis strokes. Each stream is a succession of 3D Convolutional Neural Network (CNN) blocks using attention mechanisms. Each stream processes different 4D inputs. Our method utilizes raw RGB data and pose information computed from MMPose toolbox. The pose information is treated as an image by applying the pose either on a black background or on the original RGB frame it has been computed from. Best performance is obtained by feeding raw RGB data to one stream, Pose + RGB (PRGB) information to the other stream and applying late fusion on the features. The approaches were evaluated on the provided TTStroke-21 data sets. We can report an improvement in stroke classification, reaching 87.3% of accuracy, while the detection does not outperform the baseline but still reaches an IoU of 0.349 and mAP of 0.110.

Results

TaskDatasetMetricValueModel
VideoTTStroke-21 ME22Acc0.8731RGB and PRGB
Action DetectionTTStroke-21 ME22IoU0.3491RGB and PRGB
Action DetectionTTStroke-21 ME22mAP0.1101RGB and PRGB

Related Papers

CBF-AFA: Chunk-Based Multi-SSL Fusion for Automatic Fluency Assessment2025-06-25MultiHuman-Testbench: Benchmarking Image Generation for Multiple Humans2025-06-25Brain Stroke Classification Using Wavelet Transform and MLP Neural Networks on DWI MRI Images2025-06-18Distributed Activity Detection for Cell-Free Hybrid Near-Far Field Communications2025-06-17SurgBench: A Unified Large-Scale Benchmark for Surgical Video Analysis2025-06-09From Play to Replay: Composed Video Retrieval for Temporally Fine-Grained Videos2025-06-05Speaker Diarization with Overlapping Community Detection Using Graph Attention Networks and Label Propagation Algorithm2025-06-03Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion2025-06-02