MM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of High-Intensity Surgical Environments

Ege Özsoy, Chantal Pellegrini, Tobias Czempiel, Felix Tristram, Kun Yuan, David Bani-Harouni, Ulrich Eck, Benjamin Busam, Matthias Keicher, Nassir Navab

2025-03-04CVPR 2025 1Scene Graph Generation Video Panoptic Segmentation 2D Panoptic Segmentation Graph Generation Language Modelling

Paper PDF Code(official)

Abstract

Operating rooms (ORs) are complex, high-stakes environments requiring precise understanding of interactions among medical staff, tools, and equipment for enhancing surgical assistance, situational awareness, and patient safety. Current datasets fall short in scale, realism and do not capture the multimodal nature of OR scenes, limiting progress in OR modeling. To this end, we introduce MM-OR, a realistic and large-scale multimodal spatiotemporal OR dataset, and the first dataset to enable multimodal scene graph generation. MM-OR captures comprehensive OR scenes containing RGB-D data, detail views, audio, speech transcripts, robotic logs, and tracking data and is annotated with panoptic segmentations, semantic scene graphs, and downstream task labels. Further, we propose MM2SG, the first multimodal large vision-language model for scene graph generation, and through extensive experiments, demonstrate its ability to effectively leverage multimodal inputs. Together, MM-OR and MM2SG establish a new benchmark for holistic OR understanding, and open the path towards multimodal scene analysis in complex, high-stakes environments. Our code, and data is available at https://github.com/egeozsoy/MM-OR.

Results

Task	Dataset	Metric	Value	Model
Scene Parsing	4D-OR	F1	0.901	MM2SG
Scene Parsing	MM-OR	Macro F1	0.529	MM2SG
Semantic Segmentation	4D-OR	VPQ	69.8	MM-OR-VPQ4
Semantic Segmentation	4D-OR	VPQ	69.2	MM-OR-VPQ8
Semantic Segmentation	MM-OR	VPQ	67	MM-OR-VPQ4
Semantic Segmentation	MM-OR	VPQ	66.4	MM-OR-VPQ8
2D Semantic Segmentation	4D-OR	F1	0.901	MM2SG
2D Semantic Segmentation	MM-OR	Macro F1	0.529	MM2SG
Scene Graph Generation	4D-OR	F1	0.901	MM2SG
Scene Graph Generation	MM-OR	Macro F1	0.529	MM2SG
10-shot image generation	4D-OR	VPQ	69.8	MM-OR-VPQ4
10-shot image generation	4D-OR	VPQ	69.2	MM-OR-VPQ8
10-shot image generation	MM-OR	VPQ	67	MM-OR-VPQ4
10-shot image generation	MM-OR	VPQ	66.4	MM-OR-VPQ8
Panoptic Segmentation	4D-OR	VPQ	69.8	MM-OR-VPQ4
Panoptic Segmentation	4D-OR	VPQ	69.2	MM-OR-VPQ8
Panoptic Segmentation	MM-OR	VPQ	67	MM-OR-VPQ4
Panoptic Segmentation	MM-OR	VPQ	66.4	MM-OR-VPQ8
2D Panoptic Segmentation	MM-OR	VPQ	67.5	MM-OR
2D Panoptic Segmentation	4D-OR	VPQ	71.8	MM-OR

MM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of High-Intensity Surgical Environments

Abstract

Results

Related Papers

MM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of High-Intensity Surgical Environments

Abstract

Results

Related Papers