Seyed Hamid Reza Roodabeh, Zongyu Li, Homa Alemzadeh
A unified multimodal framework is proposed for real-time error detection in robot-assisted surgery by integrating video, kinematics, and descriptive textual prompts.
Robot-assisted minimally invasive surgery increases complexity, making technical error detection crucial for patient safety. Current video-based methods often overlook fine-grained contextual descriptions of activities and errors within the hierarchical structure of surgical procedures and under-utilize complementary multimodal information.
The framework takes video, kinematics, and descriptive textual prompts as input. It integrates descriptive language for gesture-level activities, instrument-object interactions, and error definitions through 'activity prompting'. It also introduces activity-aware visual embeddings derived from vision encoders pretrained on surgical activity labels to compare contrastive language-image embeddings with traditional image-based embeddings.
The framework achieves up to 5% and 16.6% F1 score improvements over state-of-the-art baselines on the JIGSAWS and SAR-RARP50 datasets, respectively. This demonstrates the significant value of combining curated textual prompts with multimodal data for accurate error detection.