Skip to main navigation Skip to search Skip to main content

SPGrasp: Spatiotemporal Prompt-Driven Grasp Synthesis in Dynamic Scenes

  • Yunpeng Mei
  • , Hongjie Cao
  • , Wei Xiao
  • , Yinqiu Xia
  • , Zhaohan Feng
  • , Gang Wang*
  • , Jie Chen
  • *Corresponding author for this work
  • Beijing Institute of Technology
  • Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Real-time interactive grasp synthesis for dynamic objects remains challenging, as existing instance-level methods struggle to achieve low-latency inference while maintaining robust temporal consistency. To bridge this gap, we propose SPGrasp (Spatiotemporal Prompt-driven dynamic Grasp synthesis), a novel framework that extends the Segment Anything Model 2 (SAM 2) for video-stream grasp estimation. Our core innovation integrates user prompts with a spatiotemporal context module, enabling real-time interaction with end-to-end latency as low as 59 ms while preserving consistent instance identity and grasp predictions in dynamic, cluttered scenes, including object overlap and occlusion. In benchmark evaluations, SPGrasp achieves instance-level grasp accuracies of 90.6% on OCID and 93.8% on Jacquard. On the GraspNet-1Billion dataset under continuous tracking, SPGrasp reaches 92.0% accuracy with 73.1 ms per-frame latency, corresponding to a 58.5% latency reduction over the prior state-of-the-art promptable method RoG-SAM while maintaining competitive accuracy. Real-world experiments further demonstrate reliable interactive grasping under frequent occlusions, achieving a 94.8% success rate. These results suggest that SPGrasp effectively mitigates the latency-interactivity trade-off in dynamic grasp synthesis.

Original languageEnglish
Pages (from-to)202-212
Number of pages11
JournalIEEE Journal on Selected Topics in Signal Processing
Volume20
Issue number2
DOIs
Publication statusPublished - 1 Mar 2026
Externally publishedYes

Keywords

  • Dynamic grasp synthesis
  • prompt-driven grasping
  • segment anything model
  • spatiotemporal context

Fingerprint

Dive into the research topics of 'SPGrasp: Spatiotemporal Prompt-Driven Grasp Synthesis in Dynamic Scenes'. Together they form a unique fingerprint.

Cite this