Abstract
Real-time interactive grasp synthesis for dynamic objects remains challenging, as existing instance-level methods struggle to achieve low-latency inference while maintaining robust temporal consistency. To bridge this gap, we propose SPGrasp (Spatiotemporal Prompt-driven dynamic Grasp synthesis), a novel framework that extends the Segment Anything Model 2 (SAM 2) for video-stream grasp estimation. Our core innovation integrates user prompts with a spatiotemporal context module, enabling real-time interaction with end-to-end latency as low as 59 ms while preserving consistent instance identity and grasp predictions in dynamic, cluttered scenes, including object overlap and occlusion. In benchmark evaluations, SPGrasp achieves instance-level grasp accuracies of 90.6% on OCID and 93.8% on Jacquard. On the GraspNet-1Billion dataset under continuous tracking, SPGrasp reaches 92.0% accuracy with 73.1 ms per-frame latency, corresponding to a 58.5% latency reduction over the prior state-of-the-art promptable method RoG-SAM while maintaining competitive accuracy. Real-world experiments further demonstrate reliable interactive grasping under frequent occlusions, achieving a 94.8% success rate. These results suggest that SPGrasp effectively mitigates the latency-interactivity trade-off in dynamic grasp synthesis.
| Original language | English |
|---|---|
| Pages (from-to) | 202-212 |
| Number of pages | 11 |
| Journal | IEEE Journal on Selected Topics in Signal Processing |
| Volume | 20 |
| Issue number | 2 |
| DOIs | |
| Publication status | Published - 1 Mar 2026 |
| Externally published | Yes |
Keywords
- Dynamic grasp synthesis
- prompt-driven grasping
- segment anything model
- spatiotemporal context
Fingerprint
Dive into the research topics of 'SPGrasp: Spatiotemporal Prompt-Driven Grasp Synthesis in Dynamic Scenes'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver