Abstract
RGB-D-based category-level object pose estimation has achieved very good pose estimation results in robotic grasp tasks. However, these methods rely on accurate depth information, and in industrial scenes where depth information is unknown or subject to significant interference, these methods are unable to apply. Therefore, this paper investigates the problem of category-level object pose estimation based solely on RGB images. First, to accurately predict the Normalized Object Coordinate Space (NOCS) map of objects in the scene, we build an encoding-decoding structure to achieve accurate NOCS map prediction. Then, to ensure an accurate transformation from NOCS maps to Intra-class Variation-Free Consensus (IVFC) maps, we propose a new deformable convolution that reduces the network's computational load while improving the prediction accuracy of the IVFC maps. Finally, to fully utilize the information contained in multiple pose maps (NOCS maps and IVFC maps), we propose a Multi-Pose Map-based Pose (MPMPose) computation method to accurately predict the object pose. We test our method on the CAMERA25, REAL275, and Wild6D datasets, and the experimental results show that our proposed MPMPose can effectively complete the pose estimation task of unknown objects in scenes based solely on RGB images. Finally, we apply MPMPose to the robot grasping task in a real-world scenario. The experimental results show that MPMPose can effectively assist robots in completing the pose estimation task of objects in real scenes, enabling stable object grasping by robots. Note to Practitioners - Robotic grasp detection based on 6D object pose estimation is an important way to realize reliable robotic grasp. However, in some practical scenarios, the depth image of an object is significantly affected or cannot be obtained, resulting in the depth image being unusable. In this case, it is necessary to perform further pose estimation tasks based on RGB images. However, in the task of category-level object pose estimation, the object model cannot be obtained; only the category of the object and the mask of the object in the image are available. There is no distance information or geometric information about the object, and relying solely on this information to complete the pose estimation poses significant challenges. For previous methods, although they can accomplish the task of object pose estimation, their accuracy is not high, and their real-time performance is poor, which is critical for practical robotic grasping applications. To solve these problems, we design a new pose estimation method that effectively improves the accuracy of pose estimation while enhancing computational speed, and can generalize well to object pose estimation tasks in real-world scenarios, demonstrating good practical application value.
| Original language | English |
|---|---|
| Pages (from-to) | 11272-11284 |
| Number of pages | 13 |
| Journal | IEEE Transactions on Automation Science and Engineering |
| Volume | 23 |
| DOIs | |
| Publication status | Published - 2026 |
| Externally published | Yes |
Keywords
- Category-level object pose estimation
- deformable convolution
- multiple pose maps
- object detection
Fingerprint
Dive into the research topics of 'RGB-Based Category-Level Object Pose Estimation with Multi Pose Maps for Robotic Grasp Detection'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver