Sheng Yu 0009

dblp:181/2830-9 · DBLP profile ↗
← Back
16ranked-venue papers
14as first author
16since 2021 · last 2026
0009-0002-0709-1024ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 8 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 GraphGrasp: Lightweight and Efficient Graph-Guided 6-DoF Robotic Grasp Pose Estimation Network
abstract
6-DoF object grasping is a crucial skill for embodied intelligent robots. Previous methods often rely on large-scale networks for feature extraction, followed by grasp pose prediction, which increases the network's parameter count and overlooks the geometric and graph features of the point cloud. To address these challenges, we propose GraphGrasp, a graph-guided 6-DoF grasping pose prediction method. It performs graph analysis from the perspectives of scene, object, and grasping graphs. First, we introduce a graph feature embedding method based on local-global features to model the scene graph effectively. Then, we use a graph transformer strategy to represent spatial relationships between objects in the object graph. Finally, we propose a multi-metric, multi-level grasp pose evaluation algorithm to predict and explore graspable points, enabling effective construction of grasp graphs and accurate grasp pose evaluation. We test GraphGrasp on the GraspNet-1Billion dataset, and the results show that, compared to previous methods, it achieves nearly the same performance with about 1/5 of the parameters of state-of-the-art methods, significantly improving grasp pose prediction speed. Additionally, in real-world robot grasping scenarios, GraphGrasp outperforms previous methods in practical grasp pose prediction tasks.
Sheng Yu 0009, Dihua Zhai, Yuanqing Xia
AAAI1
2026 RGB-Based Category-Level Object Pose Estimation With Multi Pose Maps for Robotic Grasp Detection
abstract
RGB-D-based category-level object pose estimation has achieved very good pose estimation results in robotic grasp tasks. However, these methods rely on accurate depth information, and in industrial scenes where depth information is unknown or subject to significant interference, these methods are unable to apply. Therefore, this paper investigates the problem of category-level object pose estimation based solely on RGB images. First, to accurately predict the Normalized Object Coordinate Space (NOCS) map of objects in the scene, we build an encoding-decoding structure to achieve accurate NOCS map prediction. Then, to ensure an accurate transformation from NOCS maps to Intra-class Variation-Free Consensus (IVFC) maps, we propose a new deformable convolution that reduces the network’s computational load while improving the prediction accuracy of the IVFC maps. Finally, to fully utilize the information contained in multiple pose maps (NOCS maps and IVFC maps), we propose a Multi-Pose Map-based Pose (MPMPose) computation method to accurately predict the object pose. We test our method on the CAMERA25, REAL275, and Wild6D datasets, and the experimental results show that our proposed MPMPose can effectively complete the pose estimation task of unknown objects in scenes based solely on RGB images. Finally, we apply MPMPose to the robot grasping task in a real-world scenario. The experimental results show that MPMPose can effectively assist robots in completing the pose estimation task of objects in real scenes, enabling stable object grasping by robots.
Sheng Yu 0009, Dihua Zhai, Yunqiao Zeng, Yuanqing Xia
IEEE Trans Autom. Sci. Eng.1
2025 KeyPose: Category-Level 6D Object Pose Estimation with Self-Adaptive Keypoints
abstract
Category-level object pose estimation is an important task in computer vision. Some prior methods based on assumptions often struggle with drastic changes in object appearance. To address this challenge, we propose a new method for object pose estimation based on object-adaptive keypoints. In this paper, we first introduce a transformer-based keypoint prediction method for adaptive forecasting of point cloud keypoints. This method calculates the similarity between keypoint features and point cloud features, allowing keypoints to represent object geometry more effectively. Furthermore, to enhance the geometric feature construction of keypoints, we propose a graph-based keypoint feature aggregation method, which considers both the structural relationships between keypoints and the point cloud, strengthening the network's understanding of geometric structures. At this stage, keypoints remain at the geometric spatial level of the object and have not been predicted in NOCS. To improve the accuracy of keypoint prediction in NOCS, we design a NOCS voxelization method that divides NOCS into multiple voxels and accurately predicts NOCS keypoints within these voxels. Experimental results on multiple benchmark datasets demonstrate that our proposed KeyPose method outperforms all existing methods, achieving over 20% improvement in pose accuracy on some critical datasets.
Sheng Yu 0009, Dihua Zhai, Yuanqing Xia
AAAI1
2025 RCGNet: RGB-based Category-Level 6D Object Pose Estimation with Geometric Guidance
abstract
While most current RGB-D-based category-level object pose estimation methods achieve strong performance, they face significant challenges in scenes lacking depth information. In this paper, we propose a novel category-level object pose estimation approach that relies solely on RGB images. This method enables accurate pose estimation in real-world scenarios without the need for depth data. Specifically, we design a transformer-based neural network for category-level object pose estimation, where the transformer is employed to predict and fuse the geometric features of the target object. To ensure that these predicted geometric features faithfully capture the object’s geometry, we introduce a geometric feature-guided algorithm, which enhances the network’s ability to effectively represent the object’s geometric information. Finally, we utilize the RANSAC-PnP algorithm to compute the object’s pose, addressing the challenges associated with variable object scales in pose estimation. Experimental results on benchmark datasets demonstrate that our approach is not only highly efficient but also achieves superior accuracy compared to previous RGB-based methods. These promising results offer a new perspective for advancing category-level object pose estimation using RGB images.
Sheng Yu 0009, Dihua Zhai, Yuanqing Xia
IROS1
2025 Fast and Accurate Category-Level Object Pose Estimation Without Shape Priors for Robotic Grasp Detection
abstract
Category-level object pose estimation is crucial for enabling robot grasping. Currently, many methods rely on 3D shape priors for pose estimation, but obtaining priors for specific categories often requires a significant amount of time to generate. Although some methods do not rely on priors, However, these methods struggle to achieve a balance between speed and accuracy. Achieving fast and accurate category-level object pose estimation remains a challenging issue. In this paper, we propose an algorithm called FAPose, which aims to simultaneously achieve speed and accuracy in category-level object pose estimation without any shape priors. Firstly, we design an RGB-point cloud feature aggregation method based on transformers to fuse RGB and point cloud features. Secondly, we develop a dual-constraint object pose estimation method to effectively leverage both feature space and geometric space information. This approach constructs pose constraints in both feature space and geometric space to achieve optimal pose prediction. Finally, we validate FAPose through benchmark datasets and real robot grasping experiments. The experimental results demonstrate that our proposed method surpasses most existing state-of-the-art pose estimation methods, achieving superior pose estimation performance. The code for FAPose will be available after the paper is accepted for publication.
Sheng Yu 0009, Jian Yin 0032, Dihua Zhai, Yuanqing Xia
IEEE Trans Autom. Sci. Eng.1
2025 ZSPose: Instance-Level Zero-Shot Object Pose Estimation With Segment Anything Model
abstract
Estimating the poses of new objects is a challenging problem. Although many methods have been developed for instance-level object pose estimation, they often struggle when faced with new/unfamiliar objects. In this paper, we propose a zero-shot pose estimation method for new objects called ZSPose. We leverage SAM’s zero-shot feature to segment objects in cluttered environments and acquire masks for each object. To facilitate the matching of object masks with object models and categories, we propose a novel object matching strategy that aligns masks with the corresponding object models. Subsequently, based on the derived object masks, we produce object point clouds. Utilizing the RGB images of the objects alongside the point clouds, we present a feature-weight-based method for object pose estimation, achieving accurate pose estimation by predicting the matching weights between the model features and the point cloud features. We conduct performance testing on various instance-level object pose estimation datasets, and experimental results show that our proposed method significantly enhances the accuracy of object pose estimation. It demonstrates excellent generalization, making it applicable to pose estimation for a wide range of new objects. Finally, to validate the practical applicability of ZSPose, we apply it to real-world object pose estimation tasks and robotic grasping tasks. The experimental findings indicate that ZSPose effectively estimates the poses of new objects, assisting robots in performing practical grasping tasks, thus holding considerable practical value.
Sheng Yu 0009, Dihua Zhai, Jian Yin 0032, Yuanqing Xia
IEEE Trans Autom. Sci. Eng.1
2025 TCRNet: Transparent Object Depth Completion With Cascade Refinements
abstract
Transparent objects are commonly found in real life and industrial production. Unlike opaque objects, transparent objects are not easily identifiable in RGB images and often require depth information to determine their position in the image. However, due to the influence of other environmental factors such as reflection and refraction, the depth information of transparent objects is often inaccurate. This leads to difficulties for robots in grasping transparent objects, as incorrect depth information can result in the robot being unable to predict or predict incorrectly the grasping pose. Therefore, it is necessary to complete the depth information for transparent objects. Previous methods for depth completion of transparent objects often struggle to balance accuracy and real-time performance simultaneously. To achieve this goal, in this paper, we propose a transparent object depth completion network called TCRNet based on a cascade refinement structure, which balances accuracy and real-time performance simultaneously. First, the network incorporates a cascade refinement structure in the decoding stage to refine features multiple times, improving the accuracy of depth information. Additionally, an attention module is designed to adjust the extracted features, enabling the network to focus on depth information features in transparent object regions. Finally, a transformer-based error module is implemented in the network’s final output stage to predict and adjust the error between the depth image and the ground truth. TCRNet is trained and tested on three datasets: ClearGrasp, Omniverse Object, and TransCG. It outperforms previous methods in terms of performance. Furthermore, TCRNet is applied to existing grasp detection methods to conduct grasping experiments on transparent objects using a real Baxter robot.Note to Practitioners—With the development of RGB-D camera technology, RGB-D cameras are now widely used in various scenarios such as industrial production, autonomous driving, and robot grasping. However, in certain situations where the camera faces transparent or highly reflective objects, the depth information captured by the camera is often not accurate enough, which can lead to subsequent accidents. Therefore, it is necessary to repair and complete the depth images to achieve accurate understanding of the scene’s depth information. In recent years, with the advancement of deep learning, deep learning-based depth image processing and restoration techniques have been widely applied. In this paper, we propose a high-accuracy network for repairing depth images of transparent objects, which can accurately restore and estimate the depth information of transparent objects in various scenarios. Moreover, experimental results demonstrate that our proposed method can generalize well to other unknown scenes, achieving excellent results.
Dihua Zhai, Sheng Yu 0009, Yuyin Guan, Yuanqing Xia
IEEE Trans Autom. Sci. Eng.2
2025 Category-Level 6-D Object Pose Estimation With Shape Deformation for Robotic Grasp Detection
abstract
Category-level 6-D object pose estimation plays a crucial role in achieving reliable robotic grasp detection. However, the disparity between synthetic and real datasets hinders the direct transfer of models trained on synthetic data to real-world scenarios, leading to ineffective results. Additionally, creating large-scale real datasets is a time-consuming and labor-intensive task. To overcome these challenges, we propose CatDeform, a novel category-level object pose estimation network trained on synthetic data but capable of delivering good performance on real datasets. In our approach, we introduce a transformer-based fusion module that enables the network to leverage multiple sources of information and enhance prediction accuracy through feature fusion. To ensure proper deformation of the prior point cloud to align with scene objects, we propose a transformer-based attention module that deforms the prior point cloud from both geometric and feature perspectives. Building upon CatDeform, we design a two-branch network for supervised learning, bridging the gap between synthetic and real datasets and achieving high-precision pose estimation in real-world scenes using predominantly synthetic data supplemented with a small amount of real data. To minimize reliance on large-scale real datasets, we train the network in a self-supervised manner by estimating object poses in real scenes based on the synthetic dataset without manual annotation. We conduct training and testing on CAMERA25 and REAL275 datasets, and our experimental results demonstrate that the proposed method outperforms state-of-the-art (SOTA) techniques in both self-supervised and supervised training paradigms. Finally, we apply CatDeform to object pose estimation and robotic grasp experiments in real-world scenarios, showcasing a higher grasp success rate.
Sheng Yu 0009, Dihua Zhai, Yuyin Guan, Yuanqing Xia
IEEE Trans. Neural Networks Learn. Syst.1
2025 6-D Object Pose Estimation Based on Point Pair Matching for Robotic Grasp Detection
abstract
The 6-D pose estimation is a critical work essential to achieve reliable robotic grasping. Currently, the prevalent method is reliant on keypoint correspondence. However, this approach hinges on the determination of object keypoint locations, alongside their detection and localization in real scenes. It also employs the random sample consensus (RANSAC)-based perspective-n-point (PnP) algorithm to solve the pose. Yet, it is nondifferentiable and incapable of backpropagation with loss during the training phase. Alternatively, the direct regression method, while speedy and differentiable, falls short in terms of pose estimation performance, and thus needs enhancement. In view of these gaps, we investigate PPM6D, a new method for 6-D object pose estimation based on regression and point pair matching. Our methodology begins with a proposed cross-fusion module, designed to achieve the fusion and complementation of RGB features and point cloud features. Subsequently, an attention module adjusts the features of the object's 3-D model. Finally, we design a point pair matching module for effective matching of points and characteristics, resulting in an integral matching and fusion. PPM6D is extensively trained and tested utilizing benchmark datasets like LINEMOD, occlusion LINEMOD (LINEMOD-occ), YCB-Video, and T-LESS dataset. Experimental results prove that PPM6D can outperform many keypoint-based pose estimation methods, given its relatively rapid speed, thereby offering novel regression-based pose estimation ideas. When applied to real-world scenarios of object pose estimation tasks and grasp tasks of an actual Baxter robot, PPM6D demonstrates superior performance as compared to most alternatives.
Sheng Yu 0009, Dihua Zhai, Yufeng Zhan, Wencai Wang, Yuyin Guan, Yuanqing Xia
IEEE Trans. Neural Networks Learn. Syst.1
2024 CatFormer: Category-Level 6D Object Pose Estimation with Transformer
abstract
Although there has been significant progress in category-level object pose estimation in recent years, there is still considerable room for improvement. In this paper, we propose a novel transformer-based category-level 6D pose estimation method called CatFormer to enhance the accuracy pose estimation. CatFormer comprises three main parts: a coarse deformation part, a fine deformation part, and a recurrent refinement part. In the coarse and fine deformation sections, we introduce a transformer-based deformation module that performs point cloud deformation and completion in the feature space. Additionally, after each deformation, we incorporate a transformer-based graph module to adjust fused features and establish geometric and topological relationships between points based on these features. Furthermore, we present an end-to-end recurrent refinement module that enables the prior point cloud to deform multiple times according to real scene features. We evaluate CatFormer's performance by training and testing it on CAMERA25 and REAL275 datasets. Experimental results demonstrate that CatFormer surpasses state-of-the-art methods. Moreover, we extend the usage of CatFormer to instance-level object pose estimation on the LINEMOD dataset, as well as object pose estimation in real-world scenarios. The experimental results validate the effectiveness and generalization capabilities of CatFormer. Our code and the supplemental materials are avaliable at https://github.com/BIT-robot-group/CatFormer.
Sheng Yu 0009, Dihua Zhai, Yuanqing Xia
AAAI1
2024 FANet: Fast and Accurate Robotic Grasp Detection Based on Keypoints
abstract
In practice, the real-time and accuracy of robotic grasp detection are two very important metrics. In the past, researchers had to sacrifice the real-time nature of the detection network in order to obtain higher detection accuracy. How to make the real-time and accuracy of the network co-exist is a problem worth studying. In order to solve this problem, this paper proposes a network, FANet, based on grasp keypoints, which improves the accuracy of grasp detection while ensuring the real-time performance. The key of this paper is how to quickly and accurately detect grasped keypoints. To this end, this paper proposes a local refinement module that optimizes and de-duplicates each feature of the multi-scale feature map, enabling the network to make full use of the multi-scale features. We also propose a global feature refinement module that allows the network to make better use of global features. We also propose a grasp keypoint optimization module that predicts the offset between the actual keypoints and the predicted keypoints, enabling the network to predict the keypoints more accurately. Moreover, we develop two FANets specifically for grasp detection on CPU and GPU, both of which can accomplish real-time grasp detection in real-world scenes. We complete the training and testing of FANet on the Cornell dataset and the Jacquard dataset, achieving SOTA results on the Jacquard dataset. We also test FANet on a dataset of unknown objects, all with good results. Finally, we use the FANet in grasping experiments with an actual Baxter robot and achieve an average grasping success rate of 96%.Note to Practitioners—Real-time and accuracy are two very important metrics in robotic grasping detection. To achieve high accuracy, more time is often consumed for feature extraction. Similarly, in order to improve the real time performance, we need to reduce the time consumed in the feature extraction process, which may result in a drop in detection accuracy. How to coordinate the relationship between them, so as to have both, is a problem worth investigating. Current methods tend to focus on obtaining higher accuracy and they are willing to spend more time to achieve higher accuracy. But in some practical scenarios, such as on factory assembly lines, objects move fast, and the network needs to be able to detect the grasping position quickly, the real-time performance is more important, which makes some methods difficult to use. In addition, most of the current methods tend to focus on GPU-based robotic grasp detection methods, and in real-world scenarios we may not have such a powerful processing GPU available. In contrast, the CPU is an indispensable unit of the computer that we can use to process images without a high-performance GPU. However, compared to GPUs, the CPUs’ image processing capability is poor, making it difficult to achieve real-time processing. Faced with this situation, the problem of how to achieve real-time and high accuracy in a CPU-only robotic grasp detection network is worth studying, but most of the existing methods ignore this problem. To address these problems, we propose a Fast and Accurate robotic grasp detection Network (FANet), which not only enables the network to combine real-time and accuracy, but also enables real-time detection on CPU or GPU.
Dihua Zhai, Sheng Yu 0009, Yuanqing Xia
IEEE Trans Autom. Sci. Eng.2
2024 An Efficient Robotic Pushing and Grasping Method in Cluttered Scene
abstract
Pushing and grasping (PG) are crucial skills for intelligent robots. These skills enable robots to perform complex grasping tasks in various scenarios. These PG methods can be categorized into single-stage and multistage approaches. Single-stage methods are faster but less accurate, while multistage methods offer high accuracy at the expense of time efficiency. To address this issue, a novel end-to-end PG method called efficient PG network (EPGNet) is proposed in this article. EPGNet achieves both high accuracy and efficiency simultaneously. To optimize performance with fewer parameters, EfficientNet-B0 is used as the backbone of EPGNet. Additionally, a novel cross-fusion module is introduced to enhance network performance in robotic PG tasks. This module fuses and utilizes local and global features, aiding the network in handling objects of varying sizes in different scenes. EPGNet consists of two branches dedicated to predicting PG actions, respectively. Both branches are trained simultaneously within a Q-learning framework. Training data is collected through trial and error, involving the robot performing PG actions. To bridge the gap between simulation and reality, a unique PG dataset is proposed. Additionally, a YOLACT network is trained on the PG dataset to facilitate object detection and segmentation. A comprehensive set of experiments is conducted in simulated environments and real-world scenarios. The results demonstrate that EPGNet outperforms single-stage methods and offers competitive performance compared to multistage methods, all while utilizing fewer parameters. A video is available at https://youtu.be/HNKJjQH0MPc.
Sheng Yu 0009, Dihua Zhai, Yuanqing Xia, Yuyin Guan
IEEE Trans. Cybern.1
2024 CatTrack: Single-Stage Category-Level 6D Object Pose Tracking via Convolution and Vision Transformer
abstract
In the current research, many researchers have focused on instance-level pose tracking, which requires a 3D model of the object in advance, making it challenging to apply in practice. To address this limitation, some researchers have proposed the category-level object pose tracking method. Achieving accurate and speedy monocular category-level pose tracking is an essential research goal. In this article, we propose CatTrack, a new single-stage keypoints-based monocular category-level multi-object pose tracking network. A significant issue in object pose tracking tasks is utilizing the information from the previous frame to guide pose estimation for the next frame. However, as the object poses and camera information in each frame are different, we need to remove irrelevant information and emphasize useful features. To this end, we propose a transformer-based temporal information capture module to leverage the position information of keypoints from the previous frame. Furthermore, we propose a new keypoint matching module to enable the grouping and matching of object keypoints in complex scenes. We have successfully applied CatTrack to the Objectron dataset and achieved superior results in comparison to existing methods. Furthermore, we have also evaluated the generalization of CatTrack and successfully applied it to track the 6D pose of unseen real-world objects.
Sheng Yu 0009, Dihua Zhai, Yuanqing Xia
IEEE Trans. Multim.1
2024 A Novel Robotic Pushing and Grasping Method Based on Vision Transformer and Convolution
abstract
Robotic grasping techniques have been widely studied in recent years. However, it is always a challenging problem for robots to grasp in cluttered scenes. In this issue, objects are placed close to each other, and there is no space around for the robot to place the gripper, making it difficult to find a suitable grasping position. To solve this problem, this article proposes to use the combination of pushing and grasping (PG) actions to help grasp pose detection and robot grasping. We propose a pushing-grasping combined grasping network (GN), PG method based on transformer and convolution (PGTC). For the pushing action, we propose a vision transformer (ViT)-based object position prediction network pushing transformer network (PTNet), which can well capture the global and temporal features and can better predict the position of objects after pushing. To perform the grasping detection, we propose a cross dense fusion network (CDFNet), which can make full use of the RGB image and depth image, and fuse and refine them several times. Compared with previous networks, CDFNet is able to detect the optimal grasping position more accurately. Finally, we use the network for both simulation and actual UR3 robot grasping experiments and achieve SOTA performance. Video and dataset are available at https://youtu.be/Q58YE-Cc250.
Sheng Yu 0009, Dihua Zhai, Yuanqing Xia
IEEE Trans. Neural Networks Learn. Syst.1
2024 Robotic Grasp Detection With 6-D Pose Estimation Based on Graph Convolution and Refinement
abstract
Six-dimensional (6-D) object pose estimation plays a critical role in robotic grasp, which performs extensive usage in manufacturing. The current state-of-the-art pose estimation techniques primarily depend on matching keypoints. Typically, these methods establish a correspondence between 2-D keypoints in an image and the corresponding ones in a 3-D object model. And then they use the PnP-RANSAC algorithm to determine the 6-D pose of the object. However, this approach is not end-to-end trainable and may encounter difficulties when applied to scenarios necessitating differentiable poses. When employing a direct end-to-end regression method, the outcomes are often inferior. To tackle the mentioned problems, we present GR6D, which is a keypoint-and graph-convolution-based neural network for differentiable pose estimation based on RGB-D data. First, we propose a multiscale fusion method that utilizes convolution and graph convolution to exploit information contained in RGB and depth images. Additionally, we propose a transformer-based pose refinement module to further adjust features from RGB images and point clouds. We evaluate GR6D on three datasets: 1) LINEMOD; 2) occlusion LINEMOD; and 3) YCB-Video dataset, and it outperforms most state-of-the-art methods. Finally, we apply GR6D to pose estimation and the robotic grasping task in the real world, manifesting superior performance.
Sheng Yu 0009, Dihua Zhai, Yuanqing Xia, Chengyu Zhang 0010
IEEE Trans. Syst. Man Cybern. Syst.1
2023 SKGNet: Robotic Grasp Detection With Selective Kernel Convolution
abstract
Real-time and accuracy are important evaluation metrics of robotic grasp detection algorithms. To further improve the accuracy on the premise of ensuring real-time performance, in this paper, a new Selective Kernel convolution Grasp detection Network (SKGNet) is proposed. Compared with previous methods, the attention mechanism and multi-scale fusion features are integrated into the SKGNet, which makes the network not only pay full attention to the grasp area but also flexibly adjust the grasp area according to the scale of the object, thus effectively distinguishing the object from the background. The SKGNet is trained and tested on the Cornell dataset and the Jacquard dataset, with the accuracy of 99.1% and 95.9% respectively, which is superior to SOTA methods. Moreover, SKGNet’s detection speed has reached 28fps. To demonstrate the performance of SKGNet, comparison studies and ablation experiments are performed in this paper. Finally, the grasp experiments of Baxter robot are also performed to verify the generalization of SKGNet in the actual scene, which achieves an average grasping success rate of 96.5%. Video is available athttps://youtu.be/j07sb_ChzWQ. Note to Practitioners—Autonomous grasping is an very important skill for the robotic systems in the real world. However, due to the low grasp detection accuracy, robotic grasp is still a challenging problem. Although some methods have been developed to improve the grasp detection accuracy, the time efficiency is poor. Grasp detection with good accuracy and efficiency is worthy of further study. In this view, a novel deep learning-based grasp detection network SKGNet is proposed in this paper. It takes RGB-D images as input, trains and tests on public datasets, and finally outputs a series of grasp rectangles. Compared with the existing works, it not only achieves state-of-the-art detection accuracy, but also has high efficiency. To demonstrate the generalization performance and effectiveness, the SKGNet is also tested in the real world, and applied to perform the actual grasp task of Baxter robot. The results show that the SKGNet has good robustness and can detect the unknown objects of different sizes and shapes in the real world well.
Sheng Yu 0009, Dihua Zhai, Yuanqing Xia
IEEE Trans Autom. Sci. Eng.1