Qing Zhu 0003

dblp:74/963-3 · DBLP profile ↗
← Back
21ranked-venue papers
1as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 High-Precision Multi-Instance Registration for Stacked Objects in Bin-Picking Scenes
abstract
In industrial bin-picking, robotic systems must estimate the poses of multiple object instances, where accurate pose estimation is essential for reliable downstream manipulation and grasping. Most existing multi-instance registration methods primarily establish point correspondences based on local features to alleviate the challenges posed by occlusion and clutter. However, local features are easily disturbed by neighboring instances and lack global context, leading to unreliable correspondences and degraded registration accuracy. In addition, the absence of rotational invariance further reduces correspondence accuracy in scenes with stacked instances and highly varying object orientations. To address these challenges, we present a one-stage multi-instance point cloud registration framework for stacked-object scenes. Our framework incorporates a rotation-invariant operator to enhance the robustness of feature representations under arbitrary orientations. Then, we propose a Center-Aware Res-Masked Transformer module, which incorporates an object center embedding to enrich global instance-level context and a center-aware residual mask prediction module to balance weight distribution across objects of varying sizes during training. Extensive experiments on the challenging ROBI dataset demonstrate that our method outperforms the competitive baseline MIRETR by more than 10% in mean precision, highlighting its effectiveness in complex bin-picking scenes. Furthermore, evaluations on the unstacked Scan2CAD dataset confirm the generalizability of the proposed framework across different application scenarios.
Jiawen Zhao, Qing Zhu 0003, Yaonan Wang 0001, Weixing Peng, Jianxu Mao, Min Liu 0008, Xuebing Liu, Hui Zhang 0023
IEEE Trans. Circuits Syst. Video Technol.2
2026 Remaining Useful Life Prediction for Key Components of Transportation Vehicles: A Physics-Informed Perspective
abstract
In transportation systems, accurately estimating the remaining useful life (RUL) of critical components, such as aircraft engines, Battery Management Systems (BMSs), is crucial for the safe and reliable operation and manufacturing of transportation vehicles. However, most existing research overlooks the underlying physical information, which is vital for more precise RUL prediction. To fill this gap, this paper proposes a physics-informed method for predicting the RUL of key components of transportation vehicles. By integrating the Mamba network with a multi-head attention mechanism, we capture and emphasize key features and trends in the equipment’s operational state, improving prediction accuracy. Additionally, we introduce a Physics-Informed Neural Network (PINN) framework to model the underlying physical relationships between RUL and sensor data, incorporating these relationships as a regularization term in the loss function to enhance predictive capability and interpretability. We conducted experimental validation using the C-MAPSS aircraft engine dataset (operation) and the transportation vehicle chip manufacturing dataset (manufacture). The results show that the proposed method significantly improves the accuracy of RUL prediction, providing strong support for the intelligent maintenance and reliability management of key components in transportation vehicles.
Qing Zhu 0003, Yucong Shi, Yun Feng 0001, Ya-Zhi Zhang, Haoran Tan, Yaonan Wang 0001, Wanke Yu, Yongfu Li 0001
IEEE Trans. Intell. Transp. Syst.1
2025 Multi-range Adaptive Perception Transformer for Iterative Homography Estimation
abstract
Homography estimation is fundamental to various vision tasks. Iteration-based methods have recently achieved significant success in this field. However, errors introduced during iterations can lead to increased image deformation. Existing methods often focus on capturing local correspondences in the later stages of iteration while downplaying global ones, which may cause errors to persist and propagate into subsequent iterations, ultimately leading to error accumulation. To alleviate this issue, we propose Multi-range Adaptive Perception Transformer for Iterative Homography Estimation (MAPTHomo), which integrates Multi-range Attention (MRA) and Adaptive Perception Module (APM). Specifically, MRA captures both global and local correspondences, enabling the model to adapt to varying levels of deformation. The APM dynamically adjusts attention focus based on the current context. The combination of MRA and APM enhances the error-correction capability of the iterative process, effectively mitigating error accumulation. Extensive experiments demonstrate that MAPTHomo outperforms previous methods and exhibits strong generalization ability.
Tianming Li, Qing Zhu 0003, Zhen Zhou 0003, Jianqiao Luo, Yaonan Wang 0001
ICASSP2
2025 DRLHomo: Disentangled Representation Learning for Cross-Modal Homography Estimation
Tianming Li, Zhen Zhou 0003, Qing Zhu 0003, Jianqiao Luo, Yaonan Wang 0001
PRCV (9)3
2025 RAMPGrasp: Retentive Attention-Based Multiscale Perception Grasp Detection Network
abstract
In robotic grasp detection, challenges such as uncertainty in object type, size, and placement within the scene diminish grasping accuracy. However, the inability to effectively locate the graspable area and incomplete feature extraction for grasp detection are two key factors that hinder grasp detection accuracy and are not considered in current methods. This paper presents a novel retentive attention-based multiscale perception grasp detection network (RAMPGrasp) to address this constraint. First, we introduce retentive attention in the feature extraction module, which significantly improves the efficiency of attention score computation for long sequences in visual tasks. Second, we propose a multiscale spatial pyramid attention module, which can effectively adjust the importance of multiscale feature sequences and feature channels, while enhancing the correlation of multiscale features. Third, we design the prediction module as a coarse-to-fine framework, improving feature representation for grasp detection by considering the distribution trend of grasp poses. As a result, RAMPGrasp achieves state-of-the-art grasp detection accuracy, with 98.4% and 95.6% on the Cornell and Jacquard datasets, respectively.
Jianan Huang 0002, Xuebing Liu, Qing Zhu 0003, Yaonan Wang 0001, Mingtao Feng, Zhen Zhou 0003, Lin Chen 0034, Danwei Wang
IEEE Trans. Circuits Syst. Video Technol.3
2025 A State Space Model for Multiobject Full 3-D Information Estimation From RGB-D Images
abstract
Visual understanding of 3-D objects is essential for robotic manipulation, autonomous navigation, and augmented reality. However, existing methods struggle to perform this task efficiently and accurately in an end-to-end manner. We propose a single-shot method based on the state space model (SSM) to predict the full 3-D information (pose, size, shape) of multiple 3-D objects from a single RGB-D image in an end-to-end manner. Our method first encodes long-range semantic information from RGB and depth images separately and then combines them into an integrated latent representation that is processed by a modified SSM to infer the full 3-D information in two separate task heads within a unified model. A heatmap/detection head predicts object centers, and a 3-D information head predicts a matrix detailing the pose, size and latent code of shape for each detected object. We also propose a shape autoencoder based on the SSM, which learns canonical shape codes derived from a large database of 3-D point cloud shapes. The end-to-end framework, modified SSM block and SSM-based shape autoencoder form major contributions of this work. Our design includes different scan strategies tailored to different input data representations, such as RGB-D images and point clouds. Extensive evaluations on the REAL275, CAMERA25, and Wild6D datasets show that our method achieves state-of-the-art performance. On the large-scale Wild6D dataset, our model significantly outperforms the nearest competitor, achieving 2.6% and 5.1% improvements on the IOU-50 and 5°10 cm metrics, respectively.
Qing Zhu 0003, Yaonan Wang 0001, Mingtao Feng, Jian Liu 0014, Jianan Huang 0002, Ajmal Mian
IEEE Trans. Cybern.2
2025 Multiscale Spherical Feature Decoupling Network for Multimodal Image Registration
abstract
Multimodal image registration plays a crucial role in advancing Earth science. However, significant appearance variations and geometric deformations between multimodal images pose considerable challenges to this task. In this paper, we propose a novel multiscale spherical feature decoupling network (MSFDNet) for multimodal image registration by combining a multiscale iterative strategy with a multimodal decoupling strategy. MSFDNet adopts a multiscale architecture, with each scale incorporates a spherical feature decoupling (SFD) module with a carefully crafted three-stage decoupling strategy to bridge the modality gap. Specifically, we first introduce asymmetric shared and unique feature encoders to extract modality-shared and modality-unique features. Next, we design a spherical constraint learning (SCL) module to project the extracted features into spherical space, leveraging its regularized distance properties to enhance feature separability during the decoupling process. Finally, we propose a dual-path reconstruction mechanism that combines self-reconstruction with homography-guided cross-reconstruction to reconstruct the original multimodal images from the decoupled features, thereby simultaneously enhancing the learning of both feature decoupling and registration network. Based on the decoupled modality-shared features, we predict the registration function in a multiscale iterative manner, effectively bridging the geometric gap. Extensive experiments on multiple multimodal datasets validate the effectiveness of MSFDNet and demonstrate its state-of-the-art performance.
Tianming Li, Zhen Zhou 0003, Qing Zhu 0003, Jianqiao Luo, Yaonan Wang 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 Uncertainty Guided Deep Lucas-Kanade Homography for Multimodal Image Alignment
abstract
Homography estimation for multimodal images poses a considerable challenge in computer vision because of content disparities and the diverse feature points captured by different sensors. Existing methods typically extract feature maps using neural networks and apply the Lucas-Kanade (LK) algorithm, which is based on the brightness constancy assumption, to solve the homography matrix. However, applying this assumption across all pixel features in multimodal images can lead to inaccuracies, as these images often contain noise, such as homogeneous regions or considerable appearance variations, which can corrupt the network’s training. To address this problem, we propose an uncertainty-guided deep LK (UG-DLK) framework that integrates uncertainty predictions to enhance the network’s iterative learning process. Specifically, we employ a probabilistic approach where the network predicts the distribution of the feature map rather than fixed values. By designing an uncertainty neighborhood estimator, we unfold the cost volume along the channels into 2-D slices, allowing the model to focus on neighborhood information at specific locations, effectively reducing the interference from spatial neighborhoods in the estimation of feature uncertainty. Through uncertainty modeling, the network can accurately identify scenes and objects that comply with the brightness constancy constraint, leading to more robust learning outcomes. Additionally, we introduce a novel loss function that incorporates feature uncertainty, leading to a smoother optimization landscape near the true homography parameters and reducing convergence oscillations. Our method, which is evaluated on benchmark datasets such as Google Maps, Google Earth, MSCOCO, and DPDN, demonstrates state-of-the-art performance, confirming the robustness and adaptability of our model across various scenarios.
Zhen Zhou 0003, Jianqiao Luo, Qing Zhu 0003, Yaonan Wang 0001, Hang Zhong, Mingtao Feng, Lin Chen 0034
IEEE Trans. Geosci. Remote. Sens.3
2025 Registration of Multiview Point Clouds With Unknown Overlap
abstract
Registration of multiview point clouds obtained from 3D scanners is a common method for 3D reconstruction. However, most existing registration methods are designed to handle point clouds with known overlap relationships that are ensured by external equipment (e.g., manipulators, turntables) or acquisition sequences, which limits the application range and increases the acquisition cost. To overcome these limitations, an unknown overlap registration (UOR) method for multiview point clouds is proposed, which can estimate overlap confidence, construct a connected graph, and remove outlier point clouds automatically. First, the overlap confidence between two point clouds is estimated by calculating the average nearest neighbor feature distance within the predicted overlap region. We then construct a minimal spanning tree based on the confidence levels and search for the central node to serve as the world coordinate. Finally, the Lie algebra-based SE(3)-sensitive perturbation scheme is introduced to solve the fine transformations, in which a robust weighting function is designed to weight point correspondences. Our method can find reliable connections among point clouds, and the proposed graph can be combined with different pairwise registration methods. The experimental results on both indoor and industrial datasets demonstrate the accuracy and effectiveness of our method.
Jiawen Zhao, Qing Zhu 0003, Yaonan Wang 0001, Weixing Peng, Hui Zhang 0023, Jianxu Mao
IEEE Trans. Multim.2
2025 Deep spatial and discriminative feature enhancement network for stereo matching
Guowei An, Yaonan Wang 0001, Kai Zeng 0010, Qing Zhu 0003, Xiaofang Yuan
Vis. Comput.4
2024 Interpretable Unsupervised Homography Estimation
Zhen Zhou 0003, Qing Zhu 0003, Yaonan Wang 0001, Yang Mo, Lin Chen 0034, Jianan Huang 0002, Tianjian Jiang
PRCV (2)2
2024 3D Object Detection From Point Cloud via Voting Step Diffusion
abstract
3D object detection is a fundamental task in scene understanding. Numerous research efforts have been dedicated to better incorporate Hough voting into the 3D object detection pipeline. However, due to the noisy, cluttered, and partial nature of real 3D scans, existing voting-based methods often receive votes from the partial surfaces of individual objects together with severe noises, leading to sub-optimal detection performance. In this work, we focus on the distributional properties of point clouds and formulate the voting process as generating new points in the high-density region of the distribution of object centers. To achieve this, we propose a new method to move random 3D points toward the high-density region of the distribution by estimating the score function of the distribution with a noise conditioned score network. Specifically, we first generate a set of object center proposals to coarsely identify the high-density region of the object center distribution. To estimate the score function, we perturb the generated object center proposals by adding normalized Gaussian noise, and then jointly estimate the score function of all perturbed distributions. Finally, we generate new votes by moving random 3D points to the high-density region of the object center distribution according to the estimated score function. Extensive experiments on two large scale indoor 3D scene datasets, SUN RGB-D and ScanNet V2, demonstrate the superiority of our proposed method. The code will be released athttps://github.com/HHrEtvP/DiffVote.
Haoran Hou, Mingtao Feng, Weisheng Dong, Qing Zhu 0003, Yaonan Wang 0001, Ajmal Mian
IEEE Trans. Circuits Syst. Video Technol.5
2024 Unsupervised Homography Estimation With Pixel-Level SVDD
abstract
Homography estimation is a common image alignment method. Unsupervised learning, which uses unlabeled training and exhibits excellent performance, has attracted much attention in this field. When there are multiple planes in the scene, using features over the entire image for matching will lead to compromised results. However, existing methods for learning focused principal plane masks through deep neural networks lack explicit guidance. In this paper, we propose a novel unsupervised method to explicitly model anomaly descriptor removal and mask generation. Specifically, reliable feature descriptors are selected from a novel perspective, and regard the features that are not responsible for alignment as outliers. The pixel-level support vector data description (PL-SVDD) module is designed. This module learns the feature representation of image pixels and fits a hypersphere to exclude the feature redundancy information that is not responsible for alignment from the hypersphere, thereby optimizing the feature descriptor. Based on the optimized image features, a correlation learning (CL) module is designed. This module displays a generated mask through mathematical modeling to select reliable areas for homography estimation. Specifically, the feature descriptor of one unaligned images is modeled as a multivariate Gaussian distribution by Gaussian density estimation (GDE). Then, The Mahalanobis distance is combined with the multivariate Gaussian distribution of the model and the feature descriptor of another image to generate the mask. Experiments show that our method achieves good performance compared with previous methods.
Zhen Zhou 0003, Qing Zhu 0003, Mingtao Feng, Yaonan Wang 0001, Jianqiao Luo, Zhiqiang Miao, Lin Chen 0034, Yang Mo
IEEE Trans. Circuits Syst. Video Technol.2
2024 A Depth Adaptive Feature Extraction and Dense Prediction Network for 6-D Pose Estimation in Robotic Grasping
abstract
Estimating the 6-D pose of an object is a vital and challenging task for robot vision systems in industrial robotic grasping. With the wide use of 3-D cameras, the additional acquired depth image provides geometric information of the scene to increase the pose estimation performance but leads to a challenge, fully leveraging the two-modal data, the color image and the depth image. Previous works usually adopt two individual strategies to handle the data, which suffer from limited accuracy and efficiency since the two complementary data are not fully explored. Thus, we propose a depth adaptive feature extraction and dense prediction network that decouples the scale-dependent and the scale-invariant information from the depth image. The former guides the network to perceive the 3-D structure of the scene, and the latter, together with color image, provides the scene textures for feature extraction. The proposed network not only fuses multimodal textures but also retains their 3-D structure. In addition, a dense prediction strategy is adopted to regress the object pose; this approach can mitigate the instability caused by outliers. We conduct various evaluations on a real-world industrial dataset to illustrate the advantages of the proposed approach; and a practical robotic grasping platform is presented to demonstrate its application performance.
Xuebing Liu, Xiaofang Yuan, Qing Zhu 0003, Yaonan Wang 0001, Mingtao Feng, Zhen Zhou 0003
IEEE Trans. Ind. Informatics3
2024 PoseDiffusion: A Coarse-to-Fine Framework for Unseen Object 6-DoF Pose Estimation
abstract
Accurately estimating the six-degrees of freedom (DoF) pose of unseen objects is crucial for successful robotic manipulation in industrial automation. Some existing methods for this task rely on prior knowledge of individual objects, i.e., the model must be trained on the exact object instance or object category. Others perform unseen object pose estimation but are limited in their feature learning and pose refinement ability. To address these problems, we propose an unseen object pose estimation method that follows a coarse-to-fine framework and leverages the powerful learning ability of diffusion models. We introduce a diffusion model for generating object poses, and conduct a comparison between the generated poses and the original pose to determine the optimal one. We design a novel pose estimation module to provide coarse poses for the PoseDiffusion. This module comprises two feature extraction modules that extract global and masked features. In addition, we propose a strategy to estimate the pose by comparing the similarity between rendered and query poses. The renderings of an unseen object from various viewpoints are generated from its computer-aided design (CAD) model. Our method requires a CAD model of the unseen object only during inference, a scenario well suited to industrial applications. Experimental evaluation on benchmark datasets demonstrates that the proposed framework outperforms existing approaches, achieving state-of-the-art performance in six-DoF object pose estimation.
Qing Zhu 0003, Yaonan Wang 0001, Mingtao Feng, Chengzhong Wu, Xuebing Liu, Jianan Huang 0002, Ajmal Mian
IEEE Trans. Ind. Informatics2
2024 MG-GCT: A Motion-Guided Graph Convolutional Transformer for Traffic Gesture Recognition
abstract
For autonomous driving systems, it is crucial to recognize the actions and gestures of traffic conductors and cyclists on the road to ensure safety. However, traffic gesture recognition is more challenging than action recognition in general scenarios due to the differences in action posture and sample composition between traffic gesture datasets and general action datasets. Therefore, general action recognition methods cannot identify traffic gestures well. To overcome these problems, we propose a novel motion-guided graph convolutional transformer (MG-GCT) for traffic gesture recognition. Firstly, we proposed a two-stream network to fully utilize joint data and motion data for action recognition. Secondly, we designed and implemented a motion-guided module between two streams, which leverages the powerful spatial representation ability of the motion data to guide the learning of the joint data stream in the spatial dimension. Thirdly, we implemented a temporal transformer network to process the temporal features of the skeleton. Finally, we conducted extensive experiments on two public datasets and one dataset presented by us to demonstrate the effectiveness of our network in traffic gesture recognition, which has a significant advantage over the state-of-the-art methods.
Qing Zhu 0003, Yaonan Wang 0001, Yang Mo
IEEE Trans. Intell. Transp. Syst.2
2023 Improved YOLOv7 Based on Transformer for Object Detection in UAV-Captured Images
abstract
As the drone captures image targets at different flying altitudes, their scales may vary significantly, which can pose challenges for the object detection model to accurately detect them. Additionally, tiny objects in the image contain minimal information, making them difficult to distinguish from the background. To overcome these two challenges, we proposed a network architecture that aims to improve the accuracy of tiny object detection in drone images. Specially, we designed a tiny object detector(TOD) that can effectively extract features of tiny objects and distinguish between tiny object features and image background. Furthermore, this TOD module contains a Convolutional Visual Attention Network (CVAN) to better focus on the regions of tiny objects. Experimental results demonstrate that the proposed method achieves [email protected] accuracy of 53.9% on the VisDrone2021-test-dev dataset and improves by 2.8 % compared to YOLOv7.
Yuefan Luo, Qing Zhu 0003, Zhen Zhou 0003, Lin Chen 0034, Tianjian Jiang, Yijiang Li, Danwei Wang, Yaonan Wang 0001
SMC2
2023 Deep Confidence Propagation Stereo Network
abstract
Stereo matching depth estimation based on rectified image pairs is of great importance to many computer vision tasks such as vehicle navigation and autonomous driving. Confidence measures are typically used to refine stereo matching results, which provides robustness and efficiency for disparity estimation. However, previous learning-based confidence methods for stereo matching usually use the middle results or composition as a post-processing step to refine the stereo matching results. This cannot be optimized end-to-end and the performance is limited by the quality of the tri-modal output. To handle this issue, in this paper, we pursue an end-to-end hierarchical architecture and propose a differentiable confidence propagation (DCP) model of a cost aggregation network for stereo matching. The DCP model is integrated into an end-to-end neural network hierarchical architecture to guide matching cost volume aggregation. More specifically, to better represent the similarity of left and right feature maps, we extract unary context feature maps with an effective attention mechanism for matching cost construction. Moreover, we aggregate the cost volume with the multiple stacked DCP cost aggregation (DCPCA) networks to generate a more reliable and finer cost volume. This network suppresses multi-level disparity maps. Each output disparity is supervised with different training weights to learn in a coarse-to-fine way. Our method outperforms previous methods on the Sceneflow dataset by achieving the$0.6735px$EPE error, achieving 1.53% D1-all metric of Non-occluded pixels regions and 0.72% Non-occluded pixels of$5px$metric on KITTI 2015 and 2012 dataset. Extensive experiments carried out on the KITTI Stereo benchmarks demonstrate that our DCPCA-Net can significantly minimize the trade-off between accuracy and efficiency for stereo matching.
Kai Zeng 0010, Yaonan Wang 0001, Wei Wang 0025, Hui Zhang 0023, Jianxu Mao, Qing Zhu 0003
IEEE Trans. Intell. Transp. Syst.6
2022 Deep Progressive Fusion Stereo Network
abstract
Stereo matching depth estimation for rectified image pairs is of great importance to many compute vision tasks, specifically in autonomous driving. With the flourishing of convolution neural networks, responsible depth estimation of stereo matching with artificial intelligence is the most severe challenge for autonomous driving in recent years. Previous research on end-to-end trainable stereo matching networks has usually used cascading convolution blocks with down-sampling or pooling operations to extract the unary features required for matching cost construction. Such approaches lack a reconstruction stage for increasing feature map pixel-wise alignment and strength, factors which play an important role in representing the similarity between stereo image pairs. To address this issue, in this paper, we propose the progressive fusion stereo matching network (PFSM-Net). We exploit an encoder-decoder feature extraction network architecture for multi-stage and -scale dynamic feature extraction. Moreover, we propose a group-wise concatenation method to construct the cost volume, which provides a more efficient cost volume for cost aggregation. Furthermore, we propose the use of multi-scale cost aggregation networks with a progressive fusion strategy. The aggregated cost volume is progressively fused with the multi-stage and -scale cost volume as the size of the cost volume increases. Multi-stage and -scale outputs are supervised with and learned in a coarse-to-fine manner. Experimental results demonstrate that our method outperforms previous methods on the SceneFlow, KITTI 2012, and KITTI 2015 datasets.
Kai Zeng 0010, Yaonan Wang 0001, Qing Zhu 0003, Jianxu Mao, Hui Zhang 0023
IEEE Trans. Intell. Transp. Syst.3
2020 A Surface Defect Detection Framework for Glass Bottle Bottom Using Visual Attention Model and Wavelet Transform
abstract
Glass bottles must be thoroughly inspected before they are used for packaging. However, the vision inspection of bottle bottoms for defects remains a challenging task in quality control due to inaccurate localization, the difficulty in detecting defects in the texture region, and the intrinsically nonuniform brightness across the central panel. To overcome these problems, we propose a surface defect detection framework, which is composed of three main parts. First, a new localization method named entropy rate superpixel circle detection (ERSCD), which combines least-squares circle detection and entropy rate superpixel (ERS) with an improved randomized circle detection, is proposed to accurately obtain the region of interest (ROI) of the bottle bottom. Then, according to the structure-property, the ROI is divided into two measurement regions: central panel region and annular texture region. For the former, a defect detection method named frequency-tuned anisotropic diffusion super-pixel segmentation (FTADSP) that integrates frequency-tuned salient region detection (FT), anisotropic diffusion, and an improved superpixel segmentation is proposed to precisely detect the regions and boundaries of defects. For the latter, a defect detection strategy called wavelet transform multiscale filtering (WTMF) based on a wavelet transform and a multiscale filtering algorithm is proposed to reduce the influence of texture and to improve the robustness to localization error. The proposed framework is tested on four data sets obtained by our designed vision system. The experimental results demonstrate that our framework achieves the best performance compared with many traditional methods.
Xianen Zhou, Yaonan Wang 0001, Qing Zhu 0003, Jianxu Mao, Changyan Xiao, Xiao Lu 0002, Hui Zhang 0023
IEEE Trans. Ind. Informatics3
2019 SSG: superpixel segmentation and GrabCut-based salient object segmentation
Xianen Zhou, Yaonan Wang 0001, Qing Zhu 0003, Changyan Xiao, Xiao Lu 0002
Vis. Comput.3