Shengke Wang

dblp:33/3518 · DBLP profile ↗
← Back
32ranked-venue papers
9as first author
15since 2021 · last 2026
0000-0002-4906-8773ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 12 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Click-based interactive image segmentation with global hints and local corrections
abstract
The aim of click-based interactive image segmentation is to obtain pixel-level segmentation masks by only a small number of manual clicks. This approach streamlines the process of pixel-level annotation and image editing. Much research has focused on this area. In particular, SimpleClick, which utilizes Vision Transformers, has implemented a straightforward design that has demonstrated its effectiveness in interactive image segmentation. However, two issues remain: First, this simple design does not fully utilize the global guidance provided by the interaction map. Second, treating all clicks equally fails to maximize the benefits of the new clicks’ guidance in each iteration. To address these challenges, we propose a novel interactive image segmentation network called Glclick. Specifically, we introduce a Global Hint Module (GHM) that integrates global information from clicks into the transformer backbone. In addition, Glclick incorporates a Local Correction Module (LCM) that performs local optimization on the target masks generated by the backbone network. Extensive experiments on four generic datasets and three medical datasets demonstrate the superiority and generalizability of Glclick. • Designed a brand-new interactive image segmentation model called “Glclick”. • Designed a Global Hint Module called GHM and a Local Correction Module called LCM. • Advanced performance was achieved on generic datasets and three medical datasets.
Shengke Wang, Bajin Cai, Xiandong Wang, Fengqin Yao
Eng. Appl. Artif. Intell.1
2026 Beyond Semantics: Multiscale Interaction Network for Referring Camouflaged Object Detection
abstract
Referring camouflaged object detection (Ref-COD) is an emerging and challenging task that aims to localize camouflaged objects in complex scenes based on a small set of referring images with salient objects. However, existing methods primarily focus on semantic alignment between the referring and camouflaged objects while overlooking scale discrepancies, leading to under-response when small references guide large objects and over-response when large references guide small ones. To overcome this limitation, we propose a novel Multi-scale Interaction Network (MINet), explicitly designed to handle feature interactions across different scales in Ref-COD. MINet begins with a Dual-Source Fusion Block (DSFB) for semantic fusion between the referring and camouflaged features. Then, the Intra-scale Interaction Block (IIB) enhances local saliency within each scale by modeling contextual importance. Next, the Cross-scale Interaction Block (CIB) performs offset-guided alignment to bridge spatial gaps in multiscale feature fusion. Finally, the Cross-scale Aggregation Decoder (CAD) integrates multiscale features, effectively decoding the aggregated information to produce accurate predictions. Extensive experiments on Ref-COD datasets demonstrate that our method achieves state-of-the-art performance, highlighting the importance of scale interaction in Ref-COD.
Xiandong Wang, Tianqi Guo, Fengqin Yao, Shengke Wang, Junyu Dong, Guoqiang Zhong 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 GBNet: Gated Boundary-Aware Network for Camouflaged Object Detection
abstract
Camouflaged object detection involves identifying camouflaged objects visually blended into the surroundings, holding crucial significance in various visual applications. Existing methods primarily focus on leveraging boundary information to enhance camouflaged object detection. However, they often overlook the background interference near the object boundaries, which leads to coarse boundary predictions and results in suboptimal detection performance. In this paper, to address this problem, we propose GBNet, a gated boundary-aware network designed to enhance boundary precision and improve overall detection performance. Specifically, GBNet incorporates a boundary-enhanced module that selectively filters extraneous background information through a boundary gate block, ensuring the generation of high-quality boundary information. Additionally, a boundary-aware decoder is designed to enrich the representation ability of the decoder by injecting high-quality boundary features and aggregating contextual features. With meticulous design, GBNet excels in accurately segmenting camouflaged objects in challenging scenarios. Extensive experiments demonstrate that GBNet outperforms 19 state-of-the-art methods significantly across four widely-used benchmark datasets. The source code is publicly available at https://github.com/wooownn/GBNet.
Xiandong Wang, Fengqin Yao, Guoqiang Zhong 0001, Shengke Wang, James T. Kwok
IEEE Trans. Image Process.5
2025 Prototype Guided Multi-Scale Class Aggregation for Generalized Few-Shot Semantic Segmentation
abstract
Generalized Few-Shot Semantic Segmentation (GFSS) enables pixel-level segmentation of various classes in images when limited annotated data is available. This effectively reduces the reliance on extensive fine-grained annotated data for training segmentation models. The training process for GFSS is typically divided into two stages: base class training and novel class updating. However, due to the imbalance in the distribution between novel and base classes, this two-stage training approach may result in the misidentification of similar classes and segmentation loss caused by background interference. To address these challenges, this paper proposes a prototype guided multi-scale class aggregation method. The multi-scale design preserves more image details, while the prototype-guided strategy leverages the rich features learned from both base and novel class prototypes to guide class aggregation. This approach enhances the understanding of inter-class differences through aggregation and supports the effective learning of novel classes. Extensive experiments confirm the efficacy of our approach, with our model achieving state-of-the-art performance on COCO-20iand outperforming existing methods on most metrics in PASCAL-5i.
Shengke Wang
ICME5
2025 Multi-model Probability Information Fusion For Semi-Supervised Semantic Segmentation
Xiaohui Ye, Fengqin Yao, Lian Chen, Shengke Wang
PRCV (8)5
2025 An Edge-Guided SAM for effective complex object segmentation
Longyi Chen, Xiandong Wang, Fengqin Yao, Mingchen Song, Jiaheng Zhang, Shengke Wang
Expert Syst. Appl.6
2025 G2LNet: Global to local information communication for camouflaged object detection
Shengke Wang
Expert Syst. Appl.6
2024 SelfLoc: High Quality Unsupervised Object Localization with Self-Prompt SAM
Jiaheng Zhang, Xiandong Wang, Conghui Li, Longyi Chen, Shengke Wang
PRCV (12)5
2024 Underwater object detection in noisy imbalanced datasets
Long Chen 0019, Tengyue Li, Andy Zhou, Shengke Wang, Junyu Dong, Huiyu Zhou 0001
Pattern Recognit.4
2024 DewaterGAN: A Physics-Guided Unsupervised Image Water Removal for UAV in Coastal Zone
abstract
Accurate recognition of marine species in drone-captured images is essential in maintaining the stability of coastal zone ecosystems. Unmanned aerial vehicle (UAV) remote sensing images usually lack paired supervised signals and suffer from color distortion and blurring due to the interaction of ambient light with cross-medium transmission between air and water. However, current algorithms mainly focus on supervised training methods and also ignore the interaction involved in the cross-medium transmission of light in water. In this article, for UAV in coastal zones, we propose an unsupervised image water removal model, named DewaterGAN, which is based solely on low-tide and high-tide images without paired supervised signals and also preserves color and texture in the water removal process. Specifically, our approach involves two key steps: an unsupervised training CycleGAN network accomplishes domain transitions from low-tide level to high-tide level, and a physics-based attention module guides image water removal and maintains authenticity. Additionally, we utilize evaluation metrics of image restoration peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) to quantitatively analyze the performance of the model. We also employed several non-reference metrics (UIQM, UCIQE, NIQE, BRISQUE, LIQE, ILNIQE, and CLIPIQA) to evaluate the visual quality of the image de-watering process. Extensive experiments conducted on both our water removal dataset and public datasets validate the efficacy of our model. The code is athttps://github.com/yfq-yy/Dewater.git.
Fengqin Yao, Fuzhi Tang, Xiandong Wang, Shengke Wang, Guoqiang Zhong 0001, Jingfeng Zhang
IEEE Trans. Geosci. Remote. Sens.5
2023 Camouflaged Object Detection via Global-Edge Context and Mixed-Scale Refinement
Qilun Li, Fengqin Yao, Xiandong Wang, Shengke Wang
PRCV (7)4
2023 Video Object Counting With Scene-Aware Multi-Object Tracking
abstract
The critical challenge of video object counting is to avoid counting the same object multiple times in different frames. By comparing the appearance and motion feature information of the detection results, the authors use the multi-object tracking method to assign an independent ID number to each object. From the time the ID tag is obtained until the end of the video, each object is counted only once. However, even minor amounts of image noise can cause irreversible changes in feature information, resulting in severe tracking drifts. This paper introduces the concept of scene awareness and addresses unreasonable ID assignment caused by unreliable feature matching in the context of region division. Through the macro analysis of the scene, the authors define the region (called the transition region) where the number of objects can increase or decrease and require that all ID assignments for new objects and ID deletions for existing objects take place only in the transition region. Because the actual number of objects in the non-transition region is constant, they rematch unmatched objects with existing IDs in the region (called ID relocation) because changes in object ID are caused by feature matching failure. In this paper, the authors create algorithms for dynamically generating transition regions, detecting object increases and decreases, and relocating object IDs. Experimental results show that the method effectively improves the accuracy of video object counting.
Yongdong Li, Liang Qu, Guiyan Cai, Guoan Cheng, Yuling Dou, Fengqin Yao, Shengke Wang
J. Database Manag.8
2023 Lightweight network learning with Zero-Shot Neural Architecture Search for UAV images
Fengqin Yao, Shengke Wang, Laihui Ding, Guoqiang Zhong 0001, Leon Bevan Bullock, Junyu Dong
Knowl. Based Syst.2
2022 An accurate box localization method based on rotated-RPN with weighted edge attention for bin picking
Fengqin Yao, Shengke Wang, Long Chen 0019, Feng Gao 0005, Junyu Dong
Neurocomputing2
2022 SWIPENET: Object detection in noisy underwater scenes
abstract
Deep learning based object detection methods have achieved promising performance in controlled environments. However, these methods lack sufficient capabilities to handle underwater object detection due to these challenges: (1) images in the underwater datasets and real applications are blurry whilst accompanying severe noise that confuses the detectors and (2) objects in real applications are usually small. In this paper, we propose a Sample-WeIghted hyPEr Network (SWIPENET), and a novel training paradigm named Curriculum Multi-Class Adaboost (CMA), to address these two problems at the same time. Firstly, the backbone of SWIPENET produces multiple high resolution and semantic-rich Hyper Feature Maps, which significantly improve small object detection. Secondly, inspired by the human education process that drives the learning from easy to hard concepts, we propose the noise-robust CMA training paradigm that learns the clean data first and then move on to learns the diverse noisy data. Experiments on four underwater object detection datasets show that the proposed SWIPENET+CMA framework achieves better or competitive accuracy in object detection against several state-of-the-art approaches.
Long Chen 0019, Feixiang Zhou, Shengke Wang, Junyu Dong, Ning Li 0012, Haiping Ma, Xin Wang 0068, Huiyu Zhou 0001
Pattern Recognit.3
2020 Underwater object detection using Invert Multi-Class Adaboost with deep learning
abstract
In recent years, deep learning based methods have achieved promising performance in standard object detection. However, these methods lack sufficient capabilities to handle underwater object detection due to these challenges: (1) Objects in real applications are usually small and their images are blurry, and (2) images in the underwater datasets and real applications accompany heterogeneous noise. To address these two problems, we first propose a novel neural network architecture, namely Sample-WeIghted hyPEr Network (SWIPENet), for small object detection. SWIPENet consists of high resolution and semantic-rich Hyper Feature Maps which can significantly improve small object detection accuracy. In addition, we propose a novel sample-weighted loss function which can model sample weights for SWIPENet, which uses a novel sample re-weighting algorithm, namely Invert Multi-Class Adaboost (IMA), to reduce the influence of noise on the proposed SWIPENet. Experiments on two underwater robot picking contest datasets URPC2017 and URPC2018 show that the proposed SWIPENet+IMA framework achieves better performance in detection accuracy against several state-of-the-art object detection approaches.
Long Chen 0019, Zheheng Jiang, Shengke Wang, Junyu Dong, Huiyu Zhou 0001
IJCNN5
2019 Accurate Ulva prolifera regions extraction of UAV images with superpixel and CNNs for ocean environment monitoring
Shengke Wang, Liang Qu, Changyin Yu, Yujuan Sun, Feng Gao 0005, Junyu Dong
Neurocomputing1
2019 Learning spatiotemporal representations for human fall detection in surveillance video
Yongqiang Kong, Zhengang Wei, Shengke Wang
J. Vis. Commun. Image Represent.5
2019 Transferred Deep Learning for Sea Ice Change Detection From Synthetic-Aperture Radar Images
abstract
High-quality sea ice monitoring is crucial to navigation safety and climate research in the polar regions. In this letter, a transferred multilevel fusion network (MLFN) is proposed for sea ice change detection from synthetic-aperture radar (SAR) images. Considering the fact that training data are limited in the task of sea ice change detection, a large data set was used to train the MLFN, and the deep knowledge can be transferred to sea ice analysis. In addition, cascade dense blocks are employed to optimize the convolutional layers. Multilayer feature fusion is introduced to exploit the complementary information among low-, mid-, and high-level feature representations. Therefore, more discriminative feature extraction can be achieved by the MLFN. Furthermore, the fine-tune strategy is utilized to optimize the network parameters. The experimental results on two real sea ice data sets demonstrated that the proposed method achieved better performance than other competitive methods.
Yunhao Gao, Feng Gao 0005, Junyu Dong, Shengke Wang
IEEE Geosci. Remote. Sens. Lett.4
2019 Sea Ice Change Detection in SAR Images Based on Convolutional-Wavelet Neural Networks
abstract
Sea ice change detection from synthetic aperture radar (SAR) images can be regarded as a classification procedure, in which pixels are classified into changed and unchanged classes. However, existing methods usually suffer from the intrinsic speckle noise of multitemporal SAR images. To solve the problem, this letter presents a change detection method based on convolutional-wavelet neural networks (CWNNs). In CWNN, dual-tree complex wavelet transform is introduced into convolutional neural networks for changed and unchanged pixels' classification, and then, the effect of speckle noise is effectively reduced. In addition, a virtual sample generation scheme is employed to create samples for CWNN training, and the problem of limited samples is alleviated. Experimental results on two real SAR image data sets demonstrate the effectiveness and robustness of the proposed method.
Feng Gao 0005, Yunhao Gao, Junyu Dong, Shengke Wang
IEEE Geosci. Remote. Sens. Lett.5
2019 Cascaded one-vs-rest detection network for fine-grained recognition without part annotations
Long Chen 0019, Shengke Wang, Kin-Man Lam 0001, Huiyu Zhou 0001, Muwei Jian, Junyu Dong
Multim. Tools Appl.2
2018 Sea Ice Change Detection in SAR Images Based on Collaborative Representation
abstract
Sea ice change detection from synthetic aperture radar (SAR) images is important for navigation safety and natural resource extraction. This paper proposed a sea ice change detection method from SAR images based on collaborative representation. First, neighborhood-based ratio is used to generate a difference image (DI). Then, some reliable samples are selected from the DI by hierarchical fuzzy C-means (FCM) clustering. Finally, based upon these samples, collaborative representation method is utilized to classify pixels from the original SAR images into unchanged and changed class. From there, the final change map can be obtained. Experimental results on two real sea ice datasets demonstrate the superiority of the proposed method over two closely related methods.
Yunhao Gao, Feng Gao 0005, Junyu Dong, Shengke Wang
IGARSS4
2018 Sea Ice Classification from Hyperspectral Images Based on Self-Paced Boost Learning
abstract
Hyperspectral imagery has evident advantages for sea ice classification due to enormous spectral bands. In this paper, we proposed a novel sea ice classification framework from hyperspectral image based on self-paced boost learning (SPBL). First, the criterion of linear prediction error is used for unsupervised band selection. Then, local binary pattern (LBP) features are extracted from the selected bands. Finally, SPBL is employed as the classifier to provide probability outputs using the extracted features. The proposed framework can capture the intrinsic inter-class discriminative models while ensuring the reliability of the samples involved in learning. The experimental results in real-world dataset demonstrate that the proposed framework is superior to several closely related methods.
Feng Gao 0005, Junyu Dong, Shengke Wang
IGARSS5
2018 Asymmetric filtering-based dense convolutional neural network for person re-identification combined with Joint Bayesian and re-ranking
Shengke Wang, Long Chen 0019, Huiyu Zhou 0001, Junyu Dong
J. Vis. Commun. Image Represent.1
2018 Specular reflection removal of ocean surface remote sensing images from UAVs
Shengke Wang, Changyin Yu, Yujuan Sun, Feng Gao 0005, Junyu Dong
Multim. Tools Appl.1
2018 An improved genetic algorithm for three-dimensional reconstruction from a single uniform texture image
Yujuan Sun, Xiaofeng Zhang 0003, Muwei Jian, Shengke Wang, Zeju Wu, Qingtang Su, Beijing Chen
Soft Comput.4
2017 Person re-identification with deep dense feature representation and Joint Bayesian
abstract
Person re-identification that aims at matching individuals across multiple camera views has become indispensable in intelligent video surveillance systems. It remains challenging due to the large variations of pose, illumination, occlusion and camera viewpoint. Feature representation and metric learning are the two fundamental components in person reidentification. In this paper, we present a Special Dense Convolutional Neural Network (SD-CNN) to extract the feature and apply Joint Bayesian to measure the similarity of pedestrian image pairs. The SD-CNN can preserve more horizontal information to against viewpoint changes, maximize the feature reuse and ensure feature distributing discriminative. Joint Bayesian models the extracted feature representation as the sum of inter- and intra-personal variations, and the joint probability of two images being a same person can be obtained through log-likelihood ratio. Experiments show that our approach significantly outperforms state-of-the-art methods on several benchmarks of person re-identification.
Shengke Wang, Lianghua Duan, Junyu Dong
ICIP1
2016 Ocean internal waves features extraction by analysis of aerial oblique photography
abstract
Internal waves are a widespread geophysical phenomenon in stratified fluids and studying internal features in the coastal ocean is an important task. Using satellite imagery for studying oceanic internal waves is very popular and studying the low altitude aerial oblique photograph is a new direction. In this paper, we study the images captured from a circling aircraft which track a number of internal wave packets. The captured images are first rectified and photogram metrically mapped to ground coordinates, and then we use canny edge detector to find the internal wave propagation direction. The first several waves' ridges of each ground coordinated images are also exactly labeled and propagation speeds can be achieved. Experiment results show the performance of our algorithms.
Shengke Wang, Long Chen 0019, Jianping Yang, Muwei Jian, Lifang Lin, Junyu Dong
IGARSS1
2016 Fast pedestrian detection based on object proposals and HOG
abstract
Research on pedestrian detection still presents a lot of space for improvements, both on speed and detection accuracy. State-of-the-art object proposals approach has shown the very effective computational efficiency in object detection. In this paper, we present a framework for pedestrian detection based on the object proposals. Instead of scaling the test image to different sizes, we generate a pyramid of HOG templates by simply scaling the primary HOG descriptor. For the test image, Histogram of Oriented Gradient feature vector is calculated only in the region of object proposals without using sliding window, and the corresponding HOG template is selected according to the size of proposal region. Experiment results show that our proposed methods can achieve the equal detection rate while using only 20% time.
Shengke Wang, Lianghua Duan, Long Chen 0019, Guoan Cheng, Jingai Yu
IJCNN1
2016 Person re-identification based on deep spatio-temporal features and transfer learning
abstract
Person re-identification has become a hot research topic due to its importance in surveillance and forensics applications. The purpose of person re-identification is to find the same person from disjoint camera views at different time. Most of the existing methods try to identify the person by measuring the similarity of two still images from different camera views, which only uses intraimage features such as color, shape and texture. In this paper, we propose a person re-identification architecture which analyzes a sequence pair while not an image pair, so not only intra-image features but also the gait feature is also considered. In contrast to existing works that use handcrafted features, our method automatically learns spatio-temporal features optimal for the person re-identification task with a deep convolutional network. To learn a discriminant metric, we use a subspace and metric learning method called Cross-view Quadratic Discriminant Analysis (XQDA).The process of XQDA is also called transfer learning. Experiments show that our method significantly outperforms the state of the art on both a large dataset (CUHK03) and a medium-sized dataset (CUHK01). We also get a better performance on a small dataset (VIPeR) with a pre-trained network without fine-tuning.
Shengke Wang, Lianghua Duan, Long Chen 0019
IJCNN1
2016 Human fall detection in surveillance video based on PCANet
Shengke Wang, Long Chen 0019, Zixi Zhou, Xin Sun 0003, Junyu Dong
Multim. Tools Appl.1
2007 Fast Detection of Vehicles Based-on the Moving Region
abstract
The precise location and tracking of the moving vehicles is an important part in intelligent transportation system. This paper presents a new method for detecting vehicles which can remove the shadows fast. The main process is done in three steps: moving region detection, shadow detection, and edge detection. Firstly, a moving region of the image is achieved using the self-adaptive background updating method quickly. Then a coarse shadow area can be taken with the model-based method, and the shadow detection method based on the HSV color space is only applied on the coarse area. Finally, impose edge detection on both the moving regions and shadow areas. Subtraction of the two areas leads to the result of the detection of real vehicle. Experiment results show that this method can improve the efficiency of the moving vehicles detection greatly.
Zongshun Ma, Zhenghua Fang, Shengke Wang
CAD/Graphics4