Wangsheng Yu

dblp:132/4843 · DBLP profile ↗
← Back
31ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0002-0968-8951ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021
YearPublicationVenuePosition
2026 Visual tracking method with hybrid spatio-temporal backbone network and dual-memory mechanism
Junyi Dong, Xianxin Jia, Sugang Ma, Yang Liu 0116, Wangsheng Yu
Expert Syst. Appl.6
2026 Video object segmentation based on feature compression and attention correction
Jiale Dong, Chenxu Wang 0012, Sugang Ma, Wangsheng Yu
Signal Process. Image Commun.5
2025 Improved UAV Aerial Vehicle Detection Algorithm Based on YOLOv11n
Wangsheng Yu, Xianxiang Qin, Jinling Han, Sugang Ma
ICIG (2)2
2025 SEDNet: Real-Time Semantic Segmentation Algorithm Based on STDC
abstract
Recently, deep convolutional neural networks (DCNN) have been widely used in semantic segmentation tasks and have achieved high segmentation accuracy. However, most algorithms based on DCNN have high computational complexity, making them unsuitable for real‐time segmentation. To solve this problem, this paper proposes a real‐time semantic segmentation algorithm based on the STDC network. The algorithm adopts an “encoder–decoder” embedded in a U‐shaped architecture to realize real‐time segmentation while maintaining high accuracy. Following the encoder, a mixed pooling attention module is designed to expand the receptive field, enhancing the network model’s learning ability in complex scenarios. Then, a feature fusion module is used for combining features from different stages, and channel attention based on atrous convolution is employed to expand the receptive field and avoid dimensionality reduction learning. Finally, a Tversky‐based detail loss function is used to encode more spatial details. The proposed algorithm was extensively tested on the challenging Cityscapes and CamVid datasets, and the experimental results showed that the proposed algorithm obtained 76.4% and 72.8% of mIoU, respectively. Meanwhile, our algorithm achieves 105.2 FPS and 165.6 FPS inference speed with a single NVIDIA GTX 1080Ti GPU, meeting the real‐time segmentation requirements. The proposed algorithm can conduct real‐time segmentation while maintaining high accuracy, achieving a good balance between accuracy and speed.
Sugang Ma, Wangsheng Yu, Xiangmo Zhao
Int. J. Intell. Syst.4
2025 Lightweight video object segmentation: Integrating online knowledge distillation for fast segmentation
Chenxu Wang 0012, Sugang Ma, Jiale Dong, Yunchen Wang, Wangsheng Yu
Knowl. Based Syst.6
2024 A Global Re-detection Method Based on Feature Interaction Siamese Network
Ruoxue Han, Chentao Liu, Sugang Ma, Wangsheng Yu, Yunchen Wang
PRCV (12)5
2024 Multi-object tracking algorithm based on interactive attention network and adaptive trajectory reconnection
Sugang Ma, Shuaipeng Duan, Wangsheng Yu, Lei Pu, Xiangmo Zhao
Expert Syst. Appl.4
2024 SOCF: A correlation filter for real-time UAV tracking based on spatial disturbance suppression and object saliency-aware
Sugang Ma, Bo Zhao 0035, Wangsheng Yu, Lei Pu, Xiaobao Yang 0001
Expert Syst. Appl.4
2024 Dual-branch network object detection algorithm based on dual-modality fusion of visible and infrared images
Sugang Ma, Wangsheng Yu, Yunchen Wang
Multim. Syst.5
2023 Multi-template global re-detection based on Gumbel-Softmax in long-term visual tracking
Jingyuan Ma, Wangsheng Yu, Zhilong Yang, Sugang Ma, JiuLun Fan 0001
Appl. Intell.3
2023 Robust Visual Object Tracking Based on Feature Channel Weighting and Game Theory
abstract
Although the discriminative correlation filter‐ (DCF)‐based tracker improves tracking performance, some object representation issues can still be further optimized. On the one hand, the DCF tracker’s deep convolutional features contain many noisy channels, and assigning the same weights to multiple channels cannot distinguish the importance of different channels. On the other hand, a simple weighted fusion approach cannot fully utilize the benefits of different feature types. We propose a visual object tracking algorithm based on adaptive channel weighting and feature game fusion to solve these problems. In this study, an adaptive channel weighting strategy is designed to assign suitable weights to each channel based on the average energy ratio of the target and background regions in the feature channels and prune the channels with low weights to improve feature robustness and reduce computational complexity. Simultaneously, the game theory concept is introduced in the multifeature fusion. The handcrafted features are combined with shallow and deep convolutional features according to feature complementarity. Then, the two combined features are seen as two sides of the game, continuously gamed during the tracking process to generate a feature model with a higher representation capacity. Extensive experiments are conducted on four mainstream visual tracking benchmark datasets, including OTB2015, VOT2018, LaSOT, and UAV123. The experimental results show that the proposed algorithm performs outstandingly compared to the state‐of‐the‐art trackers.
Sugang Ma, Bo Zhao 0035, Wangsheng Yu, Lei Pu, Lei Zhang 0166
Int. J. Intell. Syst.4
2023 Object drift determination network based on dual-template joint decision-making in long-term visual tracking
Sugang Ma, Wangsheng Yu, JiuLun Fan 0001
J. Vis. Commun. Image Represent.5
2021 SiamDA: Dual attention Siamese network for real-time visual tracking
Lei Pu, Xinxi Feng, Wangsheng Yu, Yufei Zha
Signal Process. Image Commun.4
2021 Superpixel-Oriented Classification of PolSAR Images Using Complex-Valued Convolutional Neural Network Driven by Hybrid Data
abstract
Recently, convolutional neural networks (CNNs) have been successfully developed and used in the classification of polarimetric synthetic aperture radar (PolSAR) images. However, they often suffer from some problems, such as time-consuming, unsatisfactory detail-preservation, and bad effectiveness given limited training samples. Focusing on these problems, we propose a complex-valued CNN (CV-CNN)-based algorithm for PolSAR image classification in this article. On the one hand, a superpixel-oriented (SPO) scheme is employed to reduce the computational cost of the algorithm and preserve image details simultaneously, which takes superpixels instead of single pixels as classification units. In particular, to meet the input requirement of CV-CNN, three alternative methods of superpixel regularization are designed and compared. On the other hand, considering that both measured data (MD) and manually designed polarimetric features (PFs) have their own advantages, the hybrid data (HD) combining them is employed to drive CV-CNN, which is helpful to improve the effectiveness of the algorithm. We perform experiments on three actual PolSAR image data sets acquired by AIRSAR and Radarsat-2 systems as well as a semisimulated data set. The experimental results demonstrate that, compared to conventional pixel-oriented methods, the proposed SPO scheme is much more time-efficient and is also beneficial to detail preservation. Moreover, the CV-CNN driven by HD generally obtains consistently better classification results than that driven by pure MD or manually designed PFs.
Xianxiang Qin, Huanxin Zou, Wangsheng Yu, Peng Wang 0018
IEEE Trans. Geosci. Remote. Sens.3
2020 MHASiam: Mixed High-Order Attention Siamese Network for Real-Time Visual Tracking
Lei Pu, Xinxi Feng, Wangsheng Yu, Yufei Zha, Zhiqiang Jiao
PRCV (2)4
2019 Deep Correlation Filter based Real-Time Tracker
Lei Pu, Xinxi Feng, Wangsheng Yu, Yufei Zha, Sugang Ma
FUSION4
2019 Polsar Image Classification via Complex-Valued Convolutional Neural Network Combining Measured Data and Artificial Features
abstract
Recently, many deep convolutional neural networks (CNNs) have been developed for the polarimetric synthetic aperture radar (PolSAR) image classification. For them, it is often a hard task to obtain good results with limited training samples. To release this problem, some strategies, such as the feature-driven method that takes artificial features as the input of CNN, have been proposed. However, since artificial features are usually difficult to be universal, their applicability is limited to a certain extent. For this issue, this paper proposes a scheme of combining measured data and artificial features with a complex-valued CNN (CV-CNN). In our algorithm, not only the measured PolSAR data but also some discriminative artificial features are employed as the input of CV-CNN. The basic idea is that the measured data contains the fully acquired information of targets, while the artificial features include expert knowledge. Therefore, by fusing them, better and more stable performance may be obtained. The experiments performed on both actual and simulated PolSAR images have validated the effectiveness of the proposed algorithm.
Xianxiang Qin, Huanxin Zou, Wangsheng Yu, Peng Wang 0018
IGARSS4
2019 Visual tracking based on semantic and similarity learning
abstract
We present a method by combining the similarity and semantic features of a target to improve tracking performance in video sequences. Trackers based on Siamese networks have achieved success in recent competitions and databases through learning similarity according to binary labels. Unfortunately, such weak labels result in limiting the discriminative ability of the learned feature, thus it is difficult to identify the target itself from the distractors that have the same class. The authors observe that the inter‐class semantic features benefit to increase the separation between the target and the background, even distractors. Therefore, they proposed a network architecture which uses both similarity and semantic branches to obtain more discriminative features for locating the target accuracy in new frames. The large‐scale ImageNet VID dataset is employed to train the network. Even in the presence of background clutter, visual distortion, and distractors, the proposed method still maintains following the target. They test their method with the open benchmarks OTB and UAV123. The results show that their combined approach significantly improves the tracking ability relative to trackers using similarity or semantic features alone.
Yufei Zha, Zhuling Qiu, Wangsheng Yu
IET Comput. Vis.4
2019 Online Scale Adaptive Visual Tracking Based on Multilayer Convolutional Features
abstract
Convolutional neural networks can efficiently exploit sophisticated hierarchical features which have different properties for visual tracking problem. In this paper, by using multilayer convolutional features jointly and constructing a scale pyramid, we propose an online scale adaptive tracking method. We construct two separate correlation filters for translation and scale estimations. The translation filters improve the accuracy of target localization by a weighted fusion of multiple convolutional layers. Meanwhile, the separate scale filters achieve the optimal and fast scale estimation by a scale pyramid. This design decreases the mutual errors of translation and scale estimations, and reduces computational complexity efficiently. Moreover, in order to solve the problem of tracking drifts due to the severe occlusion or serious appearance changes of the target, we present a new adaptive and selective update mechanism to update the translation filters effectively. Extensive experimental results show that our proposed method achieves the excellent overall performance compared with the state-of-the-art methods.
Xin Wang 0026, Wangsheng Yu, Zefenfen Jin, Yufei Zha, Xianxiang Qin
IEEE Trans. Cybern.3
2018 Edge Detection of Polsar Images Using Statistical Distance Between Automatically Refined Samples
abstract
Region-based edge detectors are popular for edge extraction of polarimetric synthetic aperture radar (PolSAR) images, which, however, often suffer from the heterogeneous data and outliers. In this paper, an improved edge detector with a scheme of refining samples automatically is proposed. Firstly, dominant scattering mechanisms of PolSAR data are acquired by using the Freeman-Durden decomposition. Then, for each pixel, samples in the regions predicted by edge detector filter are refined according to their dominant scattering mechanisms and power. Furthermore, for filters of different orientations, statistical distances between the refined samples in two regions are calculated, and the maximum is assigned to be the corresponding edge intensity. The experiments performed on both simulated and actual PolSAR images demonstrate that the proposed approach is more robust to outliers than the classical algorithms.
Xianxiang Qin, Wangsheng Yu, Peng Wang 0018, Huanxin Zou
IGARSS3
2018 Target tracking approach via quantum genetic algorithm
abstract
Aiming at an efficient feature match and similarity search in visual tracking, this study proposes a tracking algorithm based on quantum genetic algorithm. Therein, the global optimisation ability of quantum genetic algorithm is utilised. In the framework of quantum genetic algorithm, the positions of pixels are taken as individuals in population, while scale‐invariant feature transform and colour features are taken as target model. Via defining the objective function, individual's fitness values can be measured. Visual tracking is realised when the pixel point with the biggest fitness value is searched and its corresponding position is returned. The experiment results show that the tracking algorithm the authors proposed performs more efficiently when it is compared with the state‐of‐the‐art tracking algorithms.
Zefenfen Jin, Wangsheng Yu, Xin Wang 0026
IET Comput. Vis.3
2018 Visual tracking via ensemble autoencoder
abstract
The authors present a novel online visual tracking algorithm via ensemble autoencoder (AE). In contrast to other existing deep model based trackers, the proposed algorithm is based on the theory that the image resolution has an influence on vision procedures. When the authors employ a deep neural network to represent the object, the resolution is corresponding to the network size. The authors apply a small network to represent the pattern in a relatively lower resolution and search the object in a relatively larger area of the neighbourhood. After roughly estimating the location of the object, the authors apply a large network, which can provide more detailed information, to estimate the state of the object more accurately. Thus, the authors employ a small AE mainly for position searching and a larger one mainly for scale estimating. When tracking an object, the two networks interact to operate under the framework of particle filtering. Extensive experiments on the benchmark dataset show that the proposed algorithm performs favourably compared with some state‐of‐the‐art methods.
Bo Dai 0017, Wangsheng Yu, Xin Wang 0026, Zefenfen Jin
IET Image Process.3
2018 Robust occlusion-aware part-based visual tracking with object scale adaptation
Xin Wang 0026, Wangsheng Yu, Lei Pu, Zefenfen Jin, Xianxiang Qin
Pattern Recognit.3
2017 Online Fast Deep Learning Tracker Based on Deep Sparse Neural Networks
Xin Wang 0026, Wangsheng Yu, Zefenfen Jin
ICIG (1)3
2017 Robust mean shift tracking based on refined appearance model and online update
Wangsheng Yu, Peng Wang 0018
Multim. Tools Appl.1
2015 A Scale Adaptive Tracking Algorithm Based on Kernel Ridge Regression and Fast Fourier Transform
Lang Zhang, Wangsheng Yu, Wanjun Xu
ICIG (1)3
2015 Multi-scale mean shift tracking
abstract
In this study, a three‐dimensional mean shift tracking algorithm, which combines the multi‐scale model and background weighted spatial histogram, is proposed to address the problem of scale estimation under the framework of mean shift tracking. The target template is modelled with multi‐scale model and described with three‐dimensional spatial histogram. The tracking algorithm is implemented by three‐dimensional mean shift iteration, which translates the problem of scale estimation in two‐dimensional image plane into the localisation in three‐dimensional image space. To enhance the robustness, the background weighted histogram is employed to suppress the background information in the target candidate model. Firstly, the multi‐scale model and three‐dimensional spatial histogram are introduced to represent the target template. Then, the three‐dimensional mean shift iteration formulation is derived based on the similarity measure between the target model and the target candidate model. Finally, a multi‐scale mean shift tracking algorithm combining multi‐scale model and background weighted spatial histogram is proposed. The proposed algorithm is evaluated on some challenging sequences which contain scale changed targets and other complex appearance variations in comparison with three representative mean shift based tracking algorithms. Both the qualitative results and quantitative analysis indicate that the proposed algorithm outperforms the referenced algorithms in both tracking precision and scale estimation.
Wangsheng Yu, Xiaohua Tian, Yufei Zha
IET Comput. Vis.1
2014 Robust visual tracking based on watershed regions
abstract
Robust visual tracking is a very challenging problem especially when the target undergoes large appearance variation. In this study, the authors propose an efficient and effective tracker based on watershed regions. As middle‐level visual cues, watershed regions contain more semantics information than low‐level features, and reflect more structure information than high‐level model. First, the authors manually select the target template in initial frame, and predict the target candidate in the next frame using motion prediction. Then, the authors utilise marker‐based watershed algorithm to obtain the watershed regions of target template and candidate template, and describe each region with multiple features. Next, the authors calculate the nearest neighbour in feature space to match the watershed regions and construct an affine relation from target template to candidate template. Finally, the authors resolve the affine relation to calculate the final tracking result, and update the template for the following tracking. The authors test their tracker on some challenging sequences with appearance variation range from illumination change, partial occlusion, pose change to background clutters and compare it with some state‐of‐the‐art works. Experiment results indicate that the proposed tracker is robust to the large appearance variation and exceeds the state‐of‐the‐art trackers in most situations.
Wangsheng Yu, Xiaohua Tian, Yufei Zha
IET Comput. Vis.1
2013 Color Image Optical Flow Estimation Algorithm with Shadow Suppression
abstract
A new optical flow estimation algorithm for color image is proposed to overcome the influence of shadows and improve the accuracy of optical flow estimation. Its idea is to compute optical flow of the color invariant space and then fuse the result with optical flow of the RGB space. Firstly, it computes edge strength from color invariants of two adjacent frames to build three-channel color images, and then computes optical flow of the images we have built before, finally, fuses the result of color invariant optical flow with RGB image optical flow through L¡Þ norm. The experimental results demonstrate that the target regions detected by this method are more robust and accurate under shadow and illumination conditions. Compared with some methods proposed recently, it performs better.
Guojian Wei, Wangsheng Yu
ICIG4
2013 Object Tracking Based on Particle Filter with Multi-scale Mode
abstract
In this paper, we propose a particle filter based tracker using multiscale mode to solve the scale changeable object tracking problem. Firstly, we discuss the traditional tracker's deficiency in dealing with scale changeable object and propose the multiscale mode. Then, we design our tracker under the particle filter framework using the proposed mode and "many to one" searching strategy. Finally, we compare our tracker with some existing trackers by tracking test on some video clips, which contain three categories of object scale change. Simulation results indicate that the proposed tracker distinctly improves both the tracking precision and efficiency. Compared with the traditional particle filter tracker, our tracker needs fewer particles and obtains more precise location and scale estimation.
Wangsheng Yu, Xiaohua Tian, Guojian Wei
ICIG1
2012 Swift template matching based on equivalent histogram
Wangsheng Yu, Xiaohua Tian, Chongzhao Han
FUSION1