Fasheng Wang

dblp:79/5519 · DBLP profile ↗
← Back
54ranked-venue papers
16as first author
37since 2021 · last 2026
0000-0002-0946-0789ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 7 first-author · 15 since 2021Artificial intelligence and machine learning · 22 · 7 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 4 since 2021Computer networks · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Breaking Alignment Barriers: TPS-Driven Semantic Correlation Learning for Alignment-Free RGB-T Salient Object Detection
abstract
Existing RGB-T salient object detection methods predominantly rely on manually aligned and annotated datasets, struggling to handle real-world scenarios with raw, unaligned RGB-T image pairs. In practical applications, due to significant cross-modal disparities such as spatial misalignment, scale variations, and viewpoint shifts, the performance of current methods drastically deteriorates on unaligned datasets. To address this issue, we propose an efficient RGB-T SOD method for real-world unaligned image pairs, termed Thin-Plate Spline-driven Semantic Correlation Learning Network (TPS-SCL). We employ a dual-stream MobileViT as the encoder, combined with efficient Mamba scanning mechanisms, to effectively model correlations between the two modalities while maintaining low parameter counts and computational overhead. To suppress interference from redundant background information during alignment, we design a Semantic Correlation Constraint Module (SCCM) to hierarchically constrain salient features. Furthermore, we introduce a Thin-Plate Spline Alignment Module (TPSAM) to mitigate spatial discrepancies between modalities. Additionally, a Cross-Modal Correlation Module (CMCM) is incorporated to fully explore and integrate inter-modal dependencies, enhancing detection performance. Extensive experiments on various datasets demonstrate that TPS-SCL attains state-of-the-art (SOTA) performance among existing lightweight SOD methods and outperforms mainstream RGB-T SOD approaches.
Lupiao Hu, Fasheng Wang, Fangmei Chen, Fuming Sun
AAAI2
2026 Wild Animal Tracking with High-Quality Segment Anything Model and Domain Adaptation
Ganggang Huang, Fasheng Wang, Hanwei Li, Mingshu Zhang, Mengyin Wang, Fuming Sun
Int. J. Comput. Vis.2
2026 Modality Interaction Decoupling: Spatiotemporal Memory-Driven Unified Multimodal Tracking Framework
abstract
Current designs of multimodal tracking networks primarily focus on spatial feature interaction within the backbone and lack the exploitation of temporal information. Although some approaches incorporate temporal cues by introducing sequential information from adjacent frames or employ updated temporal features during the feature extraction stage, they struggle to capture dynamic object variations and motion information in complex scenarios. To address these limitations, this paper proposes a Spatiotemporal Memory-Driven Unified Multimodal Tracking Framework (SMMTrack). Unlike existing works relying on feature interaction paradigms within the backbone network, this study innovatively introduces a decoupled backbone feature extraction framework. It deploys parallel, independent Vision Transformer (ViT) networks dedicated to extracting information from RGB and X modalities (RGB-T, RGB-D, RGB-E), abandoning the conventional intra-backbone feature interaction. Furthermore, we introduce a memory mechanism during the feature extraction stage to enable long-term object modeling. In addition, a Long-term Memory storage and retrieval module is designed to dynamically update the Memory-list, thereby allowing the model to capture object appearance variations and motion trends comprehensively. SMMTrack is a unified framework across three tasks (RGB-T, RGB-D, and RGB-E tracking). Experimental results demonstrate that SMMTrack outperforms the state-of-the-art (SOTA) models, achieving outstanding performance in diverse multimodal tracking scenarios. Codes and results are released on https://github.com/qfxb/SMMTrack.
Fasheng Wang, Mengyin Wang, Fuming Sun
IEEE Internet Things J.3
2026 Infrared and visible image fusion based on multi-modal and multi-scale cross-compensation
Meitian Li, Jing Sun 0012, Fasheng Wang, Fuming Sun
Knowl. Based Syst.4
2026 Towards a prompt-driven framework with state space models for UAV object tracking
Ziqing Yan, Fasheng Wang, Fuming Sun
Knowl. Based Syst.4
2026 LESOD: Lightweight and efficient network for RGB-D salient object detection
Mingyu Zhong, Jing Sun 0012, Fasheng Wang, Fuming Sun
Pattern Recognit.3
2025 Mirror Feature-Aware Generative Adversarial Network for RGB-T Salient Object Detection
abstract
Existing RGB-T salient object detection methods often employ asymmetric feature processing mechanisms, which lead to limited inter-modal interaction and difficulties in achieving cross-modal semantic alignment. Moreover, due to the inherent differences in imaging mechanisms, modality conflicts caused by feature prior mismatches severely restrict the efficient exploitation of complementary information. To address these challenges, we propose a Mirror Feature-aware Generative Adversarial Network (MFAGAN). We introduce adversarial learning into the multi-modal feature fusion process, transforming the implicit feature alignment assumptions in traditional fusion methods into explicit distribution consistency constraints through the dynamic game mechanism between the generator and discriminator. Specifically, MFAGAN designs a triple collaborative optimization component: 1) A symmetric two-stage encoder achieves a dynamic balance between pixel-level details and semantic-level representations through bidirectional alternating guidance; 2) A cross-modal residual decoder employs independent parameter paths to preserve modality-specific characteristics and suppress fusion bias; 3) A feature difference complementation module adaptively integrates differential information from regions with confidence conflicts. We conduct extensive experiments on three public datasets. The experimental results show that the MFAGAN achieves better performance than the competing methods. Codes and results are released on https://github.com/asd291614761/MFAGAN.
Fangkai Zhao, Fangmei Chen, Fasheng Wang, Fuming Sun
ICIP4
2025 Adaptive Partial Momentum Hamiltonian Monte Carlo
Fangkai Zhao, Fangmei Chen, Fasheng Wang
PRCV (1)4
2025 Facial attribute editing via a Balanced Simple Attention Generative Adversarial Network
Fanghui Ren, Wenpeng Liu, Fasheng Wang, Bo Wang 0167, Fuming Sun
Expert Syst. Appl.3
2025 Asymmetric cross-modality interaction network for RGB-D salient object detection
Yiming Su, Mengyin Wang, Fasheng Wang
Expert Syst. Appl.4
2025 Highly Efficient RGB-D Salient Object Detection With Adaptive Fusion and Attention Regulation
abstract
Existing RGB-D salient object detection (SOD) models have large numbers of parameters, high computational complexity, and slow inference speeds, limiting their deployment on edge devices. To address this issue, we propose a highly efficient network (HENet), focusing on developing lightweight RGB-D SOD models. Specifically, to fairly handle multimodal inputs and capture long-range dependencies of features, we employ a dual-stream structure and use MobileViT as the network encoder. We introduce the Adaptive Edge-Aware Fusion Module (AEFM) that adaptively adjusts the contribution of features during the fusion process based on the amount of feature information, and perceives the edges of the fused features at the pixel level. To compensate for the insufficient feature extraction capability of the lightweight backbone network, we propose the Dual-Branch Feature Enhancement Module (DFEM) to enhance the representation capability of the fused features. Finally, we design the Feature Attention Regulation Module (FARM) to adjust the model’s focus in real time. HENet has fewer parameters (11.9M) and lower computational complexity (10.7 GFLOPs), achieving an inference speed of 121 FPS for images with size$384\times 384$. Extensive experiments are conducted on seven challenging RGB-D SOD datasets. The experimental results demonstrate that HENet outperforms 16 state-of-the-art methods and shows great potential in downstream computer vision tasks. Codes and results are available onhttps://github.com/BojueGao/HENet.
Fasheng Wang, Mengyin Wang, Fuming Sun
IEEE Trans. Circuits Syst. Video Technol.2
2025 Rethinking How to Capture Long-Range Dependency in 3D Object Detection
abstract
LiDAR-based 3D object detection is essential for autonomous driving. Existing high-performance 3D object detectors usually design complex structures in the 3D backbone to capture long-range dependencies among features. However, introducing these complex structures into the 3D backbone significantly increases computational cost and inference latency, limiting the efficiency and feasibility of detectors in practical applications. In this work, we rethink the long-range dependency capturing problem from a new perspective, that is transferring this task from 3D backbone to 2D feature space. To accomplish this goal, we propose a Long-Range Dense Feature Capture Network (LDFCNet). LDFCNet retains the basic structure of the 3D backbone to extract preliminary 3D features but shifts the complex long-range dependency capturing task to be processed on a 2D dense feature map, thereby enhancing the detection performance while reducing the computational cost. Importantly, a robust 2D dense feature capture (2D-DFC) backbone is devised to effectively and efficiently capture the long-range dependencies. In addition, we introduce a re-parameterization technique to decouple the training and inference of the 2D backbone, further reducing inference latency. We conduct extensive experiments on the Waymo Open and nuScenes datasets and the experimental results show that LDFCNet demonstrates competitive performance. Notably, LDFCNet is$1.5\times $faster than the state-of-the-art hybrid detector HEDNet and$2.1\times $faster than the transformer-based detector DSVT. Codes and results are released onhttps://github.com/asd291614761/LDFCNet.
Fasheng Wang, Mengyin Wang, Fuming Sun
IEEE Trans. Circuits Syst. Video Technol.2
2025 ORSIDiff: Diffusion Model for Salient Object Detection in Optical Remote Sensing Images
abstract
The unique imaging conditions of satellites introduce significant uncertainties in the structure and scale of ground objects, presenting a major challenge for Optical Remote Sensing Image Salient Object Detection (ORSI-SOD). Current ORSI-SOD methods often fail to effectively differentiate between salient objects and subtle background variations, leading to suboptimal prediction outcomes. Furthermore, ORSI-SOD is a dense pixel prediction task, and existing approaches frequently depend on pixel-level probabilities, which can result in overconfident and inaccurate predictions. To address these challenges, we reformulate the ORSI-SOD task as a mask-generation problem by introducing a novel paradigm and propose a diffusion model-based method for ORSI-SOD, termed ORSIDiff. Central to our approach is the design of a powerful denoising network that enhances the model’s refinement capabilities. This network leverages the strengths of both global and local modeling, improving the handling of salient object details and enabling a deeper understanding of the distinctions between salient objects and their surroundings. Additionally, we introduce a consistency assessment strategy that aggregates multiple potential predictions during the denoising process, effectively mitigating the issue of overconfident point estimation. Extensive experimental results on two widely used ORSI-SOD datasets demonstrate that ORSIDiff achieves significant performance improvements over 20 state-of-the-art methods.
Jinyu Han, Jing Sun 0012, Fasheng Wang, Fuming Sun
IEEE Trans. Geosci. Remote. Sens.3
2025 Lightweight Edge-Aware Mamba-Fusion Network for Weakly Supervised Salient Object Detection in Optical Remote Sensing Images
abstract
Despite the significant progress made in fully supervised salient object detection in optical remote sensing images (ORSI-SOD), these methods rely heavily on pixel-level annotations, which are time-consuming and labor-intensive. This situation has driven the development of weakly supervised ORSI-SOD methods. However, existing weakly supervised ORSI-SOD methods still face excessive model parameters and high computational complexity, hindering their flexibility and deployment in edge devices. To address these challenges, we propose the LightEMNet, a scribble-based, lightweight, and high-performance edge-aware network for ORSI-SOD. The network employs MobileNetV2 as its lightweight encoder backbone. To mitigate the suboptimal feature extraction performance caused by the lightweight architecture, we design a feature refinement layer (FRL) to refine the features extracted from the backbone, thereby generating guidance information while achieving better structural awareness and object localization. To realize better detail optimization, we introduce edge information extracted by a multiscale edge perception module (MEP) to regulate high-level features. Finally, considering the shortcomings of traditional convolution in global-awareness, we propose a Mamba-based cross-scale edge-semantic interaction (CESI) module to achieve efficient alignment of semantics and edges, which consequently enhances the representation consistency of the fused features and improves the model’s adaptability to complex scenes. We verify the effectiveness of the LightEMNet through extensive experiments. The results demonstrate that the proposed LightEMNet exhibits competitive detection performance with only 4.81 M parameters. Codes and results are available athttps://github.com/xingggao/LightEMNet
Gaojie Xing, Mengyin Wang, Fasheng Wang, Fuming Sun
IEEE Trans. Geosci. Remote. Sens.3
2025 Spatial-Frequency Collaborative Learning for Camouflaged Object Detection
abstract
Camouflaged object detection (COD) is a challenging task that struggles to accurately detect the objects concealed in the surrounding environment. This is largely attributed to the intrinsic similarity of the camouflaged objects with the surrounding environment. To address this challenge, we propose a Spatial-Frequency Collaborative Learning network for COD (SFCNet). Specifically, we propose a Domain Transformation Fusion (DTF) module to handle the similarity between the camouflaged objects and the background, because when processed in the frequency domain, the features of the camouflaged object and the background become easy to discriminate. Then, we design a Cross-domain Integration Unit (CIU) to integrate the high-level features progressively through a Spatial-Frequency Coordinated Fusion (SFCF) module and a Multi-scale Feature Enhancement (MFE) module. Finally, the low-level features are combined with the high-level features from different decoding stages to correct the camouflaged objects in detail. In addition, an Edge Amplification (EA) module is designed to enable the model to pay attention to the global contour of the camouflaged object. It can facilitate the generation of prediction maps with accurate object boundaries. Extensive experiments on four benchmark COD datasets show that SFCNet outperforms state-of-the-art (SOTA) COD models. Meanwhile, it also has the characteristics of low parameters (21.01 M) and low computational complexity (24.14 G). Codes and results are released onhttps://github.com/Zhaorui328/SFCNet.
Mengyin Wang, Fasheng Wang, Fuming Sun
IEEE Trans. Multim.3
2024 OmniStyleGAN for Style-Guided Image-to-Image Translation
Qianyi Zhao, Mengyin Wang, Fasheng Wang, Fuming Sun
PRCV (11)4
2024 Unsupervised image-to-image translation with multiscale attention generative adversarial network
Fasheng Wang, Qianyi Zhao, Mengyin Wang, Fuming Sun
Appl. Intell.1
2024 Image deblurring method based on self-attention and residual wavelet transform
Bing Zhang 0023, Jing Sun 0012, Fuming Sun, Fasheng Wang
Expert Syst. Appl.4
2024 Cross-Modal Fusion and Progressive Decoding Network for RGB-D Salient Object Detection
Xihang Hu, Fuming Sun, Jing Sun 0012, Fasheng Wang
Int. J. Comput. Vis.4
2024 MAGNet: Multi-scale Awareness and Global fusion Network for RGB-D salient object detection
Mingyu Zhong, Jing Sun 0012, Fasheng Wang, Fuming Sun
Knowl. Based Syst.4
2024 Siamese Tracking Network with Multi-attention Mechanism
abstract
Object trackers based on Siamese networks view tracking as a similarity-matching process. However, the correlation operation operates as a local linear matching process, limiting the tracker’s ability to capture the intricate nonlinear relationship between the template and search region branches. Moreover, most trackers don’t update the template and often use the first frame of an image as the initial template, which will easily lead to poor tracking performance of the algorithm when facing instances of deformation, scale variation, and occlusion of the tracking target. To this end, we propose a Simases tracking network with a multi-attention mechanism, including a template branch and a search branch. To adapt to changes in target appearance, we integrate dynamic templates and multi-attention mechanisms in the template branch to obtain more effective feature representation by fusing the features of initial templates and dynamic templates. To enhance the robustness of the tracking model, we utilize a multi-attention mechanism in the search branch that shares weights with the template branch to obtain multi-scale feature representation by fusing search region features at different scales. In addition, we design a lightweight and simple feature fusion mechanism, in which the Transformer encoder structure is utilized to fuse the information of the template area and search area, and the dynamic template is updated online based on confidence. Experimental results on publicly tracking datasets show that the proposed method achieves competitive results compared to several state-of-the-art trackers.
Yuzhuo Xu, Ting Li 0002, Fasheng Wang, Fuming Sun
Neural Process. Lett.4
2024 Efficient Camouflaged Object Detection Network Based on Global Localization Perception and Local Guidance Refinement
abstract
Camouflaged Object Detection (COD) is a challenging visual task due to its complex contour, diverse scales, and high similarity to the background. Existing COD methods encounter two predicaments: One is that they are prone to falling into local perception, resulting in inaccurate object localization; Another issue is the difficulty in achieving precise object segmentation due to a lack of detailed information. In addition, most COD methods typically require larger parameter amounts and higher computational complexity in pursuit of better performance. To this end, we propose a global localization perception and local guidance refinement network (PRNet), that simultaneously addresses performance and computational costs. Through effective aggregation and use of semantic and details information, the PRNet can achieve accurate localization and refined segmentation of camouflaged objects. Specifically, with the help of a Cascaded Attention Perceptron (CAP) designed, we can effectively integrate and perceive multi-scale information to localize camouflaged objects. We also design a Guided Refinement Decoder (GRD) in a top-down manner to extract context information and aggregate details to further refine camouflaged prediction results. Extensive experimental results demonstrate that our PRNet outperforms 12 state-of-the-art models on 4 challenging datasets. Meanwhile, the PRNet has a smaller number of parameters (12.74M), lower computational complexity (10.24G), and real-time inference speed (105FPS). Source codes are available at https://github.com/hu-xh/PRNet.
Xihang Hu, Xiaoli Zhang 0001, Fasheng Wang, Jing Sun 0012, Fuming Sun
IEEE Trans. Circuits Syst. Video Technol.3
2024 Temporal Context and Environment-Aware Correlation Filter for UAV Object Tracking
abstract
In this article, we propose a temporal context and environment-aware correlation filter (CF) for unmanned aerial vehicle (UAV) object tracking. First, we exploit environmental residuals between two adjacent frames to enhance the discrimination ability and insensitivity of the tracker in complex tracking environments. Second, the historical filter model is utilized to build a temporal regularization term to prevent the filter degradation induced by the severe object appearance variations that affect the tracking performance while suppressing the boundary effect. Finally, we propose to use the predicted object position in the next frame to efficiently sample a predicted context patch for constructing the prediction context-aware regularization term, improving the ability of the tracker to cope with interference from unknown environment changes. Extensive experimental results on four challenging benchmarks show that our tracker achieves state-of-the-art (SOTA) performance compared with other advanced CF-based UAV trackers with real-time tracking speed (47 frames/s) on a single CPU.
Fasheng Wang, Fuming Sun
IEEE Trans. Geosci. Remote. Sens.2
2024 Low-light image enhancement using transformer with color fusion and channel attention
Yinbang Sun, Jing Sun 0012, Fuming Sun, Fasheng Wang
J. Supercomput.4
2024 CATNet: A Cascaded and Aggregated Transformer Network for RGB-D Salient Object Detection
abstract
Salient object detection (SOD) is an important preprocessing operation for various computer vision tasks. Most of existing RGB-D SOD models employ additive or connected strategies to directly aggregate and decode multi-scale features to predict salient maps. However, due to the large differences between the features of different scales, these aggregation strategies adopted may lead to information loss or redundancy, and few methods explicitly consider how to establish connections between features at different scales in the decoding process, which consequently deteriorates the detection performance of the models. To this end, we propose a cascaded and aggregated Transformer Network (CATNet) which consists of three key modules, i.e., attention feature enhancement module (AFEM), cross-modal fusion module (CMFM) and cascaded correction decoder (CCD). Specifically, the AFEM is designed on the basis of atrous spatial pyramid pooling to obtain multi-scale semantic information and global context information in high-level features through dilated convolution and multi-head self-attention mechanism, enhancing high-level features. The role of the CMFM is to enhance and thereafter fuse the RGB features and depth features, alleviating the problem of poor-quality depth maps. The CCD is composed of two subdecoders in a cascading fashion. It is designed to suppress noise in low-level features and mitigate the differences between features at different scales. Moreover, the CCD uses a feedback mechanism to correct and repair the output of the subdecoder by exploiting supervised features, so that the problem of information loss caused by the upsampling operation during the multi-scale features aggregation process can be mitigated. Extensive experimental results demonstrate that the proposed CATNet achieves superior performance over 14 state-of-the-art RGB-D methods on 7 challenging benchmarks.
Fuming Sun, Bowen Yin, Fasheng Wang
IEEE Trans. Multim.4
2024 Heterogeneous Fusion and Integrity Learning Network for RGB-D Salient Object Detection
abstract
While significant progress has been made in recent years in the field of salient object detection, there are still limitations in heterogeneous modality fusion and salient feature integrity learning. The former is primarily attributed to a paucity of attention from researchers to the fusion of cross-scale information between different modalities during processing multi-modal heterogeneous data, coupled with an absence of methods for adaptive control of their respective contributions. The latter constraint stems from the shortcomings in existing approaches concerning the prediction of salient region’s integrity. To address these problems, we propose a Heterogeneous Fusion and Integrity Learning Network for RGB-D Salient Object Detection (HFIL-Net). In response to the first challenge, we design an Advanced Semantic Guidance Aggregation (ASGA) module, which utilizes three fusion blocks to achieve the aggregation of three types of information: within-scale cross-modal, within-modal cross-scale, and cross-modal cross-scale. In addition, we embed the local fusion factor matrices in the ASGA module and utilize the global fusion factor matrices in the Multi-modal Information Adaptive Fusion module to control the contributions adaptively from different perspectives during the fusion process. For the second issue, we introduce the Feature Integrity Learning and Refinement Module. It leverages the idea of ”part-whole” relationships from capsule networks to learn feature integrity and further refine the learned features through attention mechanisms. Extensive experimental results demonstrate that our proposed HFIL-Net outperforms over 17 state-of-the-art detection methods in testing across seven challenging standard datasets. Codes and results are available on https://github.com/BojueGao/HFIL-Net .
Haorao Gao, Yiming Su, Fasheng Wang
ACM Trans. Multim. Comput. Commun. Appl.3
2024 Attention-guided Multi-modality Interaction Network for RGB-D Salient Object Detection
abstract
The past decade has witnessed great progress in RGB-D salient object detection (SOD). However, there are two bottlenecks that limit its further development. The first one is low-quality depth maps. Most existing methods directly use raw depth maps to perform detection, but low-quality depth images can bring negative impacts to the detection performance. Hence, it is not desirable to utilize depth maps indiscriminately. The other one is how to effectively predict salient maps with clear boundary and complete salient region. To address these problems, an Attention-Guided Multi-Modality Interaction Network (AMINet) is proposed. First, we propose a new quality enhancement strategy for unreliable depth images, named D epth E nhancement M odule ( DEM ). With respect to the second issue, we propose C ross- M odality A ttention M odule ( CMAM ) to rapidly locate salient region. The B oundary- A ware M odule ( BAM ) is designed to utilize high-level feature to guide the low-level feature generation in a top-down way to make up for the dilution of the boundary. To further improve the accuracy, we propose A trous R efined B lock ( ARB ) to adaptively compensate for the shortcoming of atrous convolution. By integrating these interactive modules, features from depth and RGB streams can be refined efficiently, which consequently boosts the detection performance. Experimental results demonstrate the proposed AMINet exceeds state-of-the-art (SOTA) methods on several public RGB-D datasets.
Fasheng Wang, Yiming Su, Jing Sun 0012, Fuming Sun
ACM Trans. Multim. Comput. Commun. Appl.2
2023 DCMNet: Discriminant and cross-modality network for RGB-D salient object detection
Fasheng Wang, Fuming Sun
Expert Syst. Appl.1
2023 WATB: Wild Animal Tracking Benchmark
Fasheng Wang, Fuming Sun
Int. J. Comput. Vis.1
2023 Attention-guided graph convolutional network for multi-behavior recommendation
Xingchen Peng, Jing Sun 0012, Mingshi Yan, Fuming Sun, Fasheng Wang
Knowl. Based Syst.5
2023 Style transfer network for complex multi-stroke text
Fangmei Chen, Fasheng Wang, Fuming Sun
Multim. Syst.4
2023 Deblurring transformer tracking with conditional cross-attention
Fuming Sun, Fasheng Wang
Multim. Syst.5
2023 SiamADT: Siamese Attention and Deformable Features Fusion Network for Visual Object Tracking
Fasheng Wang, Fuming Sun
Neural Process. Lett.1
2022 Learning Saliency-Aware Correlation Filters for Visual Tracking
abstract
Abstract Recently, visual object tracking has become a hot topic in computer vision community owing to its extensive applications and fundamental research significance. Many researchers have paid attention to the discriminative correlation filter (DCF) due to its excellent tracking performance. The background-aware CF (BACF) is developed to handle the inevitable boundary effects of DCFs that shows superior performance compared with other CF-based trackers. However, BACF has a poor occlusion handling ability. To overcome this defect, we propose a saliency-aware CF tracking (SACF) model. In SACF, a saliency map of the target is introduced on CFs to strengthen the ability to extract the target from a complex background. Meanwhile, to adjust to the rapid variation of the target appearance, we adaptively modify the learning rate of the filter template using adaptive learning rate. By incorporating the saliency map into the BACF objective function and adaptive learning rate adjustment, the robustness and accuracy of the tracker are boosted significantly. Extensive experiments validate that our method can efficaciously solve the occlusion problem and achieves state-of-the-art performance in terms of accuracy and speed on standard benchmarks, i.e. OTB-2013, OTB-2015 and LaSOT. Compared with other state-of-the-art CF tracking methods, the proposed method is competitive considering accuracy and success rate.
Yanbo Wang 0003, Fasheng Wang, Fuming Sun
Comput. J.2
2022 Aggregate interactive learning for RGB-D salient object detection
Jingyu Wu, Fuming Sun, Fasheng Wang
Expert Syst. Appl.5
2022 AMTSet: a benchmark for abrupt motion tracking
Fasheng Wang, Shuangshuang Yin, Fuming Sun, Junxing Zhang
Multim. Tools Appl.1
2022 Context and saliency aware correlation filter for visual tracking
Fasheng Wang, Shuangshuang Yin, Jimmy T. Mbelwa, Fuming Sun
Multim. Tools Appl.1
2020 Visual object tracking via iterative ant particle filtering
abstract
Visual object tracking remains a challenging task in computer vision although important progress has been made in the past decades. Particle filter (PF) is now a standard framework for solving non‐linear/non‐Gaussian problems, especially in visual object tracking. This study proposes an ant colony optimisation (ACO)‐based iterative PF for object tracking. In the proposed method, the basic idea of ACO is used to simulate the behaviour of a particle moving toward the posterior distribution. Such idea is incorporated into the particle filtering framework in order to overcome the well‐known particle impoverishment problem. An iterative unscented Kalman filter is used to design a proposal distribution for particle generation in order to generate better predicted sample states. For the likelihood model, the authors adopt the locality sensitive histogram to model the appearance of the target object, which can better handle the illumination variation during tracking. The experimental results demonstrate that the proposed tracker shows better performance than the other tracking methods.
Fasheng Wang, Yanbo Wang 0003, Fuming Sun, Xucheng Li, Junxing Zhang
IET Image Process.1
2020 Locality-constrained affine subspace coding for image classification and retrieval
Bingbing Zhang 0001, Qilong Wang 0001, Xiaoxiao Lu, Fasheng Wang, Peihua Li
Pattern Recognit.4
2020 Visual tracking tracker via object proposals and co-trained kernelized correlation filters
Jimmy T. Mbelwa, Qingjie Zhao, Fasheng Wang
Vis. Comput.3
2019 Objectness-based smoothing stochastic sampling and coherence approximate nearest neighbor for visual tracking
Jimmy T. Mbelwa, Qingjie Zhao, Yao Lu 0001, Fasheng Wang, Mercy Mbise
Vis. Comput.5
2018 Object tracking using Langevin Monte Carlo particle filter and locality sensitive histogram based likelihood model
Fasheng Wang, Baowei Lin, Junxing Zhang, Xucheng Li
Comput. Graph.1
2017 An ant particle filter for visual tracking
abstract
Sequential Monte Carlo method (also named as particle filter) is now a standard framework for solving nonlinear/non-Gaussian problems, especially in computer vision fields. This paper proposes an ant colony optimization (ACO) based iterative particle filter for visual tracking. In the proposed tracking method, the basic idea of ACO is used to simulate the behavior of particle moving toward the posterior density. Such idea is incorporated into the particle filtering framework in order to overcome the well-known problem of particle impoverishment. We design an iterative proposal distribution for particle generation in order to generate better predicted sample states. The experimental results demonstrate that the proposed tracker shows better performance than the other trackers.
Fasheng Wang, Baowei Lin, Xucheng Li
ICIS1
2017 Ordered over-relaxation based Langevin Monte Carlo sampling for visual tracking
Fasheng Wang, Peihua Li, Xucheng Li, Mingyu Lu
Neurocomputing1
2017 Boundary points based scale invariant 3D point feature
Baowei Lin, Fasheng Wang, Yi Sun 0009, Wen Qu
J. Vis. Commun. Image Represent.2
2017 Adaptive Hamiltonian MCMC sampling for robust visual tracking
Fasheng Wang, Xucheng Li, Mingyu Lu
Multim. Tools Appl.1
2015 SIPF: Scale invariant point feature for 3D point clouds
abstract
In this paper, we propose a method for detecting Scale-Invariant Point Feature(SIPF) including 3D keypoints Detector and feature descriptor. To detect SIPF, we first estimate a keyscale for point cloud, and calculate the covariance matrix of each 3D point. Keypoints are the saliency points who have a fast change speed along with all principal directions. Then the descriptors are encoded based on the shape of a border or silhouette of an object to be detected or recognized. Experimental results with the Stanford datasets demonstrate that the proposed method can be effectively used for 3D point clouds expression.
Baowei Lin, Fangda Zhao, Toru Tamaki, Fasheng Wang, Le Xiao
ICIP4
2014 Robust Abrupt Motion Tracking via Adaptive Hamiltonian Monte Carlo Sampling
Fasheng Wang, Xucheng Li, Mingyu Lu, Zhi-Bo Xiao
PRICAI1
2014 Robust particle tracker via Markov Chain Monte Carlo posterior sampling
Fasheng Wang, Mingyu Lu
Multim. Tools Appl.1
2013 Improving Particle Filter with Better Proposal Distribution for Nonlinear Filtering Problems
Fasheng Wang, Xucheng Li, Mingyu Lu
WASA1
2013 Efficient Visual Tracking via Hamiltonian Monte Carlo Markov Chain
abstract
Efficient visual tracking is a challenging task in the computer vision community due to its large motion uncertainty induced by occlusion, abrupt motion or appearance changes. In this paper, we propose a Hamiltonian Markov Chain Monte Carlo (MCMC) based tracking scheme for efficient tracking within the Bayesian filtering framework, aiming at handling full or partial occlusions, abrupt motion and appearance changes. In this tracking scheme, no complex models are built for motion uncertainties. The object states are augmented by introducing a momentum item and the Hamiltonian dynamics (HD) is integrated into the traditional MCMC-based tracking method. A new object state is proposed by computing a trajectory according to HD, implemented with the Leapfrog method. The new state can be distant from the current object state but, nevertheless, has a high acceptance probability, which consequently bypasses the slow exploration of the state space suffered by traditional random-walk proposal distribution. In addition, the proposed tracking algorithm can avoid being trapped in local maxima, which is suffered by conventional MCMC-based tracking algorithms. Experimental results reveal that our approach is efficient and effective in dealing with various types of tracking scenarios compared with several alternatives.
Fasheng Wang, Mingyu Lu
Comput. J.1
2012 Hamiltonian Monte Carlo estimator for abrupt motion tracking
Fasheng Wang, Mingyu Lu
ICPR1
2009 Financial Options Pricing Using the MKPF Algorithm
abstract
A mixture Kalman Particle Filter (MKPF) based options pricing method is proposed. The MKPF algorithm uses the unscented Kalman filter (UKF) and the extended Kalman filter (EKF) as proposal distribution to generate the importance sampling density. Each particle is firstly updated by the UKF and obtains a state estimation. Thereafter, this estimation is used as the prior of the EKF, in which the particle is updated again to gain the final estimation of the state. We use the classical B-S model in the experiment aiming at evaluating the performance of the newly proposed method and other existing algorithms. The experimental results show that the MKPF outperforms other algorithms.
Yingbo Zhang, Fasheng Wang, Yuejin Lin
SERA2
2006 Using Neural Network Technique in Vision-based Robot Curve Tracking
abstract
Robot curve tracking is needed in some industrial applications, such as automatic welding or incising. Such a robot system is usually equipped with visual sensors, which always require calibration before used. The calibration process is often complicated. In this paper a neural network is used to learn the relationship between the world coordinate information and the image information, instead of computing accurate camera parameters. The neural network is first trained based on sample data by using the 2D and 3D coordinates of some control points on a standard pattern. During the tracking stage, images captured by cameras are firstly changed into binary images. The curve is then thinned and its position on the image is recorded. From the image data, the curve's position in the world coordinate frame can be specified by using the trained neural network. The curve tracked can be arbitrary, open or closed. The experimental results illuminate that the neural network technique is satisfying and it is successfully used in the vision-based robot curve tracking
Qingjie Zhao, Fasheng Wang, Zengqi Sun
IROS2