Yufei Zha

dblp:93/2803 · DBLP profile ↗
← Back
52ranked-venue papers
9as first author
34since 2021 · last 2026
0000-0001-5013-2501ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 7 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 2 first-author · 18 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FiMLink: Enhancing sparse feature matching via FFT-in-Mamba and dynamic learnable Fourier encoding
Guancheng Jia, Boxiong Sun, Yufei Zha, Peng Zhang 0005
Knowl. Based Syst.5
2026 MiA: A plug-and-play hyperparameter-free Mamba in Attention module for spatial-temporal consistent visual tracking
Guancheng Jia, Yufei Zha, Peng Zhang 0005
Pattern Recognit.4
2026 FIST: Flow-inspired spatio-temporal target state modeling for visual tracking
Boxiong Sun, Tonghao Han, Yufei Zha, Peng Zhang 0005
Pattern Recognit.4
2026 WmLSTM: A Plug-and-Play Window-Level mLSTM-Based Temporal Encoder for Robust Visual Tracking
abstract
Temporal context modeling constitutes a fundamental issue for robust visual tracking. However, existing approaches are plagued by a critical granularity trade-off: frame-level template update mechanisms inevitably introduce background redundancy due to global frame information aggregation, while token-level propagation mechanisms undermine inherent local spatial correlations via independent feature transmission. To resolve this challenge, we propose WmLSTM, a plug-and-play Window-level mLSTM-based temporal encoder that reconfigures the temporal modeling paradigm for visual tracking. First, window-centric modeling retains intra-window spatial correlations while adaptively suppressing background clutter. Second, we pioneer the application of mLSTM in visual tracking, exploiting its explicit memory architecture that outperforms implicit sequence modeling alternatives (e.g., Mamba). Third, our plug-and-play design enables seamless integration with state-of-the-art trackers with minor computational overhead. Extensive experiments on seven benchmark datasets validate that WmLSTMTrack achieves an excellent balance among accuracy, speed, and parameter efficiency, attaining state-of-the-art accuracy on five benchmarks, superior real-time speed (GPU: 201fps, CPU: 47fps), and compact model size (8.22 M parameters). Moreover, the WMLSTM module consistently enhances the performance of diverse trackers, e.g., real-time FERMT-256: +2.3 points SR75 on GOT-10k, non-real-time EVPTrack-224: +1.8 points P on LaSOText, with merely 30 training epochs. The source code is available at https://github.com/Xiaochen918/WmLSTM.
Guancheng Jia, Yufei Zha, Peng Zhang 0005
IEEE Trans. Circuits Syst. Video Technol.4
2026 PromptTrack: Streaming Spatial-Temporal Prompt Learning for RGB-T Tracking
abstract
This paper presents a novel video-level RGB-T tracking paradigm based on prompt learning, termed PromptTrack, which establishes dense spatial-temporal associations through cross-modal interactions. The method introduces streaming temporal prompts to capture continuous target dynamics (e.g., appearance changes and motion trajectories), while leveraging multimodal spatial prompts to utilize complementary RGB and thermal-infrared (TIR) features dynamically. By propagating temporal prompts through consecutive frames and integrating bidirectional spatial interactions between modalities, PromptTrack achieves superior tracking performance in complex scenarios such as occlusion, low illumination, and distractors. The proposed framework employs a unified multimodal encoder with spatial-temporal modeling via multimodal spatial prompt blocks, enabling efficient fusion of RGB-TIR features without requiring domain-specific structure modifications. Extensive experiments on three RGB-T benchmarks (LasHeR, RGBT210, RGBT234) demonstrate that PromptTrack achieves new state-of-the-art performance, with 76.2% in precision rate on LasHeR and outperforming existing methods by ++1.9% in precision rate and +0.5% in success rate on RGBT234. Notably, its modality-agnostic design facilitates seamless generalization to RGB-D and RGB-E tracking domains achieving new benchmarks on DepthTrack, VOT-RGBD2022, and VisEvent datasets.
Hangfei Li, Guangbin Liu, Yufei Zha, Peng Zhang 0005
IEEE Trans. Multim.4
2025 Audio-visual correspondences based joint learning for instrumental playing source separation
Peng Zhang 0005, Siliang Wang, Wei Huang 0013, Yufei Zha, Yanning Zhang 0001
Neurocomputing5
2025 LINR: A Plug-and-Play Local Implicit Neural Representation Module for Visual Object Tracking
abstract
Current one-stream trackers suffer from limitations in distinguishing targets from complex backgrounds owing to their uniform token division strategy. By treating all regions equally, these methods allocate inadequate attention to crucial target details while overemphasizing redundant background information. Consequently, their performance deteriorates significantly in scenarios involving similar distractors or background clutter. In this work, we propose a Local Implicit Neural Representation (LINR) module specifically designed for local fine-grained object modeling. It consists of two key modules: (1) Local Window Selection: Leveraging template-guided CNN-based cross-correlation, it accurately identify crucial target-relevant regions, reducing background information redundant and computation burden. (2) INR-based Window Refinement: Using implicit neural networks, it optimizes token density and spatial continuity to improve local fine-grained instance-level representations, facilitating the discriminative ability between the target and the background. Moreover, the LINR module exhibits three remarkable advantages as a generalized enhancement for visual tracking. Firstly, it is plug-and-play, seamlessly integrating into existing one-stream trackers, both non-real-time and real-time ones, without architectural modifications, achieving significant performance improvements. Secondly, it is highly portable since it does not introduce new loss functions, additional training strategies or data. Thirdly, it is efficiency-friendly, having minimal impact on model parameters and tracking speed,e.g., AQATrack-LINR increases only 1.9% of the parameters and reduces the tracking speed by only 6fps. We incorporate the LINR module into two non-real-time trackers, OSTrack based on ViT-B and AQATrack based on HiViT-B, and one real-time tracker, FERMT based on ViT-tiny, respectively. The resultant OSTrack-LINR, AQATrack-LINR, and FERMT-LINR achieve state-of-the-art performance across seven widely utilized datasets, such as TrackingNet, LaSOT, and NFS30. The source code is available at https://github.com/Xiaochen918/LINR.
Guancheng Jia, Yufei Zha, Peng Zhang 0005, Yanning Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 Flexible Temperature Parallel Distillation for Dense Object Detection: Make Response-Based Knowledge Distillation Great Again
abstract
Feature-based approaches have been the focal point of previous research on knowledge distillation (KD) for dense object detection. These methods employ feature imitation and result in competitive performance. Despite being able to achieve comparable performance in image recognition, response-based KD methods can not reach the same level in dense object detection. Inspired by improving distillation performance from two key aspects: where to distill and how to distill, in this paper, a parallel distillation (PD) is introduced to fully utilize the sophisticated detection head and transfer all the output responses from the teacher to the student efficiently. In particular, the proposed PD takes an important consideration of the specific location of distillation, which is crucial for effective knowledge transfer. Regarding the discrepancies in output responses between the localization branch and the classification branch, we propose a novel Dynamic Localization Temperature (DLT) module to enhance the precision of distilling localization information. As for the classification branch, a Classification Temperature-Free (CTF) module is also designed to increase the robustness of distillation in heterogeneous networks. By incorporating the DLT and CTF into the PD framework to avoid setting temperature values manually, the Flexible Temperature Parallel Distillation (FTPD) is proposed to achieve a state-of-the-art (SOTA) performance, which can also be further combined with mainstream feature-based methods for better results. In terms of accuracy and robustness with extensive experiments, the proposed FTPD outperforms other KD methods in the task of dense object detection.
Yaoye Song, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Yanning Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 Toward Unifying Saliency Transformer for Video Saliency Prediction and Detection
abstract
Video saliency prediction and detection are thriving research domains that enable computers to simulate the distribution of visual attention akin to how humans perceive dynamic scenes. While many approaches have crafted task-specific training paradigms for either video saliency prediction or video salient object detection tasks, few attention has been devoted to devising a generalized saliency modeling framework that seamlessly bridges both these distinct tasks. In this study, we introduce the Unified Saliency Transformer (UniST) framework, which comprehensively utilizes the essential attributes of video saliency prediction and video salient object detection. In addition to extracting representations of frame sequences, a saliency-aware transformer is designed to learn the spatio-temporal representations at progressively increased resolutions, while incorporating effective cross-scale saliency information to produce a robust representation. Furthermore, task-specific decoders are proposed to perform the final prediction for each task. To the best of our knowledge, this is the first work to explore the design of a unified framework for both saliency modeling tasks. Convincible experiments demonstrate that the proposed UniST achieves superior performance across eight challenging benchmarks for two tasks, outperforming other state-of-the-art methods in most metrics. The project page ishttps://junwenxiong.github.io/UniST.
Junwen Xiong, Chuanyue Li, Peng Zhang 0005, Yue Huo, Wei Huang 0013, Yufei Zha
IEEE Trans. Circuits Syst. Video Technol.7
2024 DiffSal: Joint Audio and Video Learning for Diffusion Saliency Prediction
abstract
Audio-visual saliency prediction can draw support from diverse modality complements, but further performance enhancement is still challenged by customized architectures as well as task-specific loss functions. In recent studies, denoising diffusion models have shown more promising in unifying task frameworks owing to their inherent ability of generalization. Following this motivation, a novel Diffusion architecture for generalized audio-visual Saliency prediction (DiffSal) is proposed in this work, which formulates the prediction problem as a conditional generative task of the saliency map by utilizing input audio and video as the conditions. Based on the spatiotemporal audio-visual features, an extra network Saliency-UNet is designed to perform multimodal attention modulation for progressive refinement of the ground-truth saliency map from the noisy map. Extensive experiments demonstrate that the proposed DiffSal can achieve excellent performance across six challenging audio-visual benchmarks, with an average relative improvement of 6.3% over the previous state-of-the-art results by six metrics. The project url is htt ps: //junwenxiong. github.io/DiffSal.
Junwen Xiong, Peng Zhang 0005, Tao You, Chuanyue Li, Wei Huang 0013, Yufei Zha
CVPR6
2024 Unidirectional Cross-Modal Fusion for RGB-T Tracking
abstract
The key issue of RGB-T tracking is to obtain an effective multimodal representation of targets by utilizing complementary RGB and TIR modality information. Previous methods of template fusion or bidirectional search-template interaction potentially diminish the target representation, resulting from noise information of both templates and search regions. Meanwhile, the direct fusion of sole search features without interacting with templates cannot fully utilize target-relevant contextual information. To mitigate these issues, we present UCTrack, which fuses complementary multimodal search features conditioned on undisturbed RGB and TIR template features. Specifically, we design a Unidirectional Cross-modal Fusion (UCF) module to effectively minimize the influence of background noise on templates by pruning the unnecessary template-to-search cross-modal interaction and to mutually enhance RGB and TIR search features with target-relevant information through multimodal spatial fusion. Furthermore, this module is seamlessly integrated into different layers of a ViT backbone to facilitate feature extraction and cross-modal fusion for RGB-T tracking. Benefiting from the UCF module, UCTrack can effectively and accurately represent multimodal target features without unnecessary template-to-search interaction flow and direct template fusion, making the first proposal of unidirectional cross-modal fusion paradigm for RGB-T tracking to our best knowledge. Extensive experiments on three popular RGB-T tracking benchmarks demonstrate that our method achieves state-of-the-art performance.
Hangfei Li, Yufei Zha, Peng Zhang 0005
ECAI3
2024 How does Layer Normalization improve Batch Normalization in self-supervised sound source localization?
Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Tao You, Yanning Zhang 0001
Neurocomputing4
2024 Enhancing small object tracking with reversible rescaling networks
Yufei Zha, Hangfei Li
Image Vis. Comput.1
2024 Closed-loop unified knowledge distillation for dense object detection
Yaoye Song, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Tao You, Yanning Zhang 0001
Pattern Recognit.4
2024 Auto Diagnosis of Parkinson's Disease Via a Deep Learning Model Based on Mixed Emotional Facial Expressions
abstract
Parkinson's disease (PD) is a common degenerative disease of the nervous system in the elderly. The early diagnosis of PD is very important for potential patients to receive prompt treatment and avoid the aggravation of the disease. Recent studies have found that PD patients always suffer from emotional expression disorder, thus forming the characteristics of "masked faces". Based on this, we thus propose an auto PD diagnosis method based on mixed emotional facial expressions in the paper. Specifically, the proposed method is cast into four steps: Firstly, we synthesize virtual face images containing six basic expressions (i.e., anger, disgust, fear, happiness, sadness, and surprise) via generative adversarial learning, in order to approximate the premorbid expressions of PD patients; Secondly, we design an effective screening scheme to assess the quality of the above synthesized facial expression images and then shortlist the high-quality ones; Thirdly, we train a deep feature extractor accompanied with a facial expression classifier based on the mixture of the original facial expression images of the PD patients, the high-quality synthesized facial expression images of PD patients, and the normal facial expression images from other public face datasets; Finally, with the well-trained deep feature extractor, we thus adopt it to extract the latent expression features for six facial expression images of a potential PD patient to conduct PD/non-PD prediction. To show real-world impacts, we also collected a new facial expression dataset of PD patients in collaboration with a hospital. Extensive experiments are conducted to validate the effectiveness of the proposed method for PD diagnosis and facial expression recognition.
Wei Huang 0013, Renjie Wan, Peng Zhang 0005, Yufei Zha
IEEE J. Biomed. Health Informatics5
2023 CASP-Net: Rethinking Video Saliency Prediction from an Audio-Visual Consistency Perceptual Perspective
abstract
Incorporating the audio stream enables Video Saliency Prediction (VSP) to imitate the selective attention mechanism of human brain. By focusing on the benefits of joint auditory and visual information, most VSP methods are capable of exploiting semantic correlation between vision and audio modalities but ignoring the negative effects due to the temporal inconsistency of audio-visual intrinsics. Inspired by the biological inconsistency-correction within multi-sensory information, in this study, a consistency-aware audio-visual saliency prediction network (CASP-Net) is proposed, which takes a comprehensive consideration of the audio-visual semantic interaction and consistent perception. In addition a two-stream encoder for elegant association between video frames and corresponding sound source, a novel consistency-aware predictive coding is also designed to improve the consistency within audio and visual representations iteratively. To further aggregate the multi-scale audio-visual information, a saliency decoder is introduced for the final saliency map generation. Substantial experiments demonstrate that the proposed CASP-Net outperforms the other state-of-the-art methods on six challenging audio-visual eye-tracking datasets. For a demo of our system please see our project webpage.
Junwen Xiong, Ganglai Wang, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Guangtao Zhai
CVPR5
2023 Induction Network: Audio-Visual Modality Gap-Bridging for Self-Supervised Sound Source Localization
abstract
Self-supervised sound source localization is usually challenged by the modality inconsistency. In recent studies, contrastive learning based strategies have shown promising to establish such a consistent correspondence between audio and sound sources in visual scenarios. Unfortunately, the insufficient attention to the heterogeneity influence in the different modality features still limits this scheme to be further improved, which also becomes the motivation of our work. In this study, an Induction Network is proposed to bridge the modality gap more effectively. By decoupling the gradients of visual and audio modalities, the discriminative visual representations of sound sources can be learned with the designed Induction Vector in a bootstrap manner, which also enables the audio modality to be aligned with the visual modality consistently. In addition to a visual weighted contrastive loss, an adaptive threshold selection strategy is introduced to enhance the robustness of the Induction Network. Substantial experiments conducted on SoundNet-Flickr and VGG-Sound Source datasets have demonstrated a superior performance compared to other state-of-the-art works in different challenging scenarios. The code is available at https://github.com/Tahy1/AVIN.
Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Tao You, Yanning Zhang 0001
ACM Multimedia4
2023 Efficient thermal infrared tracking with cross-modal compress distillation
Hangfei Li, Yufei Zha, Huanyu Li 0003, Peng Zhang 0005, Wei Huang 0013
Eng. Appl. Artif. Intell.2
2023 Object detection based on cortex hierarchical activation in border sensitive mechanism and classification-GIou joint representation
Yaoye Song, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Tao You, Yanning Zhang 0001
Pattern Recognit.4
2023 Conditional invertible image re-scaling
Yufei Zha, Peng Zhang 0005, Wei Huang 0013
Pattern Recognit.1
2023 Facial Expression Guided Diagnosis of Parkinson's Disease via High-Quality Data Augmentation
abstract
Parkinson's disease (PD) is a neurodegenerative disease which is prevalent among the elder population and severely affects the life quality of patients and their families. Therefore, it is important to conduct an early diagnosis for potential patients with PD, so as to promote prompt treatment and avoid the aggravation of the disease. Recently, the in-vitro PD diagnosis based on facial expressions has received increasing attention because of its distinguishability (i.e., PD patients always possess the characteristics of “masked face”) and affordability. However, the performance of the existing facial expression-based PD diagnosis approaches is limited by: 1) the small-scale training data on PD patients' facial expressions, and 2) the weak prediction model. To address these two problems, we propose a new facial expression guided PD diagnosis method based on high-quality training data augmentation and deep neural network prediction. Specifically, the proposed method consists of three stages: Firstly, we synthesize virtual facial expression images with 6 basic emotions (i.e., anger, disgust, fear, happiness, sadness, and surprise) based on multi-domain adversarial learning to approximate the premorbid expressions of PD patients. Secondly, we introduce three facial image quality assessment (FIQA) criteria to measure the quality of these synthesized facial expression images and design a fusion screening strategy that shortlists the high-quality ones to augment the training data. Finally, we train a deep neural network prediction model based on the original and synthesized high-quality facial expression images for PD diagnosis. To show real-world impacts and evaluate the proposed method under different facial expressions, we also create a (currently largest) multiple facial expressions-based PD face dataset in collaboration with a hospital. Extensive experiments are performed to demonstrate the effectiveness of the multi-domain adversarial learning-based facial expression synthesis and the fusion screening strategy, particularly the superior performance of the proposed method for PD diagnosis.
Wei Huang 0013, Yintao Zhou, Yiu-Ming Cheung, Peng Zhang 0005, Yufei Zha
IEEE Trans. Multim.5
2023 Look&listen: Multi-Modal Correlation Learning for Active Speaker Detection and Speech Enhancement
abstract
Active speaker detection and speech enhancement have become two increasingly attractive topics in audio-visual scenario understanding. According to their respective characteristics, the scheme of independently designed architecture has been widely used in correspondence to each single task. This may lead to the representation learned by the model being task-specific, and inevitably result in the lack of generalization ability of the feature based on multi-modal modeling. More recent studies have shown that establishing cross-modal relationship between auditory and visual stream is a promising solution for the challenge of audio-visual multi-task learning. Therefore, as a motivation to bridge the multi-modal associations in audio-visual tasks, a unified framework is proposed to achieve target speaker detection and speech enhancement with joint learning of audio-visual modeling in this study. With the assistance of audio-visual channels of videos in challenging real-world scenarios, the proposed method is able to exploit inherent correlations in both audio and visual signals, which is used to further anticipate and model the temporal audio-visual relationships across spatial-temporal space via a cross-modal conformer. In addition, a plug-and-play multi-modal layer normalization is introduced to alleviate the distribution misalignment of multi-modal features. Based on cross-modal circulant fusion, the proposed model is capable to learned all audio-visual representations in a holistic process. Substantial experiments demonstrate that the correlations between different modalities and the associations among diverse tasks can be learned by the optimized model more effectively. In comparison to other state-of-the-art works, the proposed work shows a superior performance for active speaker detection and audio-visual speech enhancement on three benchmark datasets, also with a favorable generalization in diverse challenges.
Junwen Xiong, Peng Zhang 0005, Lei Xie 0001, Wei Huang 0013, Yufei Zha
IEEE Trans. Multim.6
2022 Information Lossless Multi-modal Image Generation for RGB-T Tracking
Yufei Zha, Lichao Zhang 0001, Peng Zhang 0005, Lang Chen
PRCV (4)2
2022 Semantic-aware spatial regularization correlation filter for visual tracking
abstract
Abstract Correlation filters with convolutional neural network (CNN) features have been successfully applied to visual tracking owing to their impressive combined capability for object representation. Unfortunately, further performance improvement is limited due to unwanted boundary effects of the circular structure. In this work, through an in‐depth study of the features’ characteristics, the authors propose a novel tracking strategy to achieve simultaneous filter matching and regularization with CNN features when tracking is on the fly. With a feature decomposed transform matrix, a spatial semantic regularization is generated to reduce the boundary effect effectively during filter optimization. Before each output, the regularized filter is then back performed to match with the extracted features of a search region to find the optimum candidate. Specifically, the most important advantage of the proposed spatial semantic map is to initialize only in the first frame as all the other tracking strategies. Besides, the authors design a novel updating strategy to tackle the cases where the object is occluded or disappeared in the scene. At this time, the maximum of the map is small, even negative. A substantial experiment has been carried out on the popular benchmark tracking datasets; the reliable results have demonstrated that the authors’ method is able to outperform most of the state‐of‐the‐art tracking works in both accuracy and robustness.
Yufei Zha, Peng Zhang 0005, Lei Pu, Lichao Zhang 0001
IET Comput. Vis.1
2022 One-shot Video Graph Generation for Explainable Action Reasoning
Tao Zhuo, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Yanning Zhang 0001, Mohan Kankanhalli
Neurocomputing5
2022 A novel locally-constrained GAN-based ensemble to synthesize arterial spin labeling images
Wei Huang 0013, Mingyuan Luo, Jing Li 0027, Peng Zhang 0005, Yufei Zha
Inf. Sci.5
2022 Identity-Aware Facial Expression Recognition Via Deep Metric Learning Based on Synthesized Images
abstract
Person-dependent facial expression recognition has received considerable research attention in recent years. Unfortunately, different identities can adversely influence recognition accuracy, and the recognition task becomes challenging. Other adverse factors, including limited training data and improper measures of facial expressions, can further contribute to the above dilemma. To solve these problems, a novel identity-aware method is proposed in this study. Furthermore, this study also represents the first attempt to fulfill the challenging person-dependent facial expression recognition task based on deep metric learning and facial image synthesis techniques. Technically, a StarGAN is incorporated to synthesize facial images depicting different but complete basic emotions for each identity to augment the training data. Then, a deep-convolutional-neural-network-based network is employed to automatically extract latent features from both real facial images and all synthesized facial images. Next, a Mahalanobis metric network trained based on extracted latent features outputs a learned metric that measures facial expression differences between images, and the recognition task can thus be realized. Extensive experiments based on several well-known publicly available datasets are carried out in this study for performance evaluations. Person-dependent datasets, including CK+, Oulu (all 6 subdatasets), MMI, ISAFE, ISED, etc., are all incorporated. After comparing the new method with several popular or state-of-the-art facial expression recognition methods, its superiority in person-dependent facial expression recognition can be proposed from a statistical point of view.
Wei Huang 0013, Peng Zhang 0005, Yufei Zha, Yuming Fang 0001, Yanning Zhang 0001
IEEE Trans. Multim.4
2021 Unsupervised Cross-Modal Distillation for Thermal Infrared Tracking
abstract
The target representation learned by convolutional neural networks plays an important role in Thermal Infrared (TIR) tracking. Currently, most of the top-performing TIR trackers are still employing representations learned by the model trained on the RGB data. However, this representation does not take into account the information in the TIR modality itself, limiting the performance of TIR tracking.
Jingxian Sun 0003, Lichao Zhang 0001, Yufei Zha, Abel Gonzalez-Garcia, Peng Zhang 0005, Wei Huang 0013, Yanning Zhang 0001
ACM Multimedia3
2021 Multiple object tracking based on multi-task learning with strip attention
abstract
Abstract Multiple object tracking (MOT) framework based on bifurcate strategy was usually challenged by data association of different model path, which work for object localisation and appearance embedding independently. By incorporating the re‐identification (re‐ID) as appearance embedding model, more recent studies on task combination of a single network have made a great progress in tracking performance. Unfortunately, the contributive improvement from re‐ID model is hard to balance the accuracy and efficiency for the whole framework. For more effective enhancement of the overall tracking performance, a real‐time detection needs to be taken into consideration with other auxiliary means for MOT modelling. Therefore, in this study, a one‐shot multiple object tracking is proposed based on multi‐task learning to obtain satisfactory performance in both speed and robustness. With updated re‐training strategy for the backbone model of detection, a D2LA network is proposed to achieve more characteristic fine‐grained feature extraction in branching task of pedestrian recognition. Additionally, a strip attention module is also introduced to further strengthen the feature discriminative capability of the tracking framework in occlusion. Experiments on the 2DMOT15, MOT16, MOT17, and MOT20 benchmark data sets have shown a superior performance in comparison to other state‐of‐the‐art tracking approaches.
Yaoye Song, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Tao You, Yanning Zhang 0001
IET Image Process.4
2021 Learning spatial-channel regularization jointly with correlation filter for visual tracking
Yufei Zha, Zhuling Qiu, Jingxian Sun 0003, Peng Zhang 0005, Wei Huang 0013
Neurocomputing1
2021 Full-scaled deep metric learning for pedestrian re-identification
Wei Huang 0013, Mingyuan Luo, Peng Zhang 0005, Yufei Zha
Multim. Tools Appl.4
2021 A novel multi-loss-based deep adversarial network for handling challenging cases in semi-supervised image semantic segmentation
Wei Huang 0013, Zhanfei Shao, Mingyuan Luo, Peng Zhang 0005, Yufei Zha
Pattern Recognit. Lett.5
2021 SiamDA: Dual attention Siamese network for real-time visual tracking
Lei Pu, Xinxi Feng, Wangsheng Yu, Yufei Zha
Signal Process. Image Commun.5
2021 Multiple Instance Models Regression for Robust Visual Tracking
abstract
In comparison to single-model based trackers, the model-ensembled tracking strategy has shown a substantial adaptivity in handling various tracking challenges. As the performance of the tracker has been improved by combining different model outputs linearly, the insufficient consideration of each ensemble member’s contribution still limits the tracking performance to be further enhanced. As the performance of the tracker has been improved by combining different model outputs linearly, the insufficient consideration of each ensemble member’s contribution still limits the tracking performance to be further enhanced. In this paper, a tracking strategy based on multiple instance models regression (MIMRT) is proposed with a unified ensembling scheme. By formulating the tracking initialization with an instance model, the encoding process for an object’s specific detail is performed corresponding to the samples in each frame. The advantage of this operation is to guarantee the model frame-wise discrimination of short-term training, as well as to evaluate the reliability of each instance model by utilizing the long-lifetime samples obtained throughout the whole tracking procedure. To finalize the proposed tracking, all the independent instance models attached to the learned regression coefficients are ensembled with respect to the long-lifetime samples. This also effectively bridges the instance model as a latent variable to investigate a semantic association between the tracking model and the overall samples. A comprehensive experiment has shown that the proposed tracker is able to achieve superior performances compared to the state-of-art tracking approaches on both short-term datasets (e.g., OTB2013, OTB100, VOT2016, UAV123) and long-term dataset (UAV20L).
Yufei Zha, Yuanqiang Zhang, Tao Ku, Hanqiao Huang, Wei Huang 0013, Peng Zhang 0005
IEEE Trans. Circuits Syst. Video Technol.1
2020 MHASiam: Mixed High-Order Attention Siamese Network for Real-Time Visual Tracking
Lei Pu, Xinxi Feng, Wangsheng Yu, Yufei Zha, Zhiqiang Jiao
PRCV (2)5
2020 Robust Visual Tracking based on Adversarial Unlabeled Instance Generation with Label Smoothing Loss Regularization
Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Garth Douglas Cooper, Yanning Zhang 0001
Pattern Recognit.4
2020 Ensemble Tracking Based on Diverse Collaborative Framework With Multi-Cue Dynamic Fusion
abstract
Tracking with deep neural networks has been verified to arrive at a new level accuracy in many challenging scenarios, but the tracking robustness has been still challenged by model singularity and self-learning loop mechanism. As a promising solution for the limitations, to ensemble diverse tracking strategies into a highly-interactive framework has shown a potential effectiveness in recent studies. In this work, a collaborative tracking framework is proposed by exploiting both discriminative correlation filters and deep classifiers into an ensembling framework. With a multi-cue dynamic fusion scheme performed on all the ensembled members’ outputs, a robust long-term tracking can be achieved by calculating the optimal robustness scores based on a dynamic weighted sum of multi-cue metrics. Meanwhile, the obtained reliable and diverse training samples are also utilized to adaptively update the tracker in each branch with heuristic frequency, which is able to alleviate the training samples’ contamination and model corruption. Experiments on the OTB-2015, Temple color 128, UAV123, VOT2016, and VOT2018 benchmark datasets have shown superior performance in comparison to other state-of-the-art tracking approaches.
Peng Zhang 0005, Tao Zhuo, Wei Huang 0013, Yufei Zha, Yanning Zhang 0001
IEEE Trans. Multim.5
2020 Deep Position-Sensitive Tracking
abstract
Classification-based tracking strategies often face more challenges from intra-class discrimination than from inter-class separability. Even for deep convolutional neural networks that have been widely proven to be effective in various vision tasks, their intra-class discriminative capability is still limited by the weakness of softmax loss, especially for targets not seen in the training dataset. By taking intrinsic attributes of training samples into account, in this paper, we propose a position-sensitive loss coupled with softmax loss to achieve intra-class compactness and inter-class explicitness. Particularly, two additive margins are introduced to encode the position attribute for decision boundary maximization, which is also utilized with the proposed loss to supervise the fine-tuned features on the pre-trained model. With the nearest neighbor ranking measurement in the feature embedding domain, the whole scheme is able to reach an optimized balance between the feature-level inter-class semantic separability and instance-level intra-class relative distance ranking. We evaluate the proposed work on different popular benchmarks, and experimental results demonstrate that our tracking strategy performs favorably against most of the state-of-the-art trackers in the comparison of accuracy and robustness.
Yufei Zha, Tao Ku, Yunqiang Li, Peng Zhang 0005
IEEE Trans. Multim.1
2019 Push for Quantization: Deep Fisher Hashing
Yunqiang Li, Wenjie Pei, Yufei Zha, Jan C. van Gemert
BMVC3
2019 Deep Correlation Filter based Real-Time Tracker
Lei Pu, Xinxi Feng, Wangsheng Yu, Yufei Zha, Sugang Ma
FUSION5
2019 Learning Attention Regularization Correlation Filter for Visual Tracking
Zhuling Qiu, Yufei Zha
PRCV (1)2
2019 Visual tracking based on semantic and similarity learning
abstract
We present a method by combining the similarity and semantic features of a target to improve tracking performance in video sequences. Trackers based on Siamese networks have achieved success in recent competitions and databases through learning similarity according to binary labels. Unfortunately, such weak labels result in limiting the discriminative ability of the learned feature, thus it is difficult to identify the target itself from the distractors that have the same class. The authors observe that the inter‐class semantic features benefit to increase the separation between the target and the background, even distractors. Therefore, they proposed a network architecture which uses both similarity and semantic branches to obtain more discriminative features for locating the target accuracy in new frames. The large‐scale ImageNet VID dataset is employed to train the network. Even in the presence of background clutter, visual distortion, and distractors, the proposed method still maintains following the target. They test their method with the open benchmarks OTB and UAV123. The results show that their combined approach significantly improves the tracking ability relative to trackers using similarity or semantic features alone.
Yufei Zha, Zhuling Qiu, Wangsheng Yu
IET Comput. Vis.1
2019 ACFT: adversarial correlation filter for robust tracking
abstract
Tracking based on correlation filters has demonstrated outstanding performance in recent visual object tracking studies and competitions. However, the performance is limited since the boundary effects are introduced by the intrinsic circular structure. In this study, a tracker, called adversarial correlation filter tracker (ACFT), is proposed to solve the above problem through Generative Adversarial Networks (GANs) that is specifically strong at producing realistic‐looking data from noise circumstances. Especially, a mask is generated by the GANs to assist the conventional correlation filter for the spatial regularisation. By overcoming the feature independence of current regularisation in another tracker, the GANs’ mask can be effectively used to identify the robust features for the target variations representation in the temporal domain. Also in the spatial domain, the background features can be substantially suppressed to obtain the optimisation filter for more reliable matching and updating. In verification, the authors evaluate the proposed tracker on the standard tracking benchmarks, and the experimental results show that their tracker outperforms favourably against other state‐of‐the‐art trackers in the measurements of accuracy and robustness.
Hanqiao Huang, Yufei Zha, Meiyun Zheng, Peng Zhang 0005
IET Image Process.2
2019 Online Scale Adaptive Visual Tracking Based on Multilayer Convolutional Features
abstract
Convolutional neural networks can efficiently exploit sophisticated hierarchical features which have different properties for visual tracking problem. In this paper, by using multilayer convolutional features jointly and constructing a scale pyramid, we propose an online scale adaptive tracking method. We construct two separate correlation filters for translation and scale estimations. The translation filters improve the accuracy of target localization by a weighted fusion of multiple convolutional layers. Meanwhile, the separate scale filters achieve the optimal and fast scale estimation by a scale pyramid. This design decreases the mutual errors of translation and scale estimations, and reduces computational complexity efficiently. Moreover, in order to solve the problem of tracking drifts due to the severe occlusion or serious appearance changes of the target, we present a new adaptive and selective update mechanism to update the translation filters effectively. Extensive experimental results show that our proposed method achieves the excellent overall performance compared with the state-of-the-art methods.
Xin Wang 0026, Wangsheng Yu, Zefenfen Jin, Yufei Zha, Xianxiang Qin
IEEE Trans. Cybern.5
2018 Joint Identification-Verification Model for Visual Tracking
abstract
Similarity algorithms determine the location of the target by the similarity between the template and the candidate, the most similar candidate to the template is considered as the target in visual tracking. Similarity algorithms search the most similar candidate to the template as the current estimation for visual object. In practice, most trackers only take usage of the intra-class similarity, yet the inter-class semantic separability is ignored. In this paper, a joint identification-verification model is proposed to learn the similarity with the category attribute for visual tracking. This approach constructs the cost function both on the inter-class semantic separability and intra-class similarity, firstly. Then, the training dataset is fed into the network. To the end, the discriminative features are learned in the embedding space. During tracking phase, the template and candidates are fed into the network simultaneously. Thereforce, the target will be located correctly by the similarity metric between the template and candidates in the learned embedding space. We evaluate the proposed approach on the open benchmark: OTB50 and UAV123 dataset. A large number of experimental results show that the inter-class semantic separability can increase the discrimination for the similar distractors effectively, and bootstrap the tracking performances of the trackers based on the similarity learning.
Yufei Zha, Yuanqiang Zhang, Tao Ku, Lichao Zhang 0001
ICPR2
2016 Kernelised supervised context hashing
abstract
Most existing supervised hashing methods learn the affinity‐preserving binary codes to represent the high‐dimensional data. However, each hashing code is assumed as independent and irrelevant with other codes. In practice, the authors find that there exists context association among hashing bits. This study proposes a novel hashing method dubbed kernelised supervised context hashing, which considers the hashing codes interrelation to reduce the quantisation. In this work, the kernel formulation is employed to tackle the high‐dimensional data which is mostly linear inseparable first; and then different distributions are utilised to describe the binary codes context; finally, the hashing codes can be approximated by gradient descent method iteratively. Therefore, the correlation between the hash codes is integrated to redefine the metric measurement (i.e. Hamming affinity) to preserve the data similarity in the raw space. The authors evaluate the proposed method on three image benchmarks CIFAR‐10, MNIST and NUS‐WIDE for image retrieval, and experimental results show that it achieves better performance than several other state‐of‐the‐art methods.
Yun-qiang Li, Yufei Zha, Bing Qin 0002
IET Image Process.2
2016 Robust and fast visual tracking via spatial kernel phase correlation filter
Lichao Zhang 0001, Duyan Bi, Yufei Zha, Hongxun Wang, Tao Ku
Neurocomputing3
2015 Multi-scale mean shift tracking
abstract
In this study, a three‐dimensional mean shift tracking algorithm, which combines the multi‐scale model and background weighted spatial histogram, is proposed to address the problem of scale estimation under the framework of mean shift tracking. The target template is modelled with multi‐scale model and described with three‐dimensional spatial histogram. The tracking algorithm is implemented by three‐dimensional mean shift iteration, which translates the problem of scale estimation in two‐dimensional image plane into the localisation in three‐dimensional image space. To enhance the robustness, the background weighted histogram is employed to suppress the background information in the target candidate model. Firstly, the multi‐scale model and three‐dimensional spatial histogram are introduced to represent the target template. Then, the three‐dimensional mean shift iteration formulation is derived based on the similarity measure between the target model and the target candidate model. Finally, a multi‐scale mean shift tracking algorithm combining multi‐scale model and background weighted spatial histogram is proposed. The proposed algorithm is evaluated on some challenging sequences which contain scale changed targets and other complex appearance variations in comparison with three representative mean shift based tracking algorithms. Both the qualitative results and quantitative analysis indicate that the proposed algorithm outperforms the referenced algorithms in both tracking precision and scale estimation.
Wangsheng Yu, Xiaohua Tian, Yufei Zha
IET Comput. Vis.4
2015 Robust visual tracking based on product sparse coding
Hong-tu Huang, Duyan Bi, Yufei Zha, Shiping Ma
Pattern Recognit. Lett.3
2014 Robust visual tracking based on watershed regions
abstract
Robust visual tracking is a very challenging problem especially when the target undergoes large appearance variation. In this study, the authors propose an efficient and effective tracker based on watershed regions. As middle‐level visual cues, watershed regions contain more semantics information than low‐level features, and reflect more structure information than high‐level model. First, the authors manually select the target template in initial frame, and predict the target candidate in the next frame using motion prediction. Then, the authors utilise marker‐based watershed algorithm to obtain the watershed regions of target template and candidate template, and describe each region with multiple features. Next, the authors calculate the nearest neighbour in feature space to match the watershed regions and construct an affine relation from target template to candidate template. Finally, the authors resolve the affine relation to calculate the final tracking result, and update the template for the following tracking. The authors test their tracker on some challenging sequences with appearance variation range from illumination change, partial occlusion, pose change to background clutters and compare it with some state‐of‐the‐art works. Experiment results indicate that the proposed tracker is robust to the large appearance variation and exceeds the state‐of‐the‐art trackers in most situations.
Wangsheng Yu, Xiaohua Tian, Yufei Zha
IET Comput. Vis.4
2010 Graph-based transductive learning for robust visual tracking
Yufei Zha, Duyan Bi
Pattern Recognit.1
2009 Learning complex background by multi-scale discriminative model
Yufei Zha, Duyan Bi
Pattern Recognit. Lett.1