Di Yuan 0002

dblp:09/5856-2 · DBLP profile ↗
← Back
44ranked-venue papers
16as first author
34since 2021 · last 2027
0000-0001-9403-1112ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 8 first-author · 19 since 2021Artificial intelligence and machine learning · 18 · 8 first-author · 14 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2027 Physics prior adapter tuning for thermal infrared tracking
Jingyuan Guo, Qiao Liu 0001, Kanlun Tan, Di Yuan 0002, Yunpeng Liu 0001
Expert Syst. Appl.4
2026 DINOTrack: Leverage Differential Attention for Noise-Aware Visual Tracking with DINOv3
abstract
Despite significant progress in Transformer-based visual tracking, standard self-attention mechanisms inherently lack the discriminative power required for accurate template search and matching. Due to the global normalization of softmax, the model inevitably assigns probabilistic quality to background noise terms that are semantically similar to the target. This "attention noise" not only reduces tracking accuracy but also forces the network to require substantial training data and computational costs to learn robust feature suppression. To address this, we propose DINOTrack, a noise-aware tracking framework designed to systematically eliminate background clutter. First, we construct a discriminative representation foundation using a frozen DINOv3 backbone with hierarchical multi-level feature fusion, ensuring the model captures robust semantic features essential for distinguishing the target from distractors. Second, to mitigate noise before feature interaction, we design a Class-Guided Heatmap Modulation (CGHM) module. By utilizing the template’s class label as a semantic prior, this module explicitly suppresses background interference and enhances potential target responses. Crucially, we introduce a Differential Denoising Interaction (DDI) mechanism for active noise cancellation. By performing a difference operation on dual query-key pairs, the DDI module effectively eliminates shared attention noise caused by semantic aliasing, enabling precise target matching. Extensive experiments demonstrate that DINOTrack achieves significant performance improvements on datasets such as LaSOT, GOT-10k, TNL2K, LaSOText, and TrackingNet.
Gu Geng, Di Yuan 0002, Rui Chen 0001, Qiao Liu 0001
ICMR3
2026 Multi-semantic depth collaborative multi-modality image fusion network
Huayi Zhu, Rui Chen 0001, Qiao Liu 0001, Xiaojun Chang, Di Yuan 0002
Knowl. Based Syst.6
2026 CMMDL: Cross-modal multi-domain learning method for image fusion
Di Yuan 0002, Huayi Zhu, Rui Chen 0001, Sida Zhou, Xiu Shu, Qiao Liu 0001
Neural Networks1
2026 EVSSD: Efficient visual state space decoding for 2D medical image segmentation
Di Yuan 0002, Youqiang Xiong, Xiu Shu, Xiaojun Chang
Signal Process.1
2026 Adaptive Mamba Network Guided by Mixture of Experts for Infrared and Visible Image Fusion
abstract
Current mainstream methods for multimodal image fusion rely on CNN and Transformers, which often suffer from severe information loss or high computational complexity. However, state-space models offer a promising alternative with their linear complexity and reduced historical information loss. This paper proposes a novel adaptive Mamba network, named MGMFuse, for infrared and visible image fusion, aiming to seamlessly integrate thermal targets with high-resolution texture details. Specifically, feature extraction is performed using a dual-branch architecture. We design the MoEGMamba module that uses the Mixture of Experts'prompt generation mechanism to dynamically guide Mamba for adaptive state space modeling. Furthermore, a specialized mechanism is proposed to execute modality-specific channel-wise rectification and cross-modal interaction to align deep semantic features. Experimental results on multiple datasets demonstrate that the proposed algorithm outperforms existing methods. Furthermore, it significantly enhances the performance of downstream visual tasks.
Quanrui Wen, Huayi Zhu, Rui Chen 0001, Qiao Liu 0001, Di Yuan 0002
IEEE Signal Process. Lett.6
2026 Reliable-Teacher: Uncertainty-Guided Collaborative Learning for Nighttime Object Detection
abstract
Nighttime object detection presents significant challenges due to the scarcity of large-scale, high-quality annotations across diverse nighttime scenarios. To circumvent the need for manual nighttime image annotation, researchers have explored Unsupervised Domain Adaptive Object Detection (UDA-OD), which transfers knowledge from labeled daytime datasets to unlabeled nighttime data through pseudo-labeling. While existing approaches have shown promising results, their effectiveness remains limited by the low quality of pseudo labels, restricting model adaptation to nighttime conditions. To address these limitations, we propose Reliable-Teacher, a novel mutual-learning framework that comprehensively leverages target domain knowledge through Uncertainty-Guided Collaborative Learning. Specifically, our approach consists of three key components: 1) A Collaborative Pseudo-Label Construction module that intelligently integrates reliable Teacher-generated pseudo-labels into Student proposals, significantly enhancing pseudo-label quality; 2) An Uncertainty-Guided Consistency Reasoning module that enforces inter-category consistency between Teacher and Student predictions at both anchor and bounding box levels; 3) A Reliability-Weighted Classification Loss that minimizes the influence of unreliable predictions to further enhance uncertainty-guided learning. Extensive experiments demonstrate that Reliable-Teacher significantly outperforms state-of-the-art methods, achieving performance gain of up to 3.1%, 2.2% and 1.7% mAP on BDD100K [1], SHIFT [2], and VisDrone [3] benchmarks, respectively. Upon acceptance, our code will be released to facilitate further research in this domain.
Wenjing Jia, Jiaqi Xiao, Jinchang Ren, Di Yuan 0002, Qiguang Miao, Xiangjian He
IEEE Trans. Circuits Syst. Video Technol.6
2026 PPIFuse: Physical Priors Injected Infrared and Visible Image Fusion
abstract
Existing infrared and visible image fusion methods commonly use two structurally identical networks to extract deep features from source images, followed by a handcrafted or learnable feature fusion strategy. These methods overlook the modality-specific characteristics of the two image types, impairing the model’s ability to fully exploit their complementary information. Additionally, their fusion results often exhibit issues such as texture detail loss or unclear thermal targets. This is because the fusion rules they used are either too simple or too redundant. To address these challenges, we start from the infrared physics priors that are naturally complementary to visible images and incorporate the thermal diffusion equation and Stefan-Boltzmann Law into the image fusion architecture. Based on these two physical priors, we design a Thermal Diffusion Convolution (TDC) and a Stefan Thermal Attention (STA) to better extract infrared-specific features. Specifically, the TDC module leverages the anisotropic and isotropic characteristics of thermal diffusion adaptively to sharpen the edges of thermal targets and remove infrared noise, minimizing artifacts in the fused results. By decomposing the Stefan-Boltzmann Law, STA pays more attention on thermal features while suppressing redundant information, enabling more effective aggregation of complementary modality-specific details. To make full use of layer-wise complementary features, we propose an Interactive Injection Fusion framework(IIF) that hierarchically integrates these features, enhancing the richness of fused image content. Furthermore, an energy conservation constraint is designed to ensure the fused images adhere to physical principles. Extensive experimental results on five datasets demonstrate that our method sets a new state-of-the-art. Code is available at https://github.com/QiaoLiuHit/PPIFuse.
Qianhong Zhang, Qiao Liu 0001, Di Yuan 0002, Xin Li 0034, Yunpeng Liu 0001
IEEE Trans. Circuits Syst. Video Technol.3
2026 Unsupervised Domain Adaptive Thermal Infrared Tracking
abstract
Existing deep Thermal InfraRed (TIR) trackers often use RGB datasets for training due to the lack of large-scale labeled TIR datasets. However, the performance of these methods on TIR image sequences is significantly degraded, because of the domain shift problem between the RGB and TIR datasets. To solve this problem, in this paper, we propose an unsupervised Dual-level Domain Adaptation TIR Tracking framework (DDAT), which can benefit from training on large-scale labeled RGB datasets and unlabeled TIR datasets. Specifically, to transfer the useful knowledge learned from RGB dataset to TIR tracking, we first propose an adversarial-based adaptation module on both the semantic-level and the feature-level. While the semantic-level adaptation can reduce the semantic gap between the TIR and RGB tracking tasks, the feature-level adaptation can learn domain-invariant features for more robust tracking. Second, we propose a partial domain adaptation module to alleviate the negative transfer problem because the RGB and TIR tracking domains have a non-identical class and feature spaces. Instead of aligning the entire feature space, this module adaptively selects partial similarity samples and features for alignment, thus getting more fine-grained aligned results. Third, we collect a currently largest-scale unlabeled TIR dataset to train the proposed framework. Extensive experiments on five TIR tracking benchmarks demonstrate the proposed method is effective and sets a new state-of-the-art.
Qiao Liu 0001, Xin Li 0034, Jiatian Pi, Di Yuan 0002, Yunpeng Liu 0001
IEEE Trans. Multim.5
2026 Interactive Classification and Regression for Visual Tracking with Dual Update Strategy
abstract
The current two-stage tracking method locates the target using the position with the highest confidence score, and updates the template using a carefully designed template update strategy. However, we identify two key issues with these trackers: (1) the update strategy lacks continuous, cost-free template adaptation, leading to suboptimal tracking under appearance changes and (2) the location with the highest confidence score does not always yield accurate bounding boxes, potentially resulting in incomplete target coverage. In this article, we propose a novel tracker that incorporates two key innovations. First, the tracker employs a dual update strategy that performs online template updates at both the image and feature levels. This strategy enables continuous adaptation to target appearance changes without introducing additional computational overhead. Second, we enhance the existing loss function by introducing a Classification–Regression Interaction (CRI) loss, which guides the training process to produce confidence scores that more accurately reflect the quality of the predicted bounding boxes. Extensive experiments are conducted to evaluate the performance of our tracker and the effectiveness of the proposed methods. The experimental results show that our method has achieved a comprehensive improvement over the baseline on five datasets, and achieves competitive performance compared to state-of-the-art trackers.
Di Yuan 0002, Gu Geng, Qiao Liu 0001, Xiaojun Chang, Zhenyu He 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2025 Efficient Hierarchical Domain Adaptive Thermal Infrared Tracking
abstract
Constrained by the scarcity of labeled Thermal InfraRed (TIR) training data, current TIR trackers commonly rely on pre-trained RGB trackers. However, the domain discrepancy between TIR and RGB images limits effective utilization of RGB features, significantly degrades TIR tracking performance. To solve this challenge, we propose a hierarchical domain adaptation model to transfer useful pre-trained RGB features into TIR tracking more effective and efficient. Specifically, we first design a reflectance consistency network to learn style-invariant representations. Second, we present a target-aware adversarial network to align the target semantic features of the two domains. These two modules respectively narrow the distribution gap at the stylistic and semantic levels in a hierarchical manner. Third, to solve the inefficiency problem of domain adaptive training, we also propose a Bi-rank adapter side network to accelerate this process. While significantly reducing training time by 90%, our method achieves a new state-of-the-art on four TIR tracking benchmarks.
Kanlun Tan, Qiao Liu 0001, Di Yuan 0002, Xin Li 0034, Yunpeng Liu 0001
ICASSP4
2025 A Multi-Stream Visual-Spectral-Spatial Adaptive Hyperspectral Object Tracking
abstract
Hyperspectral videos contain rich spectral and spatial information, which has enormous potential for object tracking compared to traditional RGB videos. However, the limited hyperspectral data, spectral differences across different bands, and high computational costs result in existing trackers being unable to effectively build connections between spectral and spatial information, leading to suboptimal tracking performance. To address these issues, this paper proposes a multi-stream Visual-Spectral-Spatial adaptive hyperspectral object tracking model(VSS). First, we design a multi-stream hyperspectral Transformer module to receive spectral data from different bands and extract Visual-Spectral-Spatial features across those bands. Then, to address the band differences between different modalities, we introduce the Bidirectional Visual-Spectral-Spatial Adapter module, which adaptively fuses visual, spectral, and spatial information from different bands. Finally, experiments on the HOTC dataset demonstrate the excellent performance of the VSS model.
Qiao Liu 0001, Zhenyu He 0001, Di Yuan 0002
ICMR4
2025 Overexposed infrared and visible image fusion benchmark and baseline
Renping Xie, Ming Tao 0001, Hengye Xu, Di Yuan 0002, Qiao Liu 0001
Expert Syst. Appl.5
2025 Uncertainty and diversity-based active learning for UAV tracking
Yingqin Liang, Feng Huang 0007, Zhaobing Qiu, Xiu Shu, Qiao Liu 0001, Di Yuan 0002
Neurocomputing6
2025 Variational methods with application to medical image segmentation: A survey
Xiu Shu, Zhihui Li 0001, Xiaojun Chang, Di Yuan 0002
Neurocomputing4
2025 An active learning model based on image similarity for skin lesion segmentation
Xiu Shu, Zhihui Li 0001, Chunwei Tian, Xiaojun Chang, Di Yuan 0002
Neurocomputing5
2025 D2Fusion: Dual-domain feature decoupling for infrared and visible image fusion
Yan Fan 0006, Wei Ran, Kanlun Tan, Qiao Liu 0001, Di Yuan 0002, Xin Li 0034, Yunpeng Liu 0001
Knowl. Based Syst.5
2025 Multi-Scale Adaptive Cascaded Tracking for Vision-Language Integration
abstract
Unlike traditional single-modal visual tracking, vision-language tracking leverages textual descriptions to provide semantic understanding, aiding in handling challenges such as appearance variations, occlusions, and motion blur. However, the semantic gap between textual and visual information complicates cross-modal alignment and fusion. Moreover, fixed textual descriptions may fail to reflect real-time variations in dynamic video frames, thereby affecting tracking accuracy. To address these challenges, we propose MACT, the Multi-scale Adaptive Cascaded Tracker, which enhances vision-language integration through a series of key mechanisms. First, Multi-Scale Feature Alignment (MSFA) encodes image and text features at multiple scales and dynamically aligns them using cross-modal attention, strengthening information flow and interaction between modalities. Second, Adaptive Cascaded Fusion (ACF) adaptively learns the correspondence between text descriptions and video frames, extracting appropriate hierarchical interaction features to ensure that textual information effectively aids visual tracking. Finally, Scale Recovery(SR) integrates multi-scale deep feature representations back to the original scale before passing them to the tracking head for final predictions. Extensive experiments on TNL2K, LaSOT, and OTB99 demonstrate that our proposed MACT achieves state-of-the-art tracking performance.
Guomao Guo, Gu Geng, Qiao Liu 0001, Di Yuan 0002
IEEE Signal Process. Lett.5
2025 Fine-Grained Feature and Template Reconstruction for TIR Object Tracking
abstract
Thermal infrared (TIR) object tracking is a significant subject within the field of computer vision. Currently, TIR object tracking faces challenges such as insufficient representation of object texture information and underutilization of temporal information, which severely affects the tracking accuracy of TIR tracking methods. To address these issues, we propose a TIR object tracking method (called: FFTR) based on fine-grained feature and template reconstruction. Specifically, aiming at the fine-grained information of the TIR object, we employ a frequency channel attention mechanism that transforms TIR images into the frequency domain using discrete cosine transform features. By capturing the fine-grained feature of TIR images from the frequency domain, we enhance the model’s ability to comprehend these images. To better leverage temporal information, we utilize a template region reconstruction method. This method reconstructs the template from the previous frame based on the search area of the current frame, which is then incorporated into the attention computation for the subsequent frame, thereby improving the tracking capability of TIR objects. Extensive quantitative and qualitative experiments show that our method achieves competitive tracking performance on the TIR benchmarks.
Donghai Liao, Xiu Shu, Zhihui Li 0001, Qiao Liu 0001, Di Yuan 0002, Xiaojun Chang, Zhenyu He 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 Adaptive Trajectory Correction for Underwater Object Tracking
abstract
Most extant underwater object tracking (UOT) utilize generic tracking algorithms, which lack applicability to underwater tracking scenarios. Moreover, these algorithms primarily emphasize minimizing interference from various challenging tasks to prevent target drift, but pay less attention to the strategies for mitigating target drift once it occurs. To alleviate the above problems, we propose a simple, effective, and UOT-focused adaptive trajectory correction framework, named ATCTrack. From the perspective of tracking failure, this methodology aims to promptly identify and rectify unreasonable target drift through accurate trajectory coordinate correction and trajectory template updates. Additionally, to mitigate the adverse effects of potential erroneous corrections, we implement an adaptive strategy that corrects only significant target drift, allowing for self-correction within a certain margin. Finally, we introduce an adaptive underwater image enhancement technique to improve the underwater image quality and maintain the trajectory’s stability and clarity. Our tracker achieves state-of-the-art performance on the currently prevalent UOT tracking benchmarks compared to other trackers.
Di Yuan 0002, Xiu Shu, Qiao Liu 0001, Xiaojun Chang, Zhenyu He 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 MFPNet: Mixed Feature Perception Network for Automated Skin Lesion Segmentation
Youqiang Xiong, Di Yuan 0002, Lu Li 0005, Xiu Shu
PRCV (14)2
2024 CFRNet: Road Extraction in Remote Sensing Images Based on Cascade Fusion Network
abstract
Road extraction from remote sensing images has attracted widespread attention of researchers due to its crucial role in the fields of autopilot, urban planning, navigation, and other fields. However, the task becomes challenging as the roads in remote sensing images are easily occluded by obstacles such as shadows, buildings and trees. In this letter, a cascade fusion network for road extraction (CFRNet) in remote sensing images is proposed. Considering the lightweight characteristics of MobileNet block (MbBlock), it is used as the feature extraction module of the backbone network. To enable CFRNet to generate and fuse more features at multiscale, we design several cascade stages. Each stage includes a sub-backbone for feature extraction and a triple-level adaptive feature fusion (TAFF) module for feature fusion. This structure can more deeply and effectively fuse multiscale features with most of the parameters in the entire backbone. The experimental results demonstrate that the proposed CFRNet significantly outperforms other state-of-the-art methods on the publicly available Istanbul City Road dataset and DeepGlobe Road dataset. Specifically, it achieves an intersection over union (IoU) of 89.76%, reflecting a 5.3% improvement on the Istanbul dataset, and 67.22% with a 0.98% enhancement on the DeepGlobe Road dataset. Our code is available athttps://github.com/XYQ1517/CFRNet.
Youqiang Xiong, Lu Li 0005, Di Yuan 0002, Tianliang Ma, Yuping Yang
IEEE Geosci. Remote. Sens. Lett.3
2024 Self-supervised discriminative model prediction for visual tracking
Di Yuan 0002, Gu Geng, Xiu Shu, Qiao Liu 0001, Xiaojun Chang, Zhenyu He 0001, Guangming Shi
Neural Comput. Appl.1
2024 LSOTB-TIR: A Large-Scale High-Diversity Thermal Infrared Single Object Tracking Benchmark
abstract
Unlike visual object tracking, thermal infrared (TIR) object tracking methods can track the target of interest in poor visibility such as rain, snow, and fog, or even in total darkness. This feature brings a wide range of application prospects for TIR object-tracking methods. However, this field lacks a unified and large-scale training and evaluation benchmark, which has severely hindered its development. To this end, we present a large-scale and high-diversity unified TIR single object tracking benchmark, called LSOTB-TIR, which consists of a tracking evaluation dataset and a general training dataset with a total of 1416 TIR sequences and more than 643 K frames. We annotate the bounding box of objects in every frame of all sequences and generate over 770 K bounding boxes in total. To the best of our knowledge, LSOTB-TIR is the largest and most diverse TIR object tracking benchmark to date. We spilt the evaluation dataset into a short-term tracking subset and a long-term tracking subset to evaluate trackers using different paradigms. What's more, to evaluate a tracker on different attributes, we also define four scenario attributes and 12 challenge attributes in the short-term tracking evaluation subset. By releasing LSOTB-TIR, we encourage the community to develop deep learning-based TIR trackers and evaluate them fairly and comprehensively. We evaluate and analyze 40 trackers on LSOTB-TIR to provide a series of baselines and give some insights and future research directions in TIR object tracking. Furthermore, we retrain several representative deep trackers on LSOTB-TIR, and their results demonstrate that the proposed training dataset significantly improves the performance of deep TIR trackers. Codes and dataset are available at https://github.com/QiaoLiuHit/LSOTB-TIR.
Qiao Liu 0001, Xin Li 0034, Di Yuan 0002, Xiaojun Chang, Zhenyu He 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 Active Learning for Deep Visual Tracking
abstract
Convolutional neural networks (CNNs) have been successfully applied to the single target tracking task in recent years. Generally, training a deep CNN model requires numerous labeled training samples, and the number and quality of these samples directly affect the representational capability of the trained model. However, this approach is restrictive in practice, because manually labeling such a large number of training samples is time-consuming and prohibitively expensive. In this article, we propose an active learning method for deep visual tracking, which selects and annotates the unlabeled samples to train the deep CNN model. Under the guidance of active learning, the tracker based on the trained deep CNN model can achieve competitive tracking performance while reducing the labeling cost. More specifically, to ensure the diversity of selected samples, we propose an active learning method based on multiframe collaboration to select those training samples that should be and need to be annotated. Meanwhile, considering the representativeness of these selected samples, we adopt a nearest-neighbor discrimination method based on the average nearest-neighbor distance to screen isolated samples and low-quality samples. Therefore, the training samples' subset selected based on our method requires only a given budget to maintain the diversity and representativeness of the entire sample set. Furthermore, we adopt a Tversky loss to improve the bounding box estimation of our tracker, which can ensure that the tracker achieves more accurate target states. Extensive experimental results confirm that our active-learning-based tracker (ALT) achieves competitive tracking accuracy and speed compared with state-of-the-art trackers on the seven most challenging evaluation benchmarks. Project website: https://sites.google.com/view/altrack/.
Di Yuan 0002, Xiaojun Chang, Qiao Liu 0001, Yi Yang 0001, Minglei Shu, Zhenyu He 0001, Guangming Shi
IEEE Trans. Neural Networks Learn. Syst.1
2023 Robust thermal infrared tracking via an adaptively multi-feature fusion model
Di Yuan 0002, Xiu Shu, Qiao Liu 0001, Zhenyu He 0001
Neural Comput. Appl.1
2023 Learning Dual-Level Deep Representation for Thermal Infrared Tracking
abstract
The feature models used by existing Thermal InfraRed (TIR) tracking methods are usually learned from RGB images due to the lack of a large-scale TIR image training dataset. However, these feature models are less effective in representing TIR objects and they are difficult to effectively distinguish distractors because they do not contain fine-grained discriminative information. To this end, we propose a dual-level feature model containing the TIR-specific discriminative feature and fine-grained correlation feature for robust TIR object tracking. Specifically, to distinguish inter-class TIR objects, we first design an auxiliary multi-classification network to learn the TIR-specific discriminative feature. Then, to recognize intra-class TIR objects, we propose a fine-grained aware module to learn the fine-grained correlation feature. These two kinds of features complement each other and represent TIR objects in the levels of inter-class and intra-class respectively. These two feature models are constructed using a multi-task matching framework and are jointly optimized on the TIR object tracking task. In addition, we develop a large-scale TIR image dataset to train the network for learning TIR-specific feature patterns. To the best of our knowledge, this is the largest TIR tracking training dataset with the richest object class and scenario. To verify the effectiveness of the proposed dual-level feature model, we propose an offline TIR tracker (MMNet) and an online TIR tracker (ECO-MM) based on the feature model and evaluate them on three TIR tracking benchmarks. Extensive experimental results on these benchmarks demonstrate that the proposed algorithms perform favorably against the state-of-the-art methods.
Qiao Liu 0001, Di Yuan 0002, Nana Fan, Peng Gao 0005, Xin Li 0034, Zhenyu He 0001
IEEE Trans. Multim.2
2022 Structural target-aware model for thermal infrared tracking
Di Yuan 0002, Xiu Shu, Qiao Liu 0001, Zhenyu He 0001
Neurocomputing1
2022 Accurate bounding-box regression with distance-IoU loss for visual tracking
Di Yuan 0002, Xiu Shu, Nana Fan, Xiaojun Chang, Qiao Liu 0001, Zhenyu He 0001
J. Vis. Commun. Image Represent.1
2022 HCDC-SRCF tracker: Learning an adaptively multi-feature fuse tracker in spatial regularized correlation filters framework
Bing Liu 0019, Xiaojun Chang, Di Yuan 0002
Knowl. Based Syst.3
2022 SiamCorners: Siamese Corner Networks for Visual Tracking
abstract
The current Siamese network based on region proposal network (RPN) has attracted great attention in visual tracking due to its excellent accuracy and high efficiency. However, the design of the RPN involves the selection of the number, scale, and aspect ratios of anchor boxes, which will affect the applicability and convenience of the model. Furthermore, these anchor boxes require complicated calculations, such as calculating their intersection-over-union (IoU) with ground truth bounding boxes. Due to the problems related to anchor boxes, we propose a simple yet effective anchor-free tracker (named Siamese corner networks, SiamCorners), which is end-to-end trained offline on large-scale image pairs. Specifically, we introduce a modified corner pooling layer to convert the bounding box estimate of the target into a pair of corner predictions (the bottom-right and the top-left corners). By tracking a target as a pair of corners, we avoid the need to design the anchor boxes. This will make the entire tracking algorithm more flexible and simple than anchor-based trackers. In our network design, we further introduce a layer-wise feature aggregation strategy that enables the corner pooling module to predict multiple corners for a tracking target in deep networks. We then introduce a new penalty term that is used to select an optimal tracking box in these candidate corners. Finally, SiamCorners achieves experimental results that are comparable to the state-of-art tracker while maintaining a high running speed. In particular, SiamCorners achieves a 53.7% AUC on NFS30 and a 61.4% AUC on UAV123, while still running at 42 frames per second (FPS).
Kai Yang 0018, Zhenyu He 0001, Wenjie Pei, Zikun Zhou, Xin Li 0034, Di Yuan 0002, Haijun Zhang 0002
IEEE Trans. Multim.6
2022 Learning Adaptive Spatial-Temporal Context-Aware Correlation Filters for UAV Tracking
abstract
Tracking in the unmanned aerial vehicle (UAV) scenarios is one of the main components of target-tracking tasks. Different from the target-tracking task in the general scenarios, the target-tracking task in the UAV scenarios is very challenging because of factors such as small scale and aerial view. Although the discriminative correlation filter (DCF)-based tracker has achieved good results in tracking tasks in general scenarios, the boundary effect caused by the dense sampling method will reduce the tracking accuracy, especially in UAV-tracking scenarios. In this work, we propose learning an adaptive spatial-temporal context-aware (ASTCA) model in the DCF-based tracking framework to improve the tracking accuracy and reduce the influence of boundary effect, thereby enabling our tracker to more appropriately handle UAV-tracking tasks. Specifically, our ASTCA model can learn a spatial-temporal context weight, which can precisely distinguish the target and background in the UAV-tracking scenarios. Besides, considering the small target scale and the aerial view in UAV-tracking scenarios, our ASTCA model incorporates spatial context information within the DCF-based tracker, which could effectively alleviate background interference. Extensive experiments demonstrate that our ASTCA method performs favorably against state-of-the-art tracking methods on some standard UAV datasets.
Di Yuan 0002, Xiaojun Chang, Zhihui Li 0001, Zhenyu He 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2021 Self-Supervised Deep Correlation Tracking
abstract
The training of a feature extraction network typically requires abundant manually annotated training samples, making this a time-consuming and costly process. Accordingly, we propose an effective self-supervised learning-based tracker in a deep correlation framework (named: self-SDCT). Motivated by the forward-backward tracking consistency of a robust tracker, we propose a multi-cycle consistency loss as self-supervised information for learning feature extraction network from adjacent video frames. At the training stage, we generate pseudo-labels of consecutive video frames by forward-backward prediction under a Siamese correlation tracking framework and utilize the proposed multi-cycle consistency loss to learn a feature extraction network. Furthermore, we propose a similarity dropout strategy to enable some low-quality training sample pairs to be dropped and also adopt a cycle trajectory consistency loss in each sample pair to improve the training loss function. At the tracking stage, we employ the pre-trained feature extraction network to extract features and utilize a Siamese correlation tracking framework to locate the target using forward tracking alone. Extensive experimental results indicate that the proposed self-supervised deep correlation tracker (self-SDCT) achieves competitive tracking performance contrasted to state-of-the-art supervised and unsupervised tracking methods on standard evaluation benchmarks.
Di Yuan 0002, Xiaojun Chang, Po-Yao Huang 0001, Qiao Liu 0001, Zhenyu He 0001
IEEE Trans. Image Process.1
2021 Learning Deep Multi-Level Similarity for Thermal Infrared Object Tracking
abstract
Existing deep Thermal InfraRed (TIR) trackers only use semantic features to represent the TIR object, which lack the sufficient discriminative capacity for handling distractors. This becomes worse when the feature extraction network is only trained on RGB images. To address this issue, we propose a multi-level similarity model under a Siamese framework for robust TIR object tracking. Specifically, we compute different pattern similarities using the proposed multi-level similarity network. One of them focuses on the global semantic similarity and the other computes the local structural similarity of the TIR object. These two similarities complement each other and hence enhance the discriminative capacity of the network for handling distractors. In addition, we design a simple while effective relative entropy based ensemble subnetwork to integrate the semantic and structural similarities. This subnetwork can adaptive learn the weights of the semantic and structural similarities at the training stage. To further enhance the discriminative capacity of the tracker, we propose a large-scale TIR video sequence dataset for training the proposed model. To the best of our knowledge, this is the first and the largest TIR object tracking training dataset to date. The proposed TIR dataset not only benefits the training for TIR object tracking but also can be applied to numerous TIR visual tasks. Extensive experimental results on three benchmarks demonstrate that the proposed algorithm performs favorably against the state-of-the-art methods.
Qiao Liu 0001, Xin Li 0034, Zhenyu He 0001, Nana Fan, Di Yuan 0002, Hongpeng Wang 0002
IEEE Trans. Multim.5
2020 Multi-Task Driven Feature Models for Thermal Infrared Tracking
abstract
Existing deep Thermal InfraRed (TIR) trackers usually use the feature models of RGB trackers for representation. However, these feature models learned on RGB images are neither effective in representing TIR objects nor taking fine-grained TIR information into consideration. To this end, we develop a multi-task framework to learn the TIR-specific discriminative features and fine-grained correlation features for TIR tracking. Specifically, we first use an auxiliary classification network to guide the generation of TIR-specific discriminative features for distinguishing the TIR objects belonging to different classes. Second, we design a fine-grained aware module to capture more subtle information for distinguishing the TIR objects belonging to the same class. These two kinds of features complement each other and recognize TIR objects in the levels of inter-class and intra-class respectively. These two feature models are learned using a multi-task matching framework and are jointly optimized on the TIR tracking task. In addition, we develop a large-scale TIR training dataset to train the network for adapting the model to the TIR domain. Extensive experimental results on three benchmarks show that the proposed algorithm achieves a relative gain of 10% over the baseline and performs favorably against the state-of-the-art methods. Codes and the proposed TIR dataset are available at https://github.com/QiaoLiuHit/MMNet.
Qiao Liu 0001, Xin Li 0034, Zhenyu He 0001, Nana Fan, Di Yuan 0002, Wei Liu 0065, Yongsheng Liang 0001
AAAI5
2020 LSOTB-TIR: A Large-Scale High-Diversity Thermal Infrared Object Tracking Benchmark
abstract
In this paper, we present a Large-Scale and high-diversity general Thermal InfraRed (TIR) Object Tracking Benchmark, called LSOTB-TIR, which consists of an evaluation dataset and a training dataset with a total of 1,400 TIR sequences and more than 600K frames. We annotate the bounding box of objects in every frame of all sequences and generate over 730K bounding boxes in total. To the best of our knowledge, LSOTB-TIR is the largest and most diverse TIR object tracking benchmark to date. To evaluate a tracker on different attributes, we define 4 scenario attributes and 12 challenge attributes in the evaluation dataset. By releasing LSOTB-TIR, we encourage the community to develop deep learning based TIR trackers and evaluate them fairly and comprehensively. We evaluate and analyze more than 30 trackers on LSOTB-TIR to provide a series of baselines, and the results show that deep trackers achieve promising performance. Furthermore, we re-train several representative deep trackers on LSOTB-TIR, and their results demonstrate that the proposed training dataset significantly improves the performance of deep TIR trackers. Codes and dataset are available at https://github.com/QiaoLiuHit/LSOTB-TIR.
Qiao Liu 0001, Xin Li 0034, Zhenyu He 0001, Chenglong Li 0002, Zikun Zhou, Di Yuan 0002, Jing Li 0071, Kai Yang 0018, Nana Fan, Feng Zheng 0001
ACM Multimedia7
2020 TRBACF: Learning temporal regularized correlation filters for high performance online visual object tracking
Di Yuan 0002, Xiu Shu, Zhenyu He 0001
J. Vis. Commun. Image Represent.1
2020 Learning target-focusing convolutional regression model for visual object tracking
Di Yuan 0002, Nana Fan, Zhenyu He 0001
Knowl. Based Syst.1
2020 Robust visual tracking with correlation filters and metric learning
Di Yuan 0002, Zhenyu He 0001
Knowl. Based Syst.1
2020 Visual object tracking with adaptive structural convolutional network
Di Yuan 0002, Xin Li 0034, Zhenyu He 0001, Qiao Liu 0001, Shuwei Lu
Knowl. Based Syst.1
2020 Adaptive weight part-based convolutional network for person re-identification
Xiu Shu, Di Yuan 0002, Qiao Liu 0001
Multim. Tools Appl.2
2019 Particle filter re-detection for visual tracking via correlation filters
Di Yuan 0002, Xiaohuan Lu, Yingyi Liang
Multim. Tools Appl.1
2019 A multiple feature fused model for visual object tracking via correlation filters
Di Yuan 0002
Multim. Tools Appl.1
2018 Object tracking based on online representative sample selection via non-negative least square
Weihua Ou, Di Yuan 0002, Qiao Liu 0001, Yongfeng Cao
Multim. Tools Appl.2