Dawei Zhang 0002

dblp:76/5684-2 · DBLP profile ↗
← Back
27ranked-venue papers
11as first author
23since 2021 · last 2026
0000-0002-7593-1593ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 7 first-author · 11 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 IGIANet: Illumination Guided Implicit Alignment Network for Infrared-Visible UAV Detection
abstract
Visible-Infrared (RGB-IR) Unmanned Aerial Vehicle (UAV) object detection integrates complementary cues from visible and infrared sensors, offering broad application potential. However, due to sensor parallax, it still faces the challenge of weak spatial misalignment, which significantly limits its performance in UAV-based object detection. Existing methods emphasize strict alignment, overlooking spectral heterogeneity under varying illumination. To address these issues, we propose the Illumination Guided Implicit Alignment Network (IGIANet) to mitigate modality heterogeneity without explicit alignment. Specifically, we integrate three novel modules. First, we propose an illumination-guided frequency modulation module that adaptively allocates fusion weights to visible and infrared features based on global illumination estimation, effectively alleviating modality imbalance under varying lighting conditions. Second, we introduce a frequency-guided cross-modality differential enhancement module, which computes differential cues across frequency domains to enhance complementary information and highlight weakly aligned and low-contrast regions. Finally, we introduce an implicit alignment-driven dynamic fusion module that actively estimates offsets and generates dynamic, position-adaptive fusion kernels to align and fuse modalities. Extensive experiments demonstrate that IGIANet outperforms state-of-the-art models on various benchmarks, achieving 80.9% mAP on DroneVehicle, 57.1% mAP on VEDAI, and 49.4% mAP on FLIR.
Xiangqi Chen, Dawei Zhang 0002, Li Zhao 0005, Chengzhuan Yang, Jungang Lou, Zhonglong Zheng, Sang-Woon Jeon, Hua Wang 0002
AAAI2
2026 Exploiting All Mamba Fusion for Efficient RGB-D Tracking
abstract
Despite the progress made through deep learning, existing Visual Object Tracking (VOT) frameworks struggle with real-world challenges. Recent approaches incorporate additional modalities like Depth, Thermal Infrared, and Language to enhance the robustness of VOT, particularly with the improvement of the depth sensor precision, facilitating RGB-D tracking. However, current RGB-D trackers often copy RGB tracking paradigms, leading to inefficiency due to two-stream architectures that fail to exploit heterogeneous features, and reliance on simplistic or large-parameter fusion methods. To address these challenges, we propose AMTrack, a one-stream RGB-D tracker leveraging Mamba's linear complexity for simultaneous feature extraction and two-stage cross-modal feature fusion. Our innovation also includes a low-parameter Multimodal Mix Mamba (3M) module, which optimizes deep feature fusion and reduces computational overhead. The advantage of the 3M module stems from our Multimodal State Space Model (MSSM), a multimodal feature interaction component reconstructed based on SSM. Experiments across multiple RGB-D tracking datasets indicate that AMTrack achieves superior performance with lower parameters and memory demands compared to state-of-the-arts.
Ge Ying, Dawei Zhang 0002, Chengzhuan Yang, Wei Liu 0044, Sang-Woon Jeon, Hua Wang 0002, Changqin Huang, Zhonglong Zheng
AAAI2
2026 Data density scaling for text-to-image models on small dataset
Senmao Ye, Dawei Zhang 0002, Hehe Fan, Madal Artur, Hua Wang 0002, Zhonglong Zheng
Neurocomputing3
2026 A frequency mixing single-stream framework with LoRA prompt tuning for RGBD tracking
Dawei Zhang 0002, Kaiwei Jiang, Zhou Ou, Yufan Zhu, Zenan Zhou, Xiaowei He 0003, Zhonglong Zheng, Jun Zhang 0003
Neurocomputing1
2026 Template-Free Tracking Guidance for transformer trackers
Xuan Wang 0032, Li Zhao 0005, Dawei Zhang 0002, Chengzhuan Yang, Jungang Lou, Yunliang Jiang, Jinli Cao, Zhonglong Zheng
Knowl. Based Syst.3
2026 Boosting Open-Vocabulary Multiple Object Tracking With Wavelet Convolution and Confidence-Aware Kalman
Dawei Zhang 0002, Run Li, Xin Xiao 0006, Chengzhuan Yang, Zhonglong Zheng
IEEE Signal Process. Lett.1
2025 CGReg: Classification-Guided Point Cloud Registration via Equivariant Learning
Qinpeng Wu, Chengzhuan Yang, Lincong Fang, Dawei Zhang 0002, Zhonglong Zheng
PRCV (10)4
2025 GLKA-UNet: A Global-Local Aware UNet with KAN Attention for Infrared Small Target Detection
Xiangqi Chen, Chengzhuan Yang, Dawei Zhang 0002, Zhonglong Zheng
PRCV (15)5
2025 A review of object tracking based on deep learning
Guochen Zhao, Fanyong Meng 0004, Chengzhuan Yang, Hui Wei 0001, Dawei Zhang 0002, Zhonglong Zheng
Neurocomputing5
2025 CATrack: Condition-aware multi-object tracking with temporally enhanced appearance features
abstract
Multiple Object Tracking (MOT) is a critical task in computer vision with a wide range of practical applications. However, current methods often use a uniform approach for associating all targets, overlooking the varying conditions of each target. This can lead to performance degradation, especially in crowded scenes with dense targets. To address this issue, we propose a novel Condition-Aware Tracking method (CATrack) to differentiate the appearance feature flow for targets under different conditions. Specifically, we propose three designs for data association and feature update. First, we develop an Adaptive Appearance Association Module (AAAM) that selects suitable track templates based on detection conditions, reducing association errors in long-tail cases like occlusions or motion blur. Second, we design an ambiguous track filtering Selective Update strategy (SU) that filters out potential low-quality embeddings. Thus, the noise accumulation in the maintained track feature will also be reduced. Meanwhile, we propose a confidence-based Adaptive Exponential Moving Average (AEMA) method for the feature state transition. By adaptively adjusting the weights of track and detection embeddings, our AEMA better preserves high-quality target features. By integrating the above modules, CATrack enhances the discriminative capability of appearance features and improves the robustness of appearance-based associations. Extensive experiments on the MOT17 and MOT20 benchmarks validate the effectiveness of the proposed CATrack. Notably, the state-of-the-art results on MOT20 demonstrate the superiority of our method in highly crowded scenarios.
Run Li, Dawei Zhang 0002, Minglu Li 0001, Jinli Cao, Zhonglong Zheng
Knowl. Based Syst.3
2025 A complementary dual model for weakly supervised salient object detection
Dawei Zhang 0002, Xiao Wang 0014, Chang-Dong Wang 0001, Zhonglong Zheng
Pattern Recognit.2
2025 Temporal adaptive bidirectional bridging for RGB-D tracking
Ge Ying, Dawei Zhang 0002, Zhou Ou, Xiao Wang 0014, Zhonglong Zheng
Pattern Recognit.2
2025 Convolutional Attention Fusion for RGBT Tracking
abstract
RGBT target tracking accomplishes the tracking task by fusing visible and thermal infrared information. The development of Convolutional Neural Networks (CNNs) and Transformer has greatly advanced this field. Most existing transformer-based trackers focus on global modeling while neglecting the utilization of local information. In this paper, we propose a novel Convolutional Attention Fusion Module (CAFM) for RGBT target tracking. To be specific, this module continuously slides local windows on the image like convolution, and captures context features in each window like attention. Additionally, local position embedding is added to the window and works in conjunction with global position embedding to enhance the model's understanding of spatial information. Therefore, our CAFM enhances the extraction of local features by restricting the attention area and promotes multimodal fusion through cross-attention. We extend OSTrack to RGBT tracking and integrate the proposed CAFM into it. Experimental results show that our method performs well on the LasHeR, RGBT210, and RGBT234 datasets, and is superior to other advanced trackers.
Dawei Zhang 0002, Xuan Wang 0032, Xin Xiao 0006, Zhonglong Zheng
IEEE Signal Process. Lett.1
2025 Mask-Guided Frequency Feature Fusion for Visible-Infrared Remote Sensing Object Detection
abstract
Visible-infrared remote sensing object detection aims to achieve all-weather object detection by leveraging the complementary information from paired visible and infrared (RGB-IR) images. However, modality differences and weak alignment often limit its performance. Existing methods largely neglect the frequency discrepancies between modalities and require strict alignment, increasing complexity. To address these challenges, this study proposes a novel mask-guided frequency feature fusion (MGFF) method for RGB-IR object detection in remote sensing. Specifically, we develop a feature frequency decomposition and enhancement module using wavelet transform to reduce modality differences between RGB and IR images by restructuring and enhancing their frequency components. Additionally, we introduce a mask-guided feature reconstruction module and a feature-guided consistency loss, ensuring that even under weak alignment, the focus remains on integrating the target features from different modalities. Meanwhile, this loss is used to guide the reconstruction of features from different modalities. Finally, We design a multi-directional perception cross-modality fusion module to achieve deep fusion of multimodal information, which enhances object perception from different directions across modalities. Extensive evaluations on the widely recognized RGB-IR remote sensing benchmarks, including DroneVehicle and VEDAI, as well as the RGB-IR pedestrian dataset KAIST, substantiate the effectiveness of the proposed MGFF method. The results consistently demonstrate that the MGFF achieves a superior performance in terms of detection accuracy and robustness compared to existing state-of-the-art approaches.
Xiangqi Chen, Li Zhao 0005, Chengzhuan Yang, Dawei Zhang 0002, Xiao Wang 0014, Xiaowei He 0003, Hua Wang 0002, Zhonglong Zheng
IEEE Trans. Geosci. Remote. Sens.5
2025 Open-Vocabulary Multi-Object Tracking With Domain Generalized and Temporally Adaptive Features
abstract
Open-vocabulary multi-object tracking (OVMOT) is a cutting research direction within the multi-object tracking field. It employs large multi-modal models to effectively address the challenge of tracking unseen objects within dynamic visual scenes. While models require robust domain generalization and temporal adaptability, OVTrack, the only existing open-vocabulary multi-object tracker, relies solely on static appearance information and lacks these crucial adaptive capabilities. In this paper, we propose OVSORT, a new framework designed to improve domain generalization and temporal information processing. Specifically, we first propose the Adaptive Contextual Normalization (ACN) technique in OVSORT, which dynamically adjusts the feature maps based on the dataset's statistical properties, thereby fine-tuning our model's to improve domain generalization. Then, we introduce motion cues for the first time. Using our Joint Motion and Appearance Tracking (JMAT) strategy, we obtain a joint similarity measure and subsequently apply the Hungarian algorithm for data association. Finally, our Hierarchical Adaptive Feature Update (HAFU) strategy adaptively adjusts feature updates according to the current state of each trajectory, which greatly improves the utilization of temporal information. Extensive experiments on the TAO validation set and test set confirm the superiority of OVSORT, which significantly improves the handling of novel and base classes. It surpasses existing methods in terms of accuracy and generalization, setting a new state-of-the-art for OVMOT.
Run Li, Dawei Zhang 0002, Yunliang Jiang, Zhonglong Zheng, Sang-Woon Jeon, Hua Wang 0002
IEEE Trans. Multim.2
2024 When decoupled GCN meets group discrimination: A special graph contrastive learning framework
Yinjie Gao, Dawei Zhang 0002, Jinli Cao, Zhonglong Zheng
Neurocomputing4
2024 Improved SiamCAR with ranking-based pruning and optimization for efficient UAV tracking
Xiaoqiang Jin, Dawei Zhang 0002, Qiner Wu, Xin Xiao 0006, Pengsen Zhao, Zhonglong Zheng
Image Vis. Comput.2
2024 Probabilistic Assignment With Decoupled IoU Prediction for Visual Tracking
abstract
Modern Siamese trackers mainly rely on classifying and regressing pre-defined anchor boxes or per-pixel points, which are assigned as positive and negative samples based on box intersection-over-union (IoU) or point distance with corresponding ground-truth for training. However, this rigid configuration potentially involves some noisy and ambiguous positive samples, leading to an inconsistency problem between classification and regression, which limits the tracking performance. In this paper, we propose a novel probabilistic assignment approach that dynamically determines positive/negative samples for each instance. To be specific, we first customize the confidence scores of positive candidates by comprehensively exploring the outputs from both classification and regression heads, and fit these scores as a probability distribution. Therefore, it is intuitive to conduct adaptive label assignment according to their probabilities. Then, we also consider dynamic re-weighting factor for each positive sample, jointly optimizing the classification and regression losses in a synchronized manner. Moreover, we introduce a decoupled IoU prediction branch to bridge the gap between the training and inference objectives for accurate tracking. Thanks to well-aligned procedures, our method significantly improves the performance of both CNN-based and Transformer-based trackers. Extensive experiments conducted on several tracking benchmarks including LaSOT and GOT-10k, demonstrate the effectiveness and efficiency of the proposed probabilistic assignment tracker.
Dawei Zhang 0002, Xin Xiao 0006, Zhonglong Zheng, Yunliang Jiang
IEEE Trans. Circuits Syst. Video Technol.1
2022 UAST: Uncertainty-Aware Siamese Tracking
abstract
Visual object tracking is basically formulated as target classification and bounding box estimation. Recent anchor-free Siamese trackers rely on predicting the distances to four sides for efficient regression but fail to estimate accurate bounding box in complex scenes. We argue that these approaches lack a clear probabilistic explanation, so it is desirable to model the uncertainty and ambiguity representation of target estimation. To address this issue, this paper presents an Uncertainty-Aware Siamese Tracker (UAST) by developing a novel distribution-based regression formulation with localization uncertainty. We exploit regression vectors to directly represent the discretized probability distribution for four offsets of boxes, which is general, flexible and informative. Based on the resulting distributed representation, our method is able to provide a probabilistic value of uncertainty. Furthermore, considering the high correlation between the uncertainty and regression accuracy, we propose to learn a joint representation head of classification and localization quality for reliable tracking, which also avoids the inconsistency of classification and quality estimation between training and inference. Extensive experiments on several challenging tracking benchmarks demonstrate the effectiveness of UAST and its superiority over other Siamese trackers.
Dawei Zhang 0002, Yanwei Fu 0001, Zhonglong Zheng
ICML1
2021 Visual Tracking via Hierarchical Deep Reinforcement Learning
abstract
Visual tracking has achieved great progress due to numerous different algorithms. However, deep trackers based on classification or Siamese network still have their specific limitations. In this work, we show how to teach machines to track a generic object in videos like humans, who can use a few search steps to perform tracking. By constructing a Markov decision process in Deep Reinforcement Learning (DRL), our agents can learn to determine hierarchical decisions on tracking mode and motion estimation. To be specific, our Hierarchical DRL framework is composed of a Siamese-based observation network which models the motion information of an arbitrary target, a policy network for mode switch and an actor-critic network for box regression. This tracking strategy is more in line with human behavior paradigm, and is effective and efficient to cope with fast motion, background clutter and large deformations. Extensive experiments on the GOT-10k, OTB-100, UAV-123, VOT and LaSOT tracking benchmarks, demonstrate that the proposed tracker achieves state-of-the-art performance while running in real-time.
Dawei Zhang 0002, Zhonglong Zheng, Riheng Jia, Minglu Li 0001
AAAI1
2021 Single Image Super-Resolution Via Global-Context Attention Networks
abstract
In the last few years, single image super-resolution (SISR) has benefited a lot from the rapid development of deep convolutional neural networks (CNNs), and the introduction of attention mechanisms further improves the performance of SISR. However, previous methods use one or more types of attention independently in multiple stages and ignore the correlations between different layers in the network. To address these issues, we propose a novel end-to-end architecture named global-context attention network (GCAN) for SISR, which consists of several residual global-context attention blocks (RGCABs) and an inter-group fusion module (IGFM). Specifically, the proposed RGCAB extracts representative features that capture non-local spatial interdependencies and multiple channel relations. Then the IGFM aggregates and fuses hierarchical features of multi-layers discriminatively by considering correlations among layers. Extensive experimental results demonstrate that our method achieves superior results against other state-of-the-art methods on publicly available datasets.
Pengcheng Bian, Zhonglong Zheng, Dawei Zhang 0002, Minglu Li 0001
ICIP3
2021 Light-Weight Multi-channel Aggregation Network for Image Super-Resolution
Pengcheng Bian, Zhonglong Zheng, Dawei Zhang 0002
PRCV (3)3
2021 CSART: Channel and spatial attention-guided residual learning for real-time object tracking
Dawei Zhang 0002, Zhonglong Zheng, Minglu Li 0001, Rixian Liu
Neurocomputing1
2020 High Performance Visual Tracking With Siamese Actor-Critic Network
abstract
Object tracking is one of the fundamental tasks of computer vision and it is still a major challenge that trackers can balance between real-time speed and high performance. In this paper, a novel Siamese Actor-Critic network (SiamAC) is proposed to improve the accuracy and robustness of tracking while performing with real-time. Specifically, SiamAC consists of a fully convolutional siamese matching network for similarity learning and an Actor-Critic framework trained by reinforcement learning. The Actor is aimed to infer the optimal action in continuous space, while the Critic produces a Q-value to guide effectively the offline training of both Actor and Critic networks. During inference, according to the response map produced by the matching network, the most similar positions and scaled candidate patches are selected as the input of Actor-Critic. Subsequently, Actor can effectively search more precise location of these candidates and the Critic acts as a validator to decide the final tracking results with the highest confidence. Benefiting from this refinement, traditional multi-scale test and certain hyper-parameters in Siamese trackers can be discarded. Evaluations on popular benchmarks demonstrate that the proposed SiamAC achieves state-of-the-art performance with a real-time speed.
Dawei Zhang 0002, Zhonglong Zheng
ICIP1
2020 Joint Representation Learning with Deep Quadruplet Network for Real-Time Visual Tracking
abstract
Recently, trackers based on Siamese networks have attracted spread attention in the field of tracking because of a balance between accuracy and speed. Learning powerful representation via effective offline training strategy is critical for constructing high performance Siamese trackers. However, features extracted in most networks cannot accurately distinguish a tracked target from the background with semantic information in some challenging scenes. In this paper, we develop a Fully-Convolutional deep Quadruple Network (QuadFC) to learn more expressive representation via a novel multi-task loss function composed of a differential pairwise loss for tracking and a constructed triplet loss for similarity learning, which can be trained offline in an end-to-end mode. During inference, the proposed deep architecture does not need to update model and the positive-negative branches are removed to avoid unnecessary calculations. In particular, our approach is able to extract more discriminative features and perform robust visual tracking, due to joint representation learning and taking full use of original samples via the combination of positive-negative pairs. Furthermore, theoretical analysis of QuadFC is carried out through comparing the gradients of different loss functions. Extensive experiments on several tracking benchmarks, show that the proposed tracker achieves the state-of-the-art tracking performance while running at 68 FPS. The code can be available at https://github.com/DavidZhangdw/QuadFC.
Dawei Zhang 0002, Zhonglong Zheng
IJCNN1
2020 Learning Fine-Grained Similarity Matching Networks for Visual Tracking
abstract
Recently, siamese trackers have been increasingly popular in visual tracking community. Despite great success, it is still difficult to perform robust tracking in various challenging scenarios. In this paper, we propose a novel similarity matching network, that effectively extracts fine-grained semantic features by adding a Classification branch and a Category-Aware module into the classical Siamese framework (CCASiam). More specifically, the supervision module can fully utilize the class information to obtain a loss for classification and the whole network performs tracking loss, so that the network can extract more discriminative features for each specific target. During online tracking, the classification branch is removed and the category-aware module is designed to guide the selection of target-active features using a ridge regression network, which avoids unnecessary calculations and over-fitting. Furthermore, we introduce different types of attention mechanisms to selectively emphasize important semantic information. Due to the fine-grained and category-aware features, CCASiam can perform high performance tracking efficiently. Extensive experimental results on several tracking benchmarks, show that the proposed tracker obtains the state-of-the-art performance with a real-time speed.
Dawei Zhang 0002, Zhonglong Zheng, Xiaowei He 0003, Liu Su
ICMR1
2020 Reinforced Similarity Learning: Siamese Relation Networks for Robust Object Tracking
abstract
Recently, Siamese networks based tracking algorithms have shown favorable performance. Latest work focuses on better feature embedding and target state estimation, which greatly improves the accuracy. Nevertheless, the simple cross-correlation operation of the features between a fixed template and the search region limits their robustness and discrimination capability. In this paper, we pay more attention to learn an outstanding similarity measure for robust tracking. We propose a novel relation network that can be integrated on top of previous trackers without any need for further training of the siamese networks, which achieves a superior discriminative ability. During online inference, we utilize the feedback from high-confidence tracking results to obtain an additional template and update it, which improves the robustness and generalization. We implement two versions of the proposed approach with the SiamFC-based tracker and SiamRPN-based tracker to validate the strong compatibility of our algorithm. Extensive experimental results on several tracking benchmarks indicate that the proposed method can effectively improve the performance and robustness of the underlying trackers without reducing speed too much, and performs superiorly against the state-of-the-art trackers.
Dawei Zhang 0002, Zhonglong Zheng, Minglu Li 0001, Xiaowei He 0003, Riheng Jia, Feilong Lin
ACM Multimedia1