Jie Zhao 0014

dblp:23/3168-14 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
16since 2021 · last 2025
0000-0002-3924-5537ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 CAT: A Unified Click-and-Track Framework for Realistic Tracking
Yongsheng Yuan, Jie Zhao 0014, Dong Wang 0004, Huchuan Lu
ICCV2
2025 Efficient Motion Prompt Learning for Robust Visual Tracking
abstract
Due to the challenges of processing temporal information, most trackers depend solely on visual discriminability and overlook the unique temporal coherence of video data. In this paper, we propose a lightweight and plug-and-play motion prompt tracking method. It can be easily integrated into existing vision-based trackers to build a joint tracking framework leveraging both motion and vision cues, thereby achieving robust tracking through efficient prompt learning. A motion encoder with three different positional encodings is proposed to encode the long-term motion trajectory into the visual embedding space, while a fusion decoder and an adaptive weight mechanism are designed to dynamically fuse visual and motion features. We integrate our motion module into three different trackers with five models in total. Experiments on seven challenging tracking benchmarks demonstrate that the proposed motion module significantly improves the robustness of vision-based trackers, with minimal training costs and negligible speed sacrifice. Code is available at https://github.com/zj5559/Motion-Prompt-Tracking.
Jie Zhao 0014, Xin Chen 0032, Yongsheng Yuan, Michael Felsberg, Dong Wang 0004, Huchuan Lu
ICML1
2025 Exploiting Lightweight Hierarchical ViT and Dynamic Framework for Efficient Visual Tracking
abstract
Abstract Transformer-based visual trackers have demonstrated significant advancements due to their powerful modeling capabilities. However, their practicality is limited on resource-constrained devices because of their slow processing speeds. To address this challenge, we present HiT, a novel family of efficient tracking models that achieve high performance while maintaining fast operation across various devices. The core innovation of HiT lies in its Bridge Module, which connects lightweight transformers to the tracking framework, enhancing feature representation quality. Additionally, we introduce a dual-image position encoding approach to effectively encode spatial information. HiT achieves an impressive speed of 61 frames per second (fps) on the NVIDIA Jetson AGX platform, alongside a competitive AUC of 64.6% on the LaSOT benchmark, outperforming all previous efficient trackers. Building on HiT, we propose DyHiT, an efficient dynamic tracker that flexibly adapts to scene complexity by selecting routes with varying computational requirements. DyHiT uses search area features extracted by the backbone network and inputs them into an efficient dynamic router to classify tracking scenarios. Based on the classification, DyHiT applies a divide-and-conquer strategy, selecting appropriate routes to achieve a superior trade-off between accuracy and speed. The fastest version of DyHiT achieves 111 fps on NVIDIA Jetson AGX while maintaining an AUC of 62.4% on LaSOT. Furthermore, we introduce a training-free acceleration method based on the dynamic routing architecture of DyHiT. This method significantly improves the execution speed of various high-performance trackers without sacrificing accuracy. For instance, our acceleration method enables the state-of-the-art tracker SeqTrack-B256 to achieve a $$2.68\times $$ 2.68 × speedup on an NVIDIA GeForce RTX 2080 Ti GPU while maintaining the same AUC of 69.9% on the LaSOT. Codes, models, and results are available at https://github.com/kangben258/HiT .
Ben Kang, Xin Chen 0032, Jie Zhao 0014, Chunjuan Bo, Dong Wang 0004, Huchuan Lu
Int. J. Comput. Vis.3
2025 Learning Language Prompt for Vision-Language Tracking
abstract
Vision-language object tracking integrates advanced linguistic information, enhancing its robustness and accuracy in complex scenarios. Nevertheless, current methods are constrained by a lack of sufficient vision-language data, making it challenging for the model to learn generalized knowledge. To alleviate this issue, we propose a new prompt-based framework for vision-language tracking, named ProVLT. This framework casts language information as a prompt for pretrained visionbased tracking models, thereby leveraging the knowledge from extensive tracking data. Experiments demonstrate that ProVLT achieves competitive performance while training only a fraction of parameters (approximately 29% of modal parameters). For instance, ProVLT achieves competitive performance, attaining AUC of 59.8% on TNL2K benchmark. Furthermore, we augment five mainstream vision-only tracking benchmarks with language annotations, and find that the inclusion of linguistic information consistently improves tracking performance. On these benchmarks, the linguistic information improves the performance by an average of 2.9% compared with the vision-based tracker. We will release the code, models, and benchmarks for the community.
ChengAo Zong, Jie Zhao 0014, Xin Chen 0032, Huchuan Lu, Dong Wang 0004
IEEE Trans. Circuits Syst. Video Technol.2
2025 Self-Adaptive Vision-Language Tracking With Context Prompting
abstract
Due to the substantial gap between vision and language modalities, along with the mismatch problem between fixed language descriptions and dynamic visual information, existing vision-language tracking methods exhibit performance on par with or slightly worse than vision-only tracking. Effectively exploiting the rich semantics of language to enhance tracking robustness remains an open challenge. To address these issues, we propose a self-adaptive vision-language tracking framework that leverages the pre-trained multi-modal CLIP model to obtain well-aligned visual-language representations. A novel context-aware prompting mechanism is introduced to dynamically adapt linguistic cues based on the evolving visual context during tracking. Specifically, our context prompter extracts dynamic visual features from the current search image and integrates them into the text encoding process, enabling self-updating language embeddings. Furthermore, our framework employs a unified one-stream Transformer architecture, supporting joint training for both vision-only and vision-language tracking scenarios. Our method not only bridges the modality gap but also enhances robustness by allowing language features to evolve with visual context. Extensive experiments on four vision-language tracking benchmarks demonstrate that our method effectively leverages the advantages of language to enhance visual tracking. Our large model can obtain 55.0% AUC on $\text {LaSOT}_{\text {EXT}}$ and 69.0% AUC on TNL2K. Additionally, our language-only tracking model achieves performance comparable to that of state-of-the-art vision-only tracking methods on TNL2K. Code is available at https://github.com/zj5559/SAVLT.
Jie Zhao 0014, Xin Chen 0032, Shengming Li, Chunjuan Bo, Dong Wang 0004, Huchuan Lu
IEEE Trans. Image Process.1
2025 Enhancing the Two-Stream Framework for Efficient Visual Tracking
abstract
Practical deployments, especially on resource-limited edge devices, necessitate high speed for visual object trackers. To meet this demand, we introduce a new efficient tracker with a Two-Stream architecture, named ToS. While the recent one-stream tracking framework, employing a unified backbone for simultaneous processing of both the template and search region, has demonstrated exceptional efficacy, we find the conventional two-stream tracking framework, which employs two separate backbones for the template and search region, offers inherent advantages. The two-stream tracking framework is more compatible with advanced lightweight backbones and can efficiently utilize benefits from large templates. We demonstrate that the two-stream setup can exceed the one-stream tracking model in both speed and accuracy through strategic designs. Our methodology rejuvenates the two-stream tracking paradigm with lightweight pre-trained backbones and the proposed three efficient strategies: 1) A feature-aggregation module that improves the representation capability of the backbone, 2) A channel-wise approach for feature fusion, presenting a more effective and lighter alternative to spatial concatenation techniques, and 3) An expanded template strategy to boost tracking accuracy with negligible additional computational cost. Extensive evaluations across multiple tracking benchmarks demonstrate that the proposed method sets a new state-of-the-art performance in efficient visual tracking.
ChengAo Zong, Xin Chen 0032, Jie Zhao 0014, Yang Liu 0066, Huchuan Lu, Dong Wang 0004
IEEE Trans. Image Process.3
2024 Learning Coverage Paths in Unknown Environments with Deep Reinforcement Learning
abstract
Coverage path planning (CPP) is the problem of finding a path that covers the entire free space of a confined area, with applications ranging from robotic lawn mowing to search-and-rescue. When the environment is unknown, the path needs to be planned online while mapping the environment, which cannot be addressed by offline planning methods that do not allow for a flexible path space. We investigate how suitable reinforcement learning is for this challenging problem, and analyze the involved components required to efficiently learn coverage paths, such as action space, input feature representation, neural network architecture, and reward function. We propose a computationally feasible egocentric map representation based on frontiers, and a novel reward term based on total variation to promote complete coverage. Through extensive experiments, we show that our approach surpasses the performance of both previous RL-based approaches and highly specialized methods across multiple CPP variations.
Arvi Jonnarth, Jie Zhao 0014, Michael Felsberg
ICML2
2024 Leveraging the Power of Data Augmentation for Transformer-based Tracking
abstract
Due to long-distance correlation and powerful pretrained models, transformer-based methods have initiated a breakthrough in visual object tracking performance. Previous works focus on designing effective architectures suited for tracking, but ignore that data augmentation is equally crucial for training a well-performing model. In this paper, we first explore the impact of general data augmentations on transformer-based trackers via systematic experiments, and reveal the limited effectiveness of these common strategies. Motivated by experimental observations, we then propose two data augmentation methods customized for tracking. First, we optimize existing random cropping via a dynamic search radius mechanism and simulation for boundary samples. Second, we propose a token-level feature mixing augmentation strategy, which enables the model against challenges like background interference. Extensive experiments on two transformer-based trackers and six benchmarks demonstrate the effectiveness and data efficiency of our methods, especially under challenging settings, like one-shot tracking and small image resolutions. Code is available at https://github.com/zj5559/DATr.
Jie Zhao 0014, Johan Edstedt, Michael Felsberg, Dong Wang 0004, Huchuan Lu
WACV1
2024 LGTrack: Exploiting Local and Global Properties for Robust Visual Tracking
abstract
Re-detection is a necessary capability for long-term tracking. Target candidate proposals in the whole image can provide a chance of tracking reset when tracking fails due to tracking drift or target invisibility. In this paper, we propose a unified local-global tracker based on the same transformer architecture sharing weights, which can not only search in a continuous local region but also provide target candidates of the global image in every frame. The requirements of both long-term and short-term scenarios can be addressed using a unified model. A simple proposal selection scheme is adopted to properly select the candidate proposals of re-detection, to assist tracking and obtain better performance. The scheme performs reevaluation of all high-quality proposals based on a transformer-based embedding network, once the predicted state of the local tracking is not sufficient to be accurate. To capture appearance variations brought by online updates in minimum risks, a long-term-friendly dynamic template update scheme is also designed. Extensive experiments are conducted to demonstrate the effectiveness of our proposed tracker, including three short-term tracking benchmarks and six long-term benchmarks. Our tracker can achieve results comparable to that of the state-of-the-art. The proposed tracker can also work well in balancing the performance and speed, achieving an average speed of approximately 25 fps tested on LaSOT testing set.
Chang Liu 0071, Jie Zhao 0014, Chunjuan Bo, Shengming Li, Dong Wang 0004, Huchuan Lu
IEEE Trans. Circuits Syst. Video Technol.2
2023 Predicting cancer outcomes from whole slide images via hybrid supervision learning
Xianying He, Jiahui Li 0005, Fang Yan 0002, Wen Chen 0001, Qi Duan, Hongsheng Li 0001, Shaoting Zhang 0001, Jie Zhao 0014
Neurocomputing11
2023 Robust Online Tracking With Meta-Updater
abstract
In a sequence, the appearance of both the target and background often changes dramatically. Offline-trained models may not handle huge appearance variations well, causing tracking failures. Most discriminative trackers address this issue by introducing an online update scheme, making the model dynamically adapt the changes of the target and background. Although the online update scheme plays an important role in improving the tracker's accuracy, it inevitably pollutes the model with noisy observation samples. It is necessary to reduce the risk of the online update scheme for better tracking. In this work, we propose a novel offline-trained Meta-Updater to address an important but unsolved problem: Is the tracker ready for updating in the current frame? The proposed module can effectively integrate geometric, discriminative, and appearance cues in a sequential manner, and then mine the sequential information with a designed cascaded LSTM module. Moreover, we strengthen the effect of appearance information on the module, i.e., the additional local outlier factor is introduced to integrate into a newly designed network. We integrate our meta-updater into eight different types of online update trackers. Extensive experiments on four long-term and two short-term tracking benchmarks demonstrate that our meta-updater is effective and has strong generalization ability.
Jie Zhao 0014, Kenan Dai, Dong Wang 0004, Huchuan Lu
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Visible-Thermal UAV Tracking: A Large-Scale Benchmark and New Baseline
abstract
With the popularity of multi-modal sensors, visible-thermal (RGB-T) object tracking is to achieve robust performance and wider application scenarios with the guidance of objects' temperature information. However, the lack of paired training samples is the main bottleneck for unlocking the power of RGB-T tracking. Since it is laborious to collect high-quality RGB-T sequences, recent benchmarks only provide test sequences. In this paper, we construct a large-scale benchmark with high diversity for visible-thermal UAV tracking (VTUAV), including 500 sequences with 1.7 million high-resolution (1920* 1080 pixels) frame pairs. In addition, comprehensive applications (short-term tracking, long-term tracking and segmentation mask prediction) with diverse categories and scenes are considered for exhaustive evaluation. Moreover, we provide a coarse-to-fine attribute annotation, where frame-level attributes are provided to exploit the potential of challenge-specific trackers. In addition, we design a new RGB-T baseline, named Hierarchical Multi-modal Fusion Tracker (HMFT), which fuses RGB-T data in various levels. Numerous experiments on several datasets are conducted to reveal the effectiveness of HMFT and the complement of different fusion types. The project is available at here.
Jie Zhao 0014, Dong Wang 0004, Huchuan Lu, Xiang Ruan
CVPR2
2022 DigestPath: A benchmark dataset with challenge review for the pathological detection and segmentation of digestive-system
Qian Da, Zhongyu Li 0002, Yanfei Zuo, Chenbin Zhang, Jingxin Liu 0005, Wen Chen 0001, Jiahui Li 0005, Dou Xu, Hongmei Yi, Zhe Wang 0043, Li Zhang 0040, Xianying He, Xiaofan Zhang 0002, Ke Mei, Chuang Zhu, Weizeng Lu, LinLin Shen, Jun Shi 0006, Jun Li 0106, Sreehari S, Ganapathy Krishnamurthi, Jiangcheng Yang, Tiancheng Lin 0001, Qingyu Song 0004, Xuechen Liu 0004, Simon Graham, Raja Muhammad Saad Bashir, Canqian Yang, Shaofei Qin, Xinmei Tian 0001, Jie Zhao 0014, Dimitris N. Metaxas, Hongsheng Li 0001, Chaofu Wang, Shaoting Zhang 0001
Medical Image Anal.36
2022 Vision-Based Anti-UAV Detection and Tracking
abstract
Unmanned aerial vehicles (UAV) have been widely used in various fields, and their invasion of security and privacy has aroused social concern. Several detection and tracking systems for UAVs have been introduced in recent years, but most of them are based on radio frequency, radar, and other media. We assume that the field of computer vision is mature enough to detect and track invading UAVs. Thus we propose a visible light mode dataset called Dalian University of Technology Anti-UAV dataset, DUT Anti-UAV for short. It contains a detection dataset with a total of 10,000 images and a tracking dataset with 20 videos that include short-term and long-term sequences. All frames and images are manually annotated precisely. We use this dataset to train several existing detection algorithms and evaluate the algorithms’ performance. Several tracking methods are also tested on our tracking dataset. Furthermore, we propose a clear and simple tracking algorithm combined with detection that inherits the detector’s high precision. Extensive experiments show that the tracking performance is improved considerably after fusing detection, thus providing a new attempt at UAV tracking using our dataset. The datasets and results are publicly available at:https://github.com/wangdongdut/DUT-Anti-UAV.
Jie Zhao 0014, Jingshu Zhang, Dongdong Li 0004, Dong Wang 0004
IEEE Trans. Intell. Transp. Syst.1
2021 Video Annotation for Visual Tracking via Selection and Refinement
abstract
Deep learning based visual trackers entail offline pre-training on large volumes of video datasets with accurate bounding box annotations that are labor-expensive to achieve. We present a new framework to facilitate bounding box annotations for video sequences, which investigates a selection-and-refinement strategy to automatically improve the preliminary annotations generated by tracking algorithms. A temporal assessment network (T-Assess Net) is proposed which is able to capture the temporal coherence of target locations and select reliable tracking results by measuring their quality. Meanwhile, a visual-geometry refinement network (VG-Refine Net) is also designed to further enhance the selected tracking results by considering both target appearance and temporal geometry constraints, allowing inaccurate tracking results to be corrected. The combination of the above two networks provides a principled approach to ensure the quality of automatic video annotation. Experiments on large scale tracking benchmarks demonstrate that our method can deliver highly accurate bounding box annotations and significantly reduce human labor by 94.0%, yielding an effective means to further boost tracking performance with augmented training data.
Kenan Dai, Jie Zhao 0014, Lijun Wang 0001, Dong Wang 0004, Huchuan Lu, Xuesheng Qian, Xiaoyun Yang
ICCV2
2021 Jointly Modeling Motion and Appearance Cues for Robust RGB-T Tracking
abstract
In this study, we propose a novel RGB-T tracking framework by jointly modeling both appearance and motion cues. First, to obtain a robust appearance model, we develop a novel late fusion method to infer the fusion weight maps of both RGB and thermal (T) modalities. The fusion weights are determined by using offline-trained global and local multimodal fusion networks, and then adopted to linearly combine the response maps of RGB and T modalities. Second, when the appearance cue is unreliable, we comprehensively take motion cues, i.e., target and camera motions, into account to make the tracker robust. We further propose a tracker switcher to switch the appearance and motion trackers flexibly. Numerous results on three recent RGB-T tracking datasets show that the proposed tracker performs significantly better than other state-of-the-art algorithms.
Jie Zhao 0014, Chunjuan Bo, Dong Wang 0004, Huchuan Lu, Xiaoyun Yang
IEEE Trans. Image Process.2
2020 Online Filtering Training Samples for Robust Visual Tracking
abstract
In recent years, discriminative trackers show its great tracking performance, that is mainly due to the online updating using samples collected during tracking. The model could adapt appearance changes of objects and the background well after updating. But these trackers have a serious disadvantage that wrong samples may cause severe model degradation. Most of the training samples in the tracking phase are obtained according to the tracking result of the current frame. Wrong training samples will be collected when the tracking result is inaccurate, seriously affecting the discrimination ability of the model. Besides, partial occlusion also leads to the same problem. In this paper, we propose an optimization module named MetricNet for online filtering training samples. It applies a matching network containing the classification and distance branches, and uses multiple metric methods for different type samples. MetricNet optimizes the training sample set by recognizing wrong and redundant samples, thereby improving the tracking performance. The proposed MetricNet can be regarded as an independent optimization module and integrated into all discriminative trackers updated online. Extensive experiments on three tracking datasets show its effectiveness and generalization ability. After applying MetricNet to MDNet, the tracking result is increased by 5.3% in terms of the success plot on the LaSOT dataset. Our project is available at https://github.com/zj5559/MetricNet.
Jie Zhao 0014, Kenan Dai, Dong Wang 0004, Huchuan Lu, Xiaoyun Yang
ACM Multimedia1