EDBT 2026 Demo / reviewers in the wild / expert
Shuiwang Li
dblp:160/6992
· DBLP profile ↗
37ranked-venue papers
6as first author
36since 2021 · last 2026
0000-0002-4587-513XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 3 first-author · 23 since 2021Artificial intelligence and machine learning · 17 · 2 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CapeNext: Rethinking and Refining Dynamic Support Information for Category-Agnostic Pose EstimationabstractRecent research in Category-Agnostic Pose Estimation (CAPE) has adopted fixed textual keypoint description as semantic prior for two-stage pose matching frameworks. While this paradigm enhances robustness and flexibility by disentangling the dependency of support images, our critical analysis reveals two inherent limitations of static joint embedding: (1) polysemy-induced cross-category ambiguity during the matching process(e.g., the concept "leg" exhibiting divergent visual manifestations across humans and furniture), and (2) insufficient discriminability for fine-grained intra-category variations (e.g., posture and fur discrepancies between a sleeping white cat and a standing black cat). To overcome these challenges, we propose a new framework that innovatively integrates hierarchical cross-modal interaction with dual-stream feature refinement, enhancing the joint embedding with both class-level and instance-specific cues from textual description and specific images. Experiments on the MP-100 dataset demonstrate that, regardless of the network backbone, CapeNext consistently outperforms state-of-the-art CAPE methods by a large margin. Dan Zeng 0002, Shuiwang Li, Qijun Zhao, Qiaomu Shen, Bo Tang 0016 |
AAAI | 3 |
| 2026 | Learning motion blur robust vision transformers for real-time UAV tracking
You Wu 0009, Xucheng Wang, Dan Zeng 0002, Hengzhou Ye, Xiaolan Xie 0002, Qijun Zhao, Shuiwang Li |
Expert Syst. Appl. | 7 |
| 2026 | SMTrack: End-to-End Trained Spiking Neural Networks for Multi-Object Tracking in RGB VideosabstractBrain-inspired Spiking Neural Networks (SNNs) leverage a sparse, event-driven computational paradigm and have shown great potential for low-power object tracking. However, most existing SNN-based object tracking rely on event camera data, whereas traditional RGB video remains the dominant input modality in real-world applications. Research on RGB-based SNN multi-object tracking, particularly directly trained deep SNN models, is still in its infancy. To address this, we propose SMTrack, the first directly trained deep SNN framework for end-to-end multi-object tracking on standard RGB data. To handle the challenges caused by scale and density variations among objects, we introduce an Adaptive Scale-aware Normalized Wasserstein Distance Loss (Asa-NWDLoss), which dynamically adjusts the normalization factor based on the average object size within each training batch. For the identity association stage, we integrate the TrackTrack to maintain robust and consistent trajectory tracking. Extensive experiments on BEE24, MOT17, MOT20, and DanceTrack demonstrate that SMTrack achieves comparable performance to mainstream ANN-based approaches with only a few time steps. Codes is available at https://github.com/OpenCodeGithub/SMTrack. Pengzhi Zhong, Dan Zeng 0002, Qihua Zhou, Feixiang He, Shuiwang Li |
IEEE Internet Things J. | 6 |
| 2026 | UniTrack: Unifying day and night tracking with continual learning
Jiwei Mo, Feixiang He, Pengzhi Zhong, Qijun Zhao, Dan Zeng 0002, Shuiwang Li, Xianhao Shen |
Inf. Sci. | 7 |
| 2026 | Learning an Adaptive and View-Invariant Vision Transformer for Real-Time UAV TrackingabstractTransformer-based models have improved visual tracking, but most still cannot run in real time on resource-limited devices, especially for unmanned aerial vehicle (UAV) tracking. To achieve a better balance between performance and efficiency, we propose AVTrack, an adaptive computation tracking framework that adaptively activates transformer blocks through an Activation Module (AM), which dynamically optimizes the ViT architecture by selectively engaging relevant components. To address extreme viewpoint variations, we propose to learn view-invariant representations via mutual information (MI) maximization. In addition, we propose AVTrack-MD, an enhanced tracker incorporating a novel MI maximization-based multi-teacher knowledge distillation framework. Leveraging multiple off-the-shelf AVTrack models as teachers, we maximize the MI between their aggregated softened features and the corresponding softened feature of the student model, improving the generalization and performance of the student, especially under noisy conditions. Extensive experiments show that AVTrack-MD achieves performance comparable to AVTrack’s performance while reducing model complexity and boosting average tracking speed by over 17%. Codes is available at https://github.com/wuyou3474/AVTrack. You Wu 0009, Xucheng Wang, Xiangyang Yang 0001, Hengzhou Ye, Dan Zeng 0002, Qijun Zhao, Shuiwang Li |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2026 | Toward Real-Time UAV Tracking With Adaptive and Background-Aware Vision TransformersabstractIn Unmanned Aerial Vehicle (UAV) tracking, discriminative correlation filters (DCF) are popular for their speed, making them ideal for real-time use with limited resources. Recently, lightweight convolutional neural networks (CNNs) have offered a new approach. Through filter pruning, these CNNs maintain high accuracy and efficiency, making them a strong alternative, especially for greater precision. Despite these advancements, the potential of pure vision transformers (ViTs) in UAV tracking remains largely untapped, especially based on the paradigm of conditional computation. In this work, we introduce an adaptive and background-aware Vision Transformer (Aba-ViT) and leverage it to develop a real-time UAV tracking framework called Aba-ViTrack. The proposed Aba-ViT exploits an adaptive and background-aware token computation method to reduce inference time. This approach adaptively discards tokens based on learned halting probabilities, which a priori are higher for background tokens than target ones. To further improve efficiency, this paper proposes a novel classroom-style learning (CSL) approach, where robustness knowledge is transmitted vertically from teacher to students, and generalization capability is enhanced through horizontal mutual learning among students. This method is used to compress Aba-ViTrack, resulting in Aba-ViTrack++. The upgraded version achieves a better balance between accuracy and efficiency in real-time UAV tracking. This version achieves a better balance between accuracy and efficiency for real-time UAV tracking. Extensive experiments on six UAV tracking benchmarks demonstrate that the proposed method achieves state-of-the-art performance in UAV tracking. The code is available at https://github.com/xyyang317/Aba-ViTrack. Xiangyang Yang 0001, Dan Zeng 0002, Xucheng Wang, Hengzhou Ye, Xiaolan Xie 0002, Qijun Zhao, Jihua Zhu, Shuiwang Li |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2025 | Towards Reflected Object Detection: A Benchmark
Zhongtian Wang, You Wu 0009, Shuiwang Li |
CVM (1) | 6 |
| 2025 | Learning Occlusion-Robust Vision Transformers for Real-Time UAV TrackingabstractSingle-stream architectures using Vision Transformer (ViT) backbones show great potential for real-time UAV tracking recently. However, frequent occlusions from obstacles like buildings and trees expose a major drawback: these models often lack strategies to handle occlusions effectively. New methods are needed to enhance the occlusion resilience of single-stream ViT models in aerial tracking. In this work, we propose to learn Occlusion-Robust Representations (ORR) based on ViTs for UAV tracking by enforcing an invariance of the feature representation of a target with respect to random masking operations modeled by a spatial Cox process. Hopefully, this random masking approximately simulates target occlusions, thereby enabling us to learn ViTs that are robust to target occlusion for UAV tracking. This framework is termed ORTrack. Additionally, to facilitate real-time applications, we propose an Adaptive Feature-Based Knowledge Distillation (AFKD) method to create a more compact tracker, which adaptively mimics the behavior of the teacher model ORTrack according to the task’s difficulty. This student model, dubbed ORTrack-D, retains much of ORTrack’s performance while offering higher efficiency. Extensive experiments on multiple benchmarks validate the effectiveness of our method, demonstrating its state-of-the-art performance. Codes is available at https://github.com/wuyou3474/ORTrack. You Wu 0009, Xucheng Wang, Xiangyang Yang 0001, Dan Zeng 0002, Hengzhou Ye, Shuiwang Li |
CVPR | 7 |
| 2025 | MambaNUT: Nighttime UAV Tracking via Mamba-based Adaptive Curriculum LearningabstractHarnessing low-light enhancement and domain adaptation, nighttime UAV tracking has made substantial strides. However, over-reliance on image enhancement, limited high-quality nighttime data, and a lack of integration between daytime and nighttime trackers hinder the development of an end-to-end trainable framework. Additionally, current ViT-based trackers demand heavy computational resources due to their reliance on the self-attention mechanism. In this paper, we propose a novel pure Mamba-based tracking framework (MambaNUT) that employs a state space model with linear complexity as its backbone, incorporating a single-stream architecture that integrates feature learning and template-search coupling within Vision Mamba. We introduce an adaptive curriculum learning (ACL) approach that dynamically adjusts sampling strategies and loss weights, thereby improving the model’s ability of generalization. Our ACL is composed of two levels of curriculum schedulers: (1) sampling scheduler that transforms the data distribution from imbalanced to balanced, as well as from easier (daytime) to harder (nighttime) samples; (2) loss scheduler that dynamically assigns weights based on the size of the training set and IoU of individual instances. Exhaustive experiments on multiple nighttime UAV tracking benchmarks demonstrate that the proposed MambaNUT achieves state-of-the-art performance while requiring lower computational costs. The code will be available at https://github.com/wuyou3474/MambaNUT. You Wu 0009, Xiangyang Yang 0001, Xucheng Wang, Hengzhou Ye, Dan Zeng 0002, Shuiwang Li |
IROS | 6 |
| 2025 | Camouflaged Object Tracking: A BenchmarkabstractVisual tracking has seen remarkable advancements, largely driven by the availability of large-scale training datasets that have enabled the development of highly accurate and robust algorithms. While significant progress has been made in tracking general objects, research on more challenging scenarios, such as tracking camouflaged objects, remains limited. Camouflaged objects, which blend seamlessly with their surroundings or other objects, present unique challenges for detection and tracking in complex environments. In critical fields like military, security, agriculture, and marine monitoring, accurately tracking camouflaged objects is essential. To address this gap, we introduce the Camouflaged Object Tracking Dataset (COTD), a specialized benchmark designed specifically for evaluating camouflaged object tracking methods. The COTD dataset comprises 200 sequences and approximately 80,000 frames, each annotated with detailed bounding boxes. Our evaluation of 20 existing tracking algorithms reveals significant deficiencies in their performance with camouflaged objects. To address these issues, we propose a novel tracking framework, HIPTrack-MLS, which demonstrates promising results in improving tracking performance for camouflaged objects. COTD and code are avialable at https://github.com/openat25/HIPTrack-MLS. Pengzhi Zhong, Defeng Huang, Huikai Shao, Qijun Zhao, Shuiwang Li |
ACM Multimedia | 7 |
| 2025 | Geometric self-supervision for monocular 3D animal pose estimation
Xiaowei Dai, Shuiwang Li, Qijun Zhao, Hongyu Yang 0002 |
Pattern Recognit. | 2 |
| 2025 | Adaptively bypassing vision transformer blocks for efficient visual tracking
Xiangyang Yang 0001, Dan Zeng 0002, Xucheng Wang, You Wu 0009, Hengzhou Ye, Qijun Zhao, Shuiwang Li |
Pattern Recognit. | 7 |
| 2025 | Exploiting rank-based filter pruning for real-time UAV tracking
Xucheng Wang, Dan Zeng 0002, Qijun Zhao, Shuiwang Li |
Signal Process. Image Commun. | 4 |
| 2024 | Tracking Reflected Objects: A Benchmark
Pengzhi Zhong, Lizhi Lin, Shuiwang Li |
ACCV (2) | 6 |
| 2024 | Learning Adaptive and View-Invariant Vision Transformer for Real-Time UAV TrackingabstractHarnessing transformer-based models, visual tracking has made substantial strides. However, the sluggish performance of current trackers limits their practicality on devices with constrained computational capabilities, especially for real-time unmanned aerial vehicle (UAV) tracking. Addressing this challenge, we introduce AVTrack, an adaptive computation framework tailored to selectively activate transformer blocks for real-time UAV tracking in this work. Our novel Activation Module (AM) dynamically optimizes ViT architecture, selectively engaging relevant components and enhancing inference efficiency without compromising much tracking performance. Moreover, we bolster the effectiveness of ViTs, particularly in addressing challenges arising from extreme changes in viewing angles commonly encountered in UAV tracking, by learning view-invariant representations through mutual information maximization. Extensive experiments on five tracking benchmarks affirm the effectiveness and versatility of our approach, positioning it as a state-of-the-art solution in visual tracking. Code is released at: https://github.com/wuyou3474/AVTrack. You Wu 0009, Xucheng Wang, Xiangyang Yang 0001, Shuiwang Li |
ICML | 6 |
| 2024 | Towards Labeling-free Fine-grained Animal Pose Estimation
Dan Zeng 0002, Shuiwang Li, Qijun Zhao, Qiaomu Shen, Bo Tang 0016 |
ACM Multimedia | 3 |
| 2024 | Tracking Transforming Objects: A Benchmark
You Wu 0009, Yuelong Wang, Yaxin Liao, Fuliang Wu, Hengzhou Ye, Shuiwang Li |
PRCV (13) | 6 |
| 2024 | Learning Target-Aware Vision Transformers for Real-Time UAV TrackingabstractIn recent years, the field of unmanned aerial vehicle (UAV) tracking has grown rapidly, finding numerous applications across various industries. While the discriminative correlation filters (DCF)-based trackers remain the most efficient and widely used in the UAV tracking, recently lightweight convolutional neural network (CNN)-based trackers using filter pruning have also demonstrated impressive efficiency and precision. However, the performance of these lightweight CNN-based trackers is still far from satisfactory. In the generic visual tracking, emerging vision transformer (ViT)-based trackers have shown great success by using cross-attention instead of correlation operation, enabling more effective capturing of relationships between the target and the search image. But to best of the authors’ knowledge, the UAV tracking community has not yet well explored the potential of ViTs for more effective and efficient template-search coupling for UAV tracking. In this article, we propose an efficient ViT-based tracking framework for real-time UAV tracking. Our framework integrates feature learning and template-search coupling into an efficient one-stream ViT to avoid an extra heavy relation modeling module. However, we observe that it tends to weaken the target information through transformer blocks due to the significantly more background tokens. To address this problem, we propose to maximize the mutual information (MI) between the template image and its feature representation produced by the ViT. The proposed method is dubbed TATrack. In addition, to further enhance efficiency, we introduce a novel MI maximization-based knowledge distillation, which strikes a better trade-off between accuracy and efficiency. Exhaustive experiments on five benchmarks show that the proposed tracker achieves state-of-the-art performance in UAV tracking. Code is released at:https://github.com/xyyang317/TATrack. Shuiwang Li, Xiangyang Yang 0001, Xucheng Wang, Dan Zeng 0002, Hengzhou Ye, Qijun Zhao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Unsupervised 3D Animal Canonical Pose Estimation with Geometric Self-SupervisionabstractAlthough analyzing animal shape and pose has potential applications in many fields, there is little work on 3D animal pose estimation. This can be attributed to two aspects: the lack of large-scale well-annotated datasets, and perspective ambiguities which make it difficult to map 2D space to 3D space. To address data scarcity, we propose an unsupervised method to estimate 3D animal pose, given only 2D poses. To deal with perspective ambiguities, we introduce a canonical consistency loss and a camera consistency loss to impose geometric priors in the training process, and combine the reprojection loss and the 2D pose discriminator to enable self-supervised learning. Specifically, given a 2D pose, the pose generator network generates a corresponding 3D pose and the camera network estimates a camera rotation. During training, the generated 3D pose is randomly reprojected onto camera viewpoints to synthesize a new 2D pose. The synthesized 2D pose is decomposed into a 3D pose and a camera rotation, based on which consistency losses are imposed in both 3D canonical poses and camera rotations for self-supervised training. We evaluate the proposed method on real and synthetic datasets, i.e., SMAL and AcinoSet. The experimental results demonstrate the effectiveness of the proposed method and we achieve state-of-the-art performance among unsupervised algorithms for 3D animal canonical pose estimation. Xiaowei Dai, Shuiwang Li, Qijun Zhao, Hongyu Yang 0002 |
FG | 2 |
| 2023 | Adaptive and Background-Aware Vision Transformer for Real-Time UAV TrackingabstractWhile discriminative correlation filters (DCF)-based trackers prevail in UAV tracking for their favorable efficiency, lightweight convolutional neural network (CNN)-based trackers using filter pruning have also demonstrated remarkable efficiency and precision. However, the use of pure vision transformer models (ViTs) for UAV tracking remains unexplored, which is a surprising finding given that ViTs have been shown to produce better performance and greater efficiency than CNNs in image classification. In this paper, we propose an efficient ViT-based tracking framework, Aba-ViTrack, for UAV tracking. In our framework, feature learning and template-search coupling are integrated into an efficient one-stream ViT to avoid an extra heavy relation modeling module. The proposed Aba-ViT exploits an adaptive and background-aware token computation method to reduce inference time. This approach adaptively discards tokens based on learned halting probabilities, which a priori are higher for background tokens than target ones. Extensive experiments on six UAV tracking benchmarks demonstrate that the proposed Aba-ViTrack achieves state-of-the-art performance in UAV tracking. Code is available at https://github.com/xyyang317/Aba-ViTrack. Shuiwang Li, Xiangxyang Yang, Dan Zeng 0002, Xucheng Wang |
ICCV | 1 |
| 2023 | Learning Disentangled Representation with Mutual Information Maximization for Real-Time UAV TrackingabstractEfficiency has been a critical problem in UAV tracking due to limitations in computation resources, battery capacity, and unmanned aerial vehicle maximum load. Although discriminative correlation filters (DCF)-based trackers prevail in this field for their favorable efficiency, some recently proposed lightweight deep learning (DL)-based trackers using model compression demonstrated quite remarkable CPU efficiency as well as precision. Unfortunately, the model compression methods utilized by these works, though simple, are still unable to achieve satisfying tracking precision with higher compression rates. This paper aims to exploit disentangled representation learning with mutual information maximization (DR-MIM) to further improve DL-based trackers’ precision and efficiency for UAV tracking. The proposed disentangled representation separates the feature into an identity-related and an identity-unrelated features. Only the latter is used, which enhances the effectiveness of the feature representation for subsequent classification and regression tasks. Extensive experiments on four UAV benchmarks, including UAV123@10fps, DTB70, UAVDT and VisDrone2018, show that our DR-MIM tracker significantly outperforms state-of-the-art UAV tracking methods. Xucheng Wang, Xiangyang Yang 0001, Hengzhou Ye, Shuiwang Li |
ICME | 4 |
| 2023 | Towards Discriminative Representations with Contrastive Instances for Real-Time UAV TrackingabstractMaintaining high efficiency and high precision are two fundamental challenges in UAV tracking due to the constraints of computing resources, battery capacity, and UAV maximum load. Discriminative correlation filters (DCF)-based trackers can yield high efficiency on a single CPU but with inferior precision. Lightweight Deep learning (DL)-based trackers can achieve a good balance between efficiency and precision but performance gains are limited by the compression rate. High compression rate often leads to poor discriminative representations. To this end, this paper aims to enhance the discriminative power of feature representations from a new feature-learning perspective. Specifically, we attempt to learn more disciminative representations with contrastive instances for UAV tracking in a simple yet effective manner, which not only requires no manual annotations but also allows for developing and deploying a lightweight model. We are the first to explore contrastive learning for UAV tracking. Extensive experiments on four UAV benchmarks, including UAV123@10fps, DTB70, UAVDT and VisDrone2018, show that the proposed DRCI tracker significantly outperforms state-of-the-art UAV tracking methods. Dan Zeng 0002, Mingliang Zou, Xucheng Wang, Shuiwang Li |
ICME | 4 |
| 2023 | Data-Scarce Animal Face Alignment via Bi-Directional Cross-Species Knowledge TransferabstractAnimal face alignment is challenging due to large intra- and inter-species variations and a scarcity of labeled data. Existing studies circumvent this problem by directly finetuning a human face alignment model or focusing on animal-specific face alignment~(e.g., horse, sheep). In this paper, we propose Cross-Species Knowledge Transfer, Meta-CSKT, for animal face alignment, which consists of a base network and an adaptation network. Two networks continuously complement each other through the bi-directional cross-species knowledge transfer. This is motivated by observing knowledge sharing among animals. Meta-CSKT uses a circuit feedback mechanism to improve the base network with the cognitive differences of the adaptation network between few-shot labeled and large-scale unlabeled data. In addition, we propose a positive example mining method to identify positives, semi-hard positives, and hard negatives in unlabeled data to mitigate the scarcity of labeled data and facilitate Meta-CSKT learning. Experiments show that Meta-CSKT outperforms state-of-the-art methods by a large margin on the horse facial keypoint dataset and Japanese Macaque Species dataset, while achieving comparable results to state-of-the-art methods on large-scale labeled AnimalWeb~(e.g., 18K), using only a few labeled images~(e.g., 40)1. Dan Zeng 0002, Shanchuan Hong, Shuiwang Li, Qiaomu Shen, Bo Tang 0016 |
ACM Multimedia | 3 |
| 2022 | Tracking Small and Fast Moving Objects: A Benchmark
Fuliang Wu, Yuming Qiu, Jingdong Liang, Shuiwang Li |
ACCV (7) | 5 |
| 2022 | Learning Disentangled Representation in Pruning for Real-Time UAV Tracking
Siyu Ma, Yuting Liu 0004, Dan Zeng 0002, Yaxin Liao, Shuiwang Li |
ACML | 6 |
| 2022 | Animal Pose Refinement in 2D Images with 3D Constraints
Xiaowei Dai, Shuiwang Li, Qijun Zhao, Hongyu Yang 0002 |
BMVC | 2 |
| 2022 | Global Filter Pruning with Self-Attention for Real-Time UAV Tracking
Yuelong Wang, Qiangyu Sun, Shuiwang Li |
BMVC | 4 |
| 2022 | Rank-Based Filter Pruning for Real-Time UAV TrackingabstractUnmanned aerial vehicle (UAV) tracking has wide poten-tial applications in such as agriculture, navigation, and public security. However, the limitations of computing resources, battery capacity, and maximum load of UAV hinder the de-ployment of deep learning-based tracking algorithms on UAV. Consequently, discriminative correlation filters (DCF) track-ers stand out in the UAV tracking community because of their high efficiency. However, their precision is usually much lower than trackers based on deep learning. Model compression is a promising way to narrow the gap (i.e., effciency, precision) between DCF- and deep learning- based trackers, which has not caught much attention in UAV tracking. In this paper, we propose the P-SiamFC++ tracker, which is the first to use rank-based filter pruning to compress the SiamFC++ model, achieving a remarkable balance between efficiency and precision. Our method is general and may encourage further studies on UAV tracking with model compression. Extensive experiments on four UAV benchmarks, including UAV123@10fps, DTB70, UAVDT and Vistrone2018, show that P-SiamFC++ tracker significantly outperforms state-of-the-art UAV tracking methods. Xucheng Wang, Dan Zeng 0002, Qijun Zhao, Shuiwang Li |
ICME | 4 |
| 2022 | Fisher Pruning for Real-Time UAV TrackingabstractDespite the wide prospect of applications of unmanned aerial vehicle (UAV)-based tracking in transportation, agriculture, public security, and so on, the limitations of computing resources, battery capacity and maximum load of UAV largely hinder the deployment of deep learning (DL)-based tracking algorithms on UAV. Discriminative correlation filters (DCF)-based trackers because of their high efficiency and low resource-consuming merits have thus stood out in the UAV tracking community. However, the precision of DCF-based trackers is hardly comparable to DL-based ones in complex scenarios due to the limited representation learning ability. Filter pruning is a common technique used to deploy deep networks in edge-devices with low power and constrained resources, without compromising much on the accuracy of the model. It is probably an effective means to improve the efficiency of DL-based trackers and facilitate their deployment on UAVs. However, applying filter pruning to UAV tracking has not been well explored. A simple and effective pruning criterion is very desirable at present and may draw more attention in the UAV tracking community to model compression. In this paper, we propose to exploit Fisher pruning to compress the SiamFC++ model for UAV tracking, resulting in our proposed F-SiamFC++ tracker which demonstrates a remarkable balance between efficiency and precision. Extensive experiments on four UAV benchmarks, including UAV123@10fps, DTB70, UAVDT and Vistrone2018 (VisDrone2018-test-dev), show that the proposed F-SiamFC++ tracker achieves state-of-the-art performance. Wanying Wu, Pengzhi Zhong, Shuiwang Li |
IJCNN | 3 |
| 2022 | Detection Beyond What and Where: A Benchmark for Detecting Occlusion State
Liwei Qin, Zhongtian Wang, Yuanyuan Liao, Shuiwang Li |
PRCV (4) | 6 |
| 2022 | SSPNet: Scale Selection Pyramid Network for Tiny Person Detection From UAV ImagesabstractWith the increasing demand for search and rescue, it is highly demanded to detect objects of interest in large-scale images captured by unmanned aerial vehicles (UAVs), which is quite challenging due to extremely small scales of objects. Most existing methods employed a feature pyramid network (FPN) to enrich shallow layers’ features by combining deep layers’ contextual features. However, under the limitation of the inconsistency in gradient computation across different layers, the shallow layers in FPN are not fully exploited to detect tiny objects. In this article, we propose a scale selection pyramid network (SSPNet) for tiny person detection, which consists of three components: context attention module (CAM), scale enhancement module (SEM), and scale selection module (SSM). CAM takes account of context information to produce hierarchical attention heatmaps. SEM highlights features of specific scales at different layers, leading the detector to focus on objects of specific scales instead of vast backgrounds. SSM exploits adjacent layers’ relationships to fulfill suitable feature sharing between deep layers and shallow layers, thereby avoiding the inconsistency in gradient computation across different layers. Besides, we propose a weighted negative sampling (WNS) strategy to guide the detector to select more representative samples. Experiments on the TinyPerson benchmark show that our method outperforms other state-of-the-art (SOTA) detectors. Mingbo Hong, Shuiwang Li, Feiyu Zhu 0001, Qijun Zhao |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Learning residue-aware correlation filters and refining scale for real-time UAV tracking
Shuiwang Li, Yuting Liu 0004, Qijun Zhao, Ziliang Feng |
Pattern Recognit. | 1 |
| 2021 | Learning Residue-Aware Correlation Filters and Refining Scale Estimates with the GrabCut for Real-Time UAV TrackingabstractUnmanned aerial vehicle (UAV)-based tracking is attracting increasing attention and developing rapidly in applications such as agriculture, aviation, navigation, transportation and public security. Recently, discriminative correlation filters (DCF)-based trackers have stood out in UAV tracking community for their high efficiency and appealing robustness on a single CPU. However, due to limited onboard computation resources and other challenges the efficiency and accuracy of existing DCF-based approaches is still not satisfying. In this paper, inspired by residue representation, we exploit the residue nature inherent to videos and propose residue-aware correlation filters that show better convergence properties in filter learning. Moreover, we explore using segmentation by the GrabCut to improve the wildly adopted discriminative scale estimation in DCF-based trackers, which, as a mater of fact, greatly impacts the precision and accuracy of the trackers since accumulated scale error degrades the appearance model as online updating goes on. Extensive experiments are conducted on four UAV benchmarks, namely, UAV123@10fps, DTB70, UAVDT and Vistrone2018 (VisDrone2018-test-dev). The results show that our method achieves state-of-the-art performance. Shuiwang Li, Yuting Liu 0004, Qijun Zhao, Ziliang Feng |
3DV | 1 |
| 2021 | Learning Disentangled Representation for Fine-Grained Visual Categorization
Wenjie Dang, Shuiwang Li, Qijun Zhao |
ICIG (1) | 2 |
| 2021 | Equivalence of Correlation Filter and Convolution Filter in Visual Tracking
Shuiwang Li, Qijun Zhao, Ziliang Feng |
ICIG (3) | 1 |
| 2021 | Towards Silhouette-Aware Human Detection in Depth ImagesabstractDetecting humans in depth images attracts increasing attention thanks to the advantage of depth modality in privacy protection. However, this task is challenging because the number of available training data is limited to date and the depth images, unlike RGB images, are short of rich texture features. Although many image synthesis methods and deep learning methods have been proposed and proven successful, especially for RGB images, it is unsatisfactory to directly apply them to depth images because of the intrinsical differences between the modalities. In view of that silhouette of humans becomes an essential discriminative cue in depth images in the absence of texture information, which is not well utilized by existing methods, in this paper, we thus propose a silhouette-aware network (SAN) to train the detection model and a depth image synthesis method that represses spurious silhouette to augment the training data. Besides, to further increase the diversity of training data, we collect a dataset of scene depth images (SDI), including both indoor and outdoor scenes, as background images when synthesizing training data. Experimental results show that (i) our proposed synthesis method can generate more realistic depth images and thus benefits the training of detection models, (ii) our collected SDI dataset can effectively enhance data diversity and thus improves the effectiveness of the obtained detection models, and (iii) our proposed silhouette-aware network (SAN) can effectively boost the human detection accuracy. Our dataset is available at https://pan.baidu.com/s/13hpuziavBNjS8KATClpXww, password: r9id Shuiwang Li, Qijun Zhao |
IJCNN | 2 |
| 2020 | Asymmetric discriminative correlation filters for visual trackingabstractDiscriminative correlation filters (DCF) are efficient in visual tracking and have advanced the field significantly. However, the symmetry of correlation (or convolution) operator results in computational problems and does harm to the generalized translation equivariance. The former problem has been approached in many ways, whereas the latter one has not been well recognized. In this paper, we analyze the problems with the symmetry of circular convolution and propose an asymmetric one, which as a generalization of the former has a weak generalized translation equivariance property. With this operator, we propose a tracker called the asymmetric discriminative correlation filter (ADCF), which is more sensitive to translations of targets. Its asymmetry allows the filter and the samples to have different sizes. This flexibility makes the computational complexity of ADCF more controllable in the sense that the number of filter parameters will not grow with the sample size. Moreover, the normal matrix of ADCF is a block matrix with each block being a two-level block Toeplitz matrix. With this well-structured normal matrix, we design an algorithm for multiplying an N × N two-level block Toeplitz matrix by a vector with time complexity O ( N log N ) and space complexity O ( N ), instead of O ( N 2 ). Unlike DCF-based trackers, introducing spatial or temporal regularization does not increase the essential computational complexity of ADCF. Comparative experiments are performed on a synthetic dataset and four benchmarks, including OTB-2013, OTB-2015, VOT-2016, and Temple-Color, and the results show that our method achieves state-of-the-art visual tracking performance. Shuiwang Li, Qianbo Jiang, Qijun Zhao, Ziliang Feng |
Frontiers Inf. Technol. Electron. Eng. | 1 |