Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Qiang Wang 0051

dblp:64/5630-51 · DBLP profile ↗
← Back
21ranked-venue papers
4as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
12 papers
Video understanding and tracking · 86% Deep learning architectures and training · 11% Representation and self-supervised learning · 3%
Computer networks
1 paper
Vehicular, aerial and satellite networks · 100%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
object tracking
4.6112023
Anti-UAV: A Large-Scale Benchmark for Vision-Based UAV Tracking · IEEE Trans. Multim. 2023
SiamMask: A Framework for Fast Online Object Tracking and Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Deep Spatial and Temporal Network for Robust Visual Object Tracking · IEEE Trans. Image Process. 2020
Computer vision › Video understanding and tracking
video object segmentation
1.432023
SiamMask: A Framework for Fast Online Object Tracking and Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Anchor Diffusion for Unsupervised Video Object Segmentation · ICCV 2019
Fast Online Object Tracking and Segmentation: A Unifying Approach · CVPR 2019
Computer vision › Video understanding and tracking › object tracking › deep tracking
siamese tracking
1.022023
SiamMask: A Framework for Fast Online Object Tracking and Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2023
SiamRPN++: Evolution of Siamese Visual Tracking With Very Deep Networks · CVPR 2019
Computer vision › Video understanding and tracking › object tracking
UAV tracking
0.712023
Anti-UAV: A Large-Scale Benchmark for Vision-Based UAV Tracking · IEEE Trans. Multim. 2023
Machine learning › Deep learning architectures and training
siamese network
0.422018
Learning Attentions: Residual Attentional Siamese Network for High Performance Online Visual Tracking · CVPR 2018
Distractor-Aware Siamese Networks for Visual Object Tracking · ECCV (9) 2018
Machine learning › Deep learning architectures and training › sequence modeling
long-range dependency modeling
0.412019
Anchor Diffusion for Unsupervised Video Object Segmentation · ICCV 2019
Computer vision › Video understanding and tracking › video object segmentation
semi-supervised video object segmentation
0.412019
Fast Online Object Tracking and Segmentation: A Unifying Approach · CVPR 2019
Computer vision › Video understanding and tracking
temporal alignment
0.412019
Anchor Diffusion for Unsupervised Video Object Segmentation · ICCV 2019
Computer vision › Video understanding and tracking › video object segmentation
unsupervised video object segmentation
0.412019
Anchor Diffusion for Unsupervised Video Object Segmentation · ICCV 2019
Computer vision › Video understanding and tracking › object tracking › discriminative tracking
correlation filter tracking
0.312018
Visual Tracking via Spatially Aligned Correlation Filters Network · ECCV (3) 2018
Machine learning › Representation and self-supervised learning › representation learning › joint representation learning
multi-task representation learning
0.312018
Do not Lose the Details: Reinforced Representation Learning for High Performance Visual Tracking · IJCAI 2018
Computer vision › Video understanding and tracking
multi-object tracking
0.212023
SiamMask: A Framework for Fast Online Object Tracking and Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Vehicular, aerial and satellite networks
unmanned aerial vehicles
0.212023
Anti-UAV: A Large-Scale Benchmark for Vision-Based UAV Tracking · IEEE Trans. Multim. 2023
Image and video processing › image filtering
correlation filters
0.112018
Do not Lose the Details: Reinforced Representation Learning for High Performance Visual Tracking · IJCAI 2018

Methods — techniques the papers use, named apart from their topics

correlation filter · 1.4semantic flow · 1.3multi-task learning · 1.3dual-flow semantic consistency · 1.3binary segmentation loss · 1.0deep learning · 0.8fully-convolutional siamese network · 0.7particle filter · 0.4offline training · 0.4gaussian process regression · 0.4encoder-decoder network · 0.3convolutional neural network · 0.3
YearPublicationVenuePosition
2025 CMRFusion: Efficient Feature Decomposition for RGB-T Fusion via Cross Modality Mask Reconstruction
abstract
The objective of infrared-visible image fusion is to effectively combine the complementary information from both modalities and preserve the significant features within their respective imaging domains. In this paper, we propose a novel fusion method for infrared-visible images named CMRFusion, which employs mask reconstruction to disentangle cross-modality features. To be specific, we firstly transform the infrared and visual images to the modality-common (the semantic information of the scene) and modality-unique (the intensity attributes of the imaging sensors) feature subspaces separately by reconstructing the masked patches from those two modalities under meticulously designed constraints. A simple yet effective generator is then trained to integrate the above decomposed features into the final fused image. Extensive quantitative experiments on the public TNO and Roadscene datasets demonstrate that CMRFusion mostly outperforms the existing state-of-the-art methods. The visualization results also prove the effectiveness of CMRFusion in terms of decomposing the shared and distinctive features present within the two modalities.
Qiang Wang 0051, Zhenyu He 0001
ICME4
2025 Universal Federated Domain Adaptation Through One-vs-All Self-Supervision for Internet of Things
abstract
In practical Internet of Things (IoT) applications, deep neural networks (DNNs) often encounter challenges arising from covariate shifts (differences in feature distributions) and category shifts (discrepancies in label spaces), which significantly degrade their generalization performance. To mitigate these issues, universal federated domain adaptation (UFDA) techniques have been proposed to train a global model that can classify known and unknown categories while keeping data private. Nevertheless, most existing methods still struggle to precisely identify samples belonging to unknown classes in the target domain due to the unavailability of data from the source domain clients. To address these challenges, we propose a novel method, termed one-vs-all self-supervision (OSS) for IoT scenario. Specifically, OSS mainly consists of following three components. First, one-vs-all pseudo-label generation is proposed to generate high-quality pseudo-labels by leveraging source client models. Subsequently, we design a category-diverse strategy to aggregate the source models by assigning appropriate weights to each source domain client. Finally, we implement a target self-supervised learning strategy to refine feature alignment with respect to cluster centers. Comprehensive experiments are performed on four benchmark datasets: Office-31, Office-Home, VisDA-2017+ImageCLEF-DA, and Digits. The results show that our proposed OSS method achieves state-of-the-art performance in UFDA, significantly enhancing the recognition accuracy.
Haojin Liao, Qiang Wang 0051, Sicheng Zhao, Tengfei Xing, Runbo Hu
IEEE Internet Things J.2
2024 DCFNet: Discriminant Correlation Filters Network for Visual Tracking
Weiming Hu 0004, Qiang Wang 0051, Bing Li 0001, Stephen J. Maybank
J. Comput. Sci. Technol.2
2023 Domain consensual contrastive learning for few-shot universal domain adaptation
Haojin Liao, Qiang Wang 0051, Sicheng Zhao, Tengfei Xing, Runbo Hu
Appl. Intell.2
2023 SiamMask: A Framework for Fast Online Object Tracking and Segmentation
abstract
In this article, we introduce SiamMask, a framework to perform both visual object tracking and video object segmentation, in real-time, with the same simple method. We improve the offline training procedure of popular fully-convolutional Siamese approaches by augmenting their losses with a binary segmentation task. Once the offline training is completed, SiamMask only requires a single bounding box for initialization and can simultaneously carry out visual object tracking and segmentation at high frame-rates. Moreover, we show that it is possible to extend the framework to handle multiple object tracking and segmentation by simply re-using the multi-task model in a cascaded fashion. Experimental results show that our approach has high processing efficiency, at around 55 frames per second. It yields real-time state-of-the art results on visual-object tracking benchmarks, while at the same time demonstrating competitive performance at a high speed for video object segmentation benchmarks.
Weiming Hu 0004, Qiang Wang 0051, Li Zhang 0040, Luca Bertinetto, Philip Torr 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Anti-UAV: A Large-Scale Benchmark for Vision-Based UAV Tracking
abstract
Unmanned Aerial Vehicles (UAV) have many applications in both commerce and recreation. However, irresponsibly operated UAVs will pose a threat to public safety. Therefore, developing our understanding of UAVs and their uses is of particular interest. This paper considers tracking UAVs, which provide multifaceted information around location, paths and trajectories. To facilitate research on this topic, we introduce a new benchmark, herein referred to as Anti-UAV, which provides a novel direction for UAV tracking with more than 300 video pairs containing over 580 k manually annotated bounding boxes. Addressing anti-UAV research challenges could help to design anti-UAV systems, which in turn may improve surveillance. Accordingly, we have proposed a simple yet effective approach, called dual-flow semantic consistency (DFSC) is proposed for UAV tracking. Modulated by the semantic flow across video sequences, tracker learns more robust class-level semantic information and obtains more discriminative instance-level features. Experiments highlight significant performance gain with the proposed approach over state-of-the-art trackers and the challenging aspects of Anti-UAV. The Anti-UAV benchmark and the code for the proposed approach have been made publicly available athttps://github.com/ucas-vg/Anti-UAVandhttps://github.com/ZhaoJ9014/Anti-UAV.
Kuiran Wang, Xiaoke Peng, Xuehui Yu, Qiang Wang 0051, Junliang Xing, Guorong Li, Guodong Guo, Qixiang Ye, Jianbin Jiao, Jian Zhao 0006, Zhenjun Han
IEEE Trans. Multim.5
2020 End-to-End Temporal Feature Aggregation for Siamese Trackers
abstract
While siamese networks have demonstrated the significant improvement on object tracking performances, how to utilize the temporal information in siamese trackers has not been widely studied yet. In this paper, we introduce a novel siamese tracking architecture equipped with a temporal aggregation module, which improves the per-frame features by aggregating temporal information from adjacent frames. This temporal fusion strategy enables the siamese trackers to handle poor object appearance like motion blur, occlusion, etc. Furthermore, we incorporate the adversarial dropout module in the siamese network for computing discriminative target features in an end-to-end-fashion. Comprehensive experiments demonstrate that the proposed tracker performs favorably against state-of-the-art trackers.
Zhenbang Li, Qiang Wang 0051, Bing Li 0001, Weiming Hu 0004
ICIP2
2020 Globally Spatial-Temporal Perception: a Long-Term Tracking System
abstract
Although siamese trackers have achieved superior performance, these kinds of approaches tend to favour the local search mechanism and are thus prone to accumulating inaccuracies of predicted positions, leading to tracking drift over time, especially in long-term tracking scenario. To solve these problems, we propose a siamese tracker in the spirit of the faster RCNN's two-stage detection paradigm. This new tracker is dedicated to reducing cumulative inaccuracies and improving robustness based on a global perception mechanism, which allows the target to be retrieved in time spatially over the whole image plane. Since the very deep network can be enabled for feature learning in this two-stage tracking framework, the power of discrimination is guaranteed. What's more, we also add a CNN-based trajectory prediction module exploiting the target's temporal motion information to mitigate the interference of distractors. These two spatial and temporal modules exploit both the high-level appearance information and complementary trajectory information to improve the tracking robustness. Comprehensive experiments demonstrate that the proposed Globally Spatial-Temporal Perception-based tracking system performs favorably against state-of-the-art trackers.
Zhenbang Li, Qiang Wang 0051, Bing Li 0001, Weiming Hu 0004
ICIP2
2020 Tracking-by-Fusion via Gaussian Process Regression Extended to Transfer Learning
abstract
This paper presents a new Gaussian Processes (GPs)-based particle filter tracking framework. The framework non-trivially extends Gaussian process regression (GPR) to transfer learning, and, following the tracking-by-fusion strategy, integrates closely two tracking components, namely a GPs component and a CFs one. First, the GPs component analyzes and models the probability distribution of the object appearance by exploiting GPs. It categorizes the labeled samples into auxiliary and target ones, and explores unlabeled samples in transfer learning. The GPs component thus captures rich appearance information over object samples across time. On the other hand, to sample an initial particle set in regions of high likelihood through the direct simulation method in particle filtering, the powerful yet efficient correlation filters (CFs) are integrated, leading to the CFs component. In fact, the CFs component not only boosts the sampling quality, but also benefits from the GPs component, which provides re-weighted knowledge as latent variables for determining the impact of each correlation filter template from the auxiliary samples. In this way, the transfer learning based fusion enables effective interactions between the two components. Superior performance on four object tracking benchmarks (OTB-2015, Temple-Color, and VOT2015/2016), and in comparison with baselines and recent state-of-the-art trackers, has demonstrated clearly the effectiveness of the proposed framework.
Qiang Wang 0051, Junliang Xing, Haibin Ling, Weiming Hu 0004, Stephen J. Maybank
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 Distractor-aware discrimination learning for online multiple object tracking
Zongwei Zhou, Wenhan Luo, Qiang Wang 0051, Junliang Xing, Weiming Hu 0004
Pattern Recognit.3
2020 Deep Spatial and Temporal Network for Robust Visual Object Tracking
abstract
There are two key components that can be leveraged for visual tracking: (a) object appearances; and (b) object motions. Many existing techniques have recently employed deep learning to enhance visual tracking due to its superior representation power and strong learning ability, where most of them employed object appearances but few of them exploited object motions. In this work, a deep spatial and temporal network (DSTN) is developed for visual tracking by explicitly exploiting both the object representations from each frame and their dynamics along multiple frames in a video, such that it can seamlessly integrate the object appearances with their motions to produce compact object appearances and capture their temporal variations effectively. Our DSTN method, which is deployed into a tracking pipeline in a coarse-to-fine form, can perceive the subtle differences on spatial and temporal variations of the target (object being tracked), and thus it benefits from both off-line training and online fine-tuning. We have also conducted our experiments over four largest tracking benchmarks, including OTB-2013, OTB-2015, VOT2015, and VOT2017, and our experimental results have demonstrated that our DSTN method can achieve competitive performance as compared with the state-of-the-art techniques. The source code, trained models, and all the experimental results of this work will be made public available to facilitate further studies on this problem.
Zhu Teng, Junliang Xing, Qiang Wang 0051, Baopeng Zhang, Jianping Fan 0001
IEEE Trans. Image Process.3
2019 SiamRPN++: Evolution of Siamese Visual Tracking With Very Deep Networks
abstract
Siamese network based trackers formulate tracking as convolutional feature cross-correlation between target template and searching region. However, Siamese trackers still have accuracy gap compared with state-of-the-art algorithms and they cannot take advantage of feature from deep networks, such as ResNet-50 or deeper. In this work we prove the core reason comes from the lack of strict translation invariance. By comprehensive theoretical analysis and experimental validations, we break this restriction through a simple yet effective spatial aware sampling strategy and successfully train a ResNet-driven Siamese tracker with significant performance gain. Moreover, we propose a new model architecture to perform depth-wise and layer-wise aggregations, which not only further improves the accuracy but also reduces the model size. We conduct extensive ablation studies to demonstrate the effectiveness of the proposed tracker, which obtains currently the best results on four large tracking benchmarks, including OTB2015, VOT2018, UAV123, and LaSOT. Our model will be released to facilitate further studies based on this problem.
Bo Li 0114, Wei Wu 0021, Qiang Wang 0051, Fangyi Zhang, Junliang Xing
CVPR3
2019 Fast Online Object Tracking and Segmentation: A Unifying Approach
abstract
In this paper we illustrate how to perform both visual object tracking and semi-supervised video object segmentation, in real-time, with a single simple approach. Our method, dubbed SiamMask, improves the offline training procedure of popular fully-convolutional Siamese approaches for object tracking by augmenting their loss with a binary segmentation task. Once trained, SiamMask solely relies on a single bounding box initialisation and operates online, producing class-agnostic object segmentation masks and rotated bounding boxes at 55 frames per second. Despite its simplicity, versatility and fast speed, our strategy allows us to establish a new state-of-the-art among real-time trackers on VOT-2018, while at the same time demonstrating competitive performance and the best speed for the semi-supervised video object segmentation task on DAVIS-2016 and DAVIS-2017.
Qiang Wang 0051, Li Zhang 0040, Luca Bertinetto, Weiming Hu 0004, Philip Torr 0001
CVPR1
2019 Anchor Diffusion for Unsupervised Video Object Segmentation
abstract
Unsupervised video object segmentation has often been tackled by methods based on recurrent neural networks and optical flow. Despite their complexity, these kinds of approach tend to favour short-term temporal dependencies and are thus prone to accumulating inaccuracies, which cause drift over time. Moreover, simple (static) image segmentation models, alone, can perform competitively against these methods, which further suggests that the way temporal dependencies are modelled should be reconsidered. Motivated by these observations, in this paper we explore simple yet effective strategies to model long-term temporal dependencies. Inspired by the non-local operators, we introduce a technique to establish dense correspondences between pixel embeddings of a reference "anchor" frame and the current one. This allows the learning of pairwise dependencies at arbitrarily long distances without conditioning on intermediate frames. Without online supervision, our approach can suppress the background and precisely segment the foreground object even in challenging scenarios, while maintaining consistent performance over time. With a mean IoU of 81.7%, our method ranks first on the DAVIS-2016 leaderboard of unsupervised methods, while still being competitive against state-of-the-art online semi-supervised approaches. We further evaluate our method on the FBMS dataset and the video saliency dataset ViSal, showing results competitive with the state of the art.
Zhao Yang 0002, Qiang Wang 0051, Luca Bertinetto, Song Bai 0001, Weiming Hu 0004, Philip Torr 0001
ICCV2
2018 Learning Attentions: Residual Attentional Siamese Network for High Performance Online Visual Tracking
abstract
Offline training for object tracking has recently shown great potentials in balancing tracking accuracy and speed. However, it is still difficult to adapt an offline trained model to a target tracked online. This work presents a Residual Attentional Siamese Network (RASNet) for high performance object tracking. The RASNet model reformulates the correlation filter within a Siamese tracking framework, and introduces different kinds of the attention mechanisms to adapt the model without updating the model online. In particular, by exploiting the offline trained general attention, the target adapted residual attention, and the channel favored feature attention, the RASNet not only mitigates the over-fitting problem in deep network training, but also enhances its discriminative capacity and adaptability due to the separation of representation learning and discriminator learning. The proposed deep architecture is trained from end to end and takes full advantage of the rich spatial temporal information to achieve robust visual tracking. Experimental results on two latest benchmarks, OTB-2015 and VOT2017, show that the RASNet tracker has the state-of-the-art tracking accuracy while runs at more than 80 frames per second.
Qiang Wang 0051, Zhu Teng, Junliang Xing, Weiming Hu 0004, Stephen J. Maybank
CVPR1
2018 Visual Tracking via Spatially Aligned Correlation Filters Network
Mengdan Zhang, Qiang Wang 0051, Junliang Xing, Peixi Peng, Weiming Hu 0004, Stephen J. Maybank
ECCV (3)2
2018 Distractor-Aware Siamese Networks for Visual Object Tracking
Qiang Wang 0051, Bo Li 0114, Wei Wu 0021, Weiming Hu 0004
ECCV (9)2
2018 SPCNet: Scale Position Correlation Network for End-to-End Visual Tracking
abstract
We present a novel Scale Position Correlation Network (SPCNet) for learning to track objects robustly and efficiently. Different from most previous Correlation Filter (CF) based tracking models, SPCNet unifies the feature representation learning and CF based appearance modeling within one end-to-end learnable framework. In particular, SPCNet learns to track objects within a joint scale-position space, and is very effective in learning features for the accurate prediction of object scale and position. To learn our model from end to end, the SPCNet introduces a differentiable correlation filter layer into a Siamese architecture. Therefore, the localization error can be effectively back-propagated through the whole network, enabling fast adaptation of feature learning and appearance modeling for the objects to be tracked. Such task driven feature learning admits a very lightweight design that can be efficiently pre-trained. In addition, the dense appearance modeling in the joint scale-position space is also efficient. It benefits from the computation of gradients within the Fourier frequency domain. Such careful architecture design ensures that SPCNet is effective and efficient with a small model size. Extensive experimental analyses and evaluations on three largest benchmarks, OTB-2013, OTB-2015, and VOT2015, demonstrate its superiority over many state-of-the-art algorithms.
Qiang Wang 0051, Mengdan Zhang, Junliang Xing, Weiming Hu 0004
ICPR1
2018 Do not Lose the Details: Reinforced Representation Learning for High Performance Visual Tracking
abstract
This work presents a novel end-to-end trainable CNN model for high performance visual object tracking. It learns both low-level fine-grained representations and a high-level semantic embedding space in a mutual reinforced way, and a multi-task learning strategy is proposed to perform the correlation analysis on representations from both levels. In particular, a fully convolutional encoder-decoder network is designed to reconstruct the original visual features from the semantic projections to preserve all the geometric information. Moreover, the correlation filter layer working on the fine-grained representations leverages a global context constraint for accurate object appearance modeling. The correlation filter in this layer is updated online efficiently without network fine-tuning. Therefore, the proposed tracker benefits from two complementary effects: the adaptability of the fine-grained correlation analysis and the generalization capability of the semantic embedding. Extensive experimental evaluations on four popular benchmarks demonstrate its state-of-the-art performance.
Qiang Wang 0051, Mengdan Zhang, Junliang Xing, Weiming Hu 0004, Stephen J. Maybank
IJCAI1
2017 Robust Object Tracking Based on Temporal and Spatial Deep Networks
abstract
Recently deep neural networks have been widely employed to deal with the visual tracking problem. In this work, we present a new deep architecture which incorporates the temporal and spatial information to boost the tracking performance. Our deep architecture contains three networks, a Feature Net, a Temporal Net, and a Spatial Net. The Feature Net extracts general feature representations of the target. With these feature representations, the Temporal Net encodes the trajectory of the target and directly learns temporal correspondences to estimate the object state from a global perspective. Based on the learning results of the Temporal Net, the Spatial Net further refines the object tracking state using local spatial object information. Extensive experiments on four of the largest tracking benchmarks, including VOT2014, VOT2016, OTB50, and OTB100, demonstrate competing performance of the proposed tracker over a number of state-of-the-art algorithms.
Zhu Teng, Junliang Xing, Qiang Wang 0051, Congyan Lang, Songhe Feng, Yi Jin 0001
ICCV3
2008 Avatar motion control by natural body movement via camera
Chun Chen 0001, Qiang Wang 0051, Mingli Song, Dacheng Tao, Xuelong Li 0001
Neurocomputing3