Qiangqiang Wu

dblp:193/1415 · DBLP profile ↗
← Back
26ranked-venue papers
8as first author
22since 2021 · last 2026
0000-0002-3847-7838ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 8 first-author · 19 since 2021Artificial intelligence and machine learning · 14 · 6 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 SAM2-OV: A Novel Detection-Only Tuning Paradigm for Open-Vocabulary Multi-Object Tracking
abstract
Open-vocabulary multi-object tracking (OV-MOT) aims to track objects with unseen categories beyond the training set. While existing methods rely on pseudo video sequences synthesized from static images, they struggle to model realistic motion patterns, resulting in limited association performance in real-world scenarios. To alleviate these issues, we propose SAM2-OV, a novel association learning-free OV-MOT method that adopts a detection-only tuning paradigm, eliminating the need for synthetic sequences or spatiotemporal supervision and substantially reducing the overall learnable parameters. The core of our method is a Unified Detection Module (UDM), which effectively provides object-level prompts to enable SAM2 for OV-MOT. Enabled by UDM, SAM2-OV is the first to integrate SAM2 for OV-MOT, fully unleashing its zero-shot cross-frame association ability. To further enhance object association under occlusion and abrupt motion, we introduce a Motion Prior Assistance Module (MPAM) that incorporates motion cues into the mask selection process. In addition, a Semantic Enhancement Adapter (SEA) distilled from CLIP is used to improve classification generalization. A sparse prompting strategy is also adopted to reduce computational redundancy by triggering detection only on selected keyframes. As only the detection module is tuned on static images, the overall training process remains simple and efficient. Experiments on the TAO dataset demonstrate that SAM2-OV achieves state-of-the-art performance under the TETA metric, particularly on novel categories. Evaluations on the KITTI dataset show the strong zero-shot cross-domain transferability of our SAM2-OV.
Yangkai Chen, Qiangqiang Wu, Junlong Gao, Guanglin Niu, Hanzi Wang
AAAI2
2026 SAVER: Mitigating Hallucinations in Large Vision-Language Models via Style-Aware Visual Early Revision
abstract
Large Vision-Language Models (LVLMs) recently achieve significant breakthroughs in understanding complex visual-textual contexts. However, hallucination issues still limit their real-world applicability. Although previous mitigation methods effectively reduce hallucinations in photographic images, they largely overlook the potential risks posed by stylized images, which play crucial roles in critical scenarios such as game scene understanding, art education, and medical analysis. In this work, we first construct a dataset comprising photographic images and their corresponding stylized versions with carefully annotated caption labels. We then conduct head-to-head comparisons on both discriminative and generative tasks by benchmarking 13 advanced LVLMs on the collected datasets. Our findings reveal that stylized images tend to induce significantly more hallucinations than their photographic counterparts. To address this issue, we propose Style-Aware Visual Early Revision (SAVER), a novel mechanism that dynamically adjusts LVLMs' final outputs based on the token-level visual attention patterns, leveraging early-layer feedback to mitigate hallucinations caused by stylized images. Extensive experiments demonstrate that SAVER achieves state-of-the-art performance in hallucination mitigation across various models, datasets, and tasks.
Zhaoxu Li, Chenqi Kong, Yi Yu 0011, Qiangqiang Wu, Xinghao Jiang, Ngai-Man Cheung, Bihan Wen, Alex Chichung Kot, Xudong Jiang 0001
AAAI4
2026 DPENet: A Dual Prototype-Enhanced Network for Few-Shot Object Detection
abstract
Existing meta-learning based few-shot object detection methods suffer from limitations in learning representative prototypes. Specifically, directly aggregating bounding box contents from support images into prototypes renders these methods vulnerable to background noise and the morphological intricacies of objects. Furthermore, these methods neglect the varied contributions of intra-class image-specific prototypes and fail to leverage semantic information effectively during prototype generation, resulting in suboptimal class representations due to naive average aggregation. To address these issues, we propose a Dual Prototype-Enhancement Network (DPENet), designed to optimize prototypes by improving support feature representation and enhancing prototype discriminability. Specifically, we introduce an Object Enhancement Module (OEM) based on dynamic hypergraph construction. This module employs hypergraph convolution to adaptively capture complex high-order semantic interactions among highly similar regions within support features, thereby highlighting salient features of target regions, suppressing background noise, and enhancing support feature representation. Moreover, we propose a Semantic Fusion Perception Module (SFPM) that generates more discriminative class-specific prototypes by integrating weighted intra-class prototype representations with text-based semantic embeddings. Experimental results demonstrate that DPENet significantly outperforms existing methods on the PASCAL VOC and MS COCO datasets.
Jingling Huang, Hanzi Wang, Qiangqiang Wu, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Image Process.3
2025 DistinctAD: Distinctive Audio Description Generation in Contexts
abstract
Audio Descriptions (ADs) aim to provide a narration of a movie in text form, describing non-dialogue-related narratives, such as characters, actions, or scene establishment. Automatic generation of ADs remains challenging due to: i) the domain gap between movie-AD data and existing data used to train vision-language models, and ii) the issue of contextual redundancy arising from highly similar neighboring visual clips in a long movie. In this work, we propose DistinctAD, a novel two-stage framework for generating ADs that emphasize distinctiveness to produce better narratives. To address the domain gap, we introduce a CLIP-AD adaptation strategy that does not require additional AD corpora, enabling more effective alignment between movie and AD modalities at both global and finegrained levels. In Stage-II, DistinctAD incorporates two key innovations: (i) a Contextual Expectation-Maximization Attention (EMA) module that reduces redundancy by extracting common bases from consecutive video clips, and (ii) an explicit distinctive word prediction loss that filters out repeated words in the context, ensuring the prediction of unique terms specific to the current AD. Comprehensive evaluations on MAD-Eval, CMD-AD, and TV-AD benchmarks demonstrate the superiority of DistinctAD, with the model consistently outperforming baselines, particularly in Recall@k/N, highlighting its effectiveness in producing high-quality, distinctive ADs.
Bo Fang 0003, Qiangqiang Wu, YuXin Song 0001, Antoni B. Chan
CVPR3
2025 Temporal Unlearnable Examples: Preventing Personal Video Data from Unauthorized Exploitation by Object Tracking
abstract
With the rise of social media, vast amounts of user-uploaded videos (e.g., YouTube) are utilized as training data for Visual Object Tracking (VOT). However, the VOT community has largely overlooked video data-privacy issues, as many private videos have been collected and used for training commercial models without authorization. To alleviate these issues, this paper presents the first investigation on preventing personal video data from unauthorized exploitation by deep trackers. Existing methods for preventing unauthorized data use primarily focus on image-based tasks (e.g., image classification), directly applying them to videos reveals several limitations, including inefficiency, limited effectiveness, and poor generalizability. To address these issues, we propose a novel generative framework for generating Temporal Unlearnable Examples (TUEs), and whose efficient computation makes it scalable for usage on large-scale video datasets. The trackers trained w/ TUEs heavily rely on unlearnable noises for temporal matching, ignoring the original data structure and thus ensuring training video data-privacy. To enhance the effectiveness of TUEs, we introduce a temporal contrastive loss, which further corrupts the learning of existing trackers when using our TUEs for training. Extensive experiments demonstrate that our approach achieves state-of-the-art performance in video data-privacy protection, with strong transferability across VOT models, datasets, and temporal matching tasks.
Qiangqiang Wu, Yi Yu 0011, Chenqi Kong, Ziquan Liu, Jia Wan 0001, Haoliang Li, Alex Chichung Kot, Antoni B. Chan
ICCV1
2025 MGCA-Net: Multi-Graph Contextual Attention Network for Two-View Correspondence Learning
abstract
Two-view correspondence learning is a key task in computer vision, which aims to establish reliable matching relationships for applications such as camera pose estimation and 3D reconstruction. However, existing methods have limitations in local geometric modeling and cross-stage information optimization, which make it difficult to accurately capture the geometric constraints of matched pairs and thus reduce the robustness of the model. To address these challenges, we propose a Multi-Graph Contextual Attention Network (MGCA-Net), which consists of a Contextual Geometric Attention (CGA) module and a Cross-Stage Multi-Graph Consensus (CSMGC) module. Specifically, CGA dynamically integrates spatial position and feature information via an adaptive attention mechanism and enhances the capability to capture both local and global geometric relationships. Meanwhile, CSMGC establishes geometric consensus via a cross-stage sparse graph network, ensuring the consistency of geometric information across different stages. Experimental results on two representative YFCC100M and SUN3D datasets show that MGCA-Net significantly outperforms existing SOTA methods in the outlier rejection and camera pose estimation tasks. Source code is available at http://www.linshuyuan.com.
Shuyuan Lin, Mengtin Lo, Haosheng Chen 0001, Qiangqiang Wu
IJCAI5
2025 Progressive Semantic-Visual Alignment and Refinement for Vision-Language Tracking
abstract
In recent years, vision-language tracking has drawn emerging attention in the tracking field. The critical challenge for the task is to fuse semantic representations of language information and visual representations of vision information. For this purpose, several vision-language tracking methods perform early or late fusion to fuse visual and semantic features. However, these methods cannot take full advantage of the transformer architecture to excavate useful cross-modal context at various levels. To this end, we propose a new progressive joint vision-language transformer (PJVLT) to progressively align and refine visual embedding with semantic embedding for vision-language tracking. Specifically, to align visual signals with semantic signals, we propose to insert a semantic-aware instance encoder layer (SAIEL) into each intermediate layer of transformer encoder to perform progressive alignment of visual and semantic features. Furthermore, to highlight the multi-modal feature channels and patches corresponding to target objects, we propose a unified channel communication patch interaction layer (CCPIL), which is plugged into each intermediate layer of transformer encoder to progressively activate target-aware channels and patches of aligned multi-modal features for fine-grained tracking. In general, by progressively aligning and refining visual features with semantic features in the transformer encoder, our PJVLT can adaptively excavate well-aligned vision-language context at coarse-to-fine levels, therefore highlighting target objects at various levels for more discriminative tracking. Experiments on several tracking datasets show that the proposed PJVLT can achieve favorable performance in comparison with both conventional trackers and other vision-language trackers.
Qiangqiang Wu, Changqun Xia, Jia Li 0003
IEEE Trans. Circuits Syst. Video Technol.2
2025 CGATracker: Correlation-Aware Graph Alignment for Referring Multi-Object Tracking
abstract
Referring multi-object tracking (RMOT) aims to identify specific targets based on sentence descriptions. To enhance multi-modal learning, previous works typically relied on a simple fusion module at early or late stages. However, those methods frequently underutilize textual semantics and struggle to model the relationships between region-level features and word-level features. To address these limitations, we propose CGATracker, a correlation-aware graph alignment method for RMOT, which facilitates precise relationship modeling through relational scoring. Specifically, we design a Language-driven Relational Alignment (LRA) module, which establishes two connection graphs to generate positive and negative samples for the visual-textual alignment. Additionally, to effectively leverage referring information, we introduce a Semantic Clarify Booster (SCBooster) module based on a semantic infusion mechanism and a bias-aware verification mechanism for interactions with different modalities. Moreover, by designing a Multi-level Cross-modal Fusion (MCF) module, our method aggregates contextual features at multiple depths to enable the creation of the enriched correlation-aware graph. Extensive experiments conducted on the Refer-KITTI and Refer-KITTI-V2 datasets demonstrate the effectiveness of CGATracker.
Siping Zhuang, Qiangqiang Wu, Yang Lu 0009, Hai-Miao Hu, Hanzi Wang
IEEE Trans. Circuits Syst. Video Technol.3
2024 Robust Zero-Shot Crowd Counting and Localization With Adaptive Resolution SAM
Jia Wan 0001, Qiangqiang Wu, Wei Lin 0018, Antoni B. Chan
ECCV (57)2
2024 Boosting 3D Single Object Tracking with 2D Matching Distillation and 3D Pre-training
Qiangqiang Wu, Yan Xia 0003, Jia Wan 0001, Antoni B. Chan
ECCV (12)1
2024 Joint Spatio-Temporal Similarity and Discrimination Learning for Visual Tracking
abstract
Visual tracking is a task of localizing a target unceasingly in a video with an initial target state at the first frame. The limited target information makes this problem an extremely challenging task. Existing tracking methods either perform matching based similarity learning or optimization based discrimination reasoning. However, these two types of tracking methods suffer from the problem of ineffectiveness for distinguishing target objects from background distractors and the problem of insufficiency in maintaining spatio-temporal consistency among successive frames, respectively. In this paper, we design a joint spatio-temporal similarity and discrimination learning (STSDL) framework for accurate and robust tracking. The designed framework is composed of two complementary branches: a similarity learning branch and a discrimination learning branch. The similarity learning branch uses an effective transformer encoder-decoder to gather rich spatio-temporal context information to generate a similarity map. In parallel, the discrimination learning branch exploits an efficient model predictor to train a target model to produce a discriminative map. Finally, the similarity map and the discriminative map are adaptively fused for accurate and robust target localization. Experimental results on six prevalent datasets demonstrate that the proposed STSDL can obtain satisfactory results, while it retains a real-time tracking speed of 50 FPS on a single GPU.
Haosheng Chen 0001, Qiangqiang Wu, Changqun Xia, Jia Li 0003
IEEE Trans. Circuits Syst. Video Technol.3
2024 DHM-Net: Deep Hypergraph Modeling for Robust Feature Matching
abstract
We present a novel deep hypergraph modeling architecture (called DHM-Net) for feature matching in this paper. Our network focuses on learning reliable correspondences between two sets of initial feature points by establishing a dynamic hypergraph structure that models group-wise relationships and assigns weights to each node. Compared to existing feature matching methods that only consider pair-wise relationships via a simple graph, our dynamic hypergraph is capable of modeling nonlinear higher-order group-wise relationships among correspondences in an interaction capturing and attention representation learning fashion. Specifically, we propose a novel Deep Hypergraph Modeling block, which initializes an overall hypergraph by utilizing neighbor information, and then adopts node-to-hyperedge and hyperedge-to-node strategies to propagate interaction information among correspondences while assigning weights based on hypergraph attention. In addition, we propose a Differentiation Correspondence-Aware Attention mechanism to optimize the hypergraph for promoting representation learning. The proposed mechanism is able to effectively locate the exact position of the object of importance via the correspondence aware encoding and simple feature gating mechanism to distinguish candidates of inliers. In short, we learn such a dynamic hypergraph format that embeds deep group-wise interactions to explicitly infer categories of correspondences. To demonstrate the effectiveness of DHM-Net, we perform extensive experiments on both real-world outdoor and indoor datasets. Particularly, experimental results show that DHM-Net surpasses the state-of-the-art method by a sizable margin. Our approach obtains an 11.65% improvement under error threshold of 5° for relative pose estimation task on YFCC100M dataset. Code will be released at https://github.com/CSX777/DHM-Net.
Shunxing Chen, Guobao Xiao, Junwen Guo, Qiangqiang Wu, Jiayi Ma 0001
IEEE Trans. Image Process.4
2023 DropMAE: Masked Autoencoders with Spatial-Attention Dropout for Tracking Tasks
abstract
In this paper, we study masked autoencoder (MAE) pretraining on videos for matching-based downstream tasks, including visual object tracking (VOT) and video object segmentation (VOS). A simple extension of MAE is to randomly mask out frame patches in videos and reconstruct the frame pixels. However, we find that this simple baseline heavily relies on spatial cues while ignoring temporal relations for frame reconstruction, thus leading to sub-optimal temporal matching representations for VOT and VOS. To alleviate this problem, we propose DropMAE, which adaptively performs spatial-attention dropout in the frame reconstruction to facilitate temporal correspondence learning in videos. We show that our DropMAE is a strong and efficient temporal matching learner, which achieves better finetuning results on matching-based tasks than the ImageNet-based MAE with$2\times$faster pre-training speed. Moreover, we also find that motion diversity in pre-training videos is more important than scene diversity for improving the performance on VOT and VOS. Our pre-trained DropMAE model can be directly loaded in existing ViT-based trackers for fine-tuning without further modifications. Notably, DropMAE sets new state-of-the-art performance on 8 out of 9 highly competitive video tracking and segmentation datasets. Our code and pre-trained models are available at https://github.com/jimmy-dq/DropMAE.git.
Qiangqiang Wu, Tianyu Yang 0003, Ziquan Liu, Baoyuan Wu, Ying Shan, Antoni B. Chan
CVPR1
2023 TORE: Token Reduction for Efficient Human Mesh Recovery with Transformer
abstract
In this paper, we introduce a set of simple yet effective TOken REduction (TORE) strategies for Transformer-based Human Mesh Recovery from monocular images. Current SOTA performance is achieved by Transformer-based structures. However, they suffer from high model complexity and computation cost caused by redundant tokens. We propose token reduction strategies based on two important aspects, i.e., the 3D geometry structure and 2D image feature, where we hierarchically recover the mesh geometry with priors from body structure and conduct token clustering to pass fewer but more discriminative image feature tokens to the Transformer. Our method massively reduces the number of tokens involved in high-complexity interactions in the Transformer. This leads to a significantly reduced computational cost while still achieving competitive or even higher accuracy in shape recovery. Extensive experiments across a wide range of benchmarks validate the superior effectiveness of the proposed method. We further demonstrate the generalizability of our method on hand mesh recovery. Visit our project page at https://frank-zy-dou.github.io/projects/Tore/index.html.
Zhiyang Dou, Qingxuan Wu, Cheng Lin 0001, Zeyu Cao, Qiangqiang Wu, Weilin Wan 0001, Taku Komura, Wenping Wang 0001
ICCV5
2023 Scalable Video Object Segmentation with Simplified Framework
abstract
The current popular methods for video object segmentation (VOS) implement feature matching through several hand-crafted modules that separately perform feature extraction and matching. However, the above hand-crafted designs empirically cause insufficient target interaction, thus limiting the dynamic target-aware feature learning in VOS. To tackle these limitations, this paper presents a scalable Simplified VOS (SimVOS) framework to perform joint feature extraction and matching by leveraging a single transformer backbone. Specifically, SimVOS employs a scalable ViT backbone for simultaneous feature extraction and matching between query and reference features. This design enables SimVOS to learn better target-ware features for accurate mask prediction. More importantly, SimVOS could directly apply well-pretrained ViT backbones (e.g., MAE [21]) for VOS, which bridges the gap between VOS and large-scale self-supervised pre-training. To achieve a better performance-speed trade-off, we further explore within-frame attention and propose a new token refinement module to improve the running speed and save computational cost. Experimentally, our SimVOS achieves state-of-the-art results on popular video object segmentation benchmarks, i.e., DAVIS-2017 (88.0% $\mathcal{J}\& \mathcal{F}$), DAVIS-2016 (92.9% $\mathcal{J}\& \mathcal{F}$) and YouTube-VOS 2019 (84.2% $\mathcal{J}\& \mathcal{F}$), without applying any synthetic video or BL30K pre-training used in previous VOS approaches. Our code and models are available at https://github.com/jimmy-dq/SimVOS.git.
Qiangqiang Wu, Tianyu Yang 0003, Antoni B. Chan
ICCV1
2023 Modeling Noisy Annotations for Point-Wise Supervision
abstract
Point-wise supervision is widely adopted in computer vision tasks such as crowd counting and human pose estimation. In practice, the noise in point annotations may affect the performance and robustness of algorithm significantly. In this paper, we investigate the effect of annotation noise in point-wise supervision and propose a series of robust loss functions for different tasks. In particular, the point annotation noise includes spatial-shift noise, missing-point noise, and duplicate-point noise. The spatial-shift noise is the most common one, and exists in crowd counting, pose estimation, visual tracking, etc, while the missing-point and duplicate-point noises usually appear in dense annotations, such as crowd counting. In this paper, we first consider the shift noise by modeling the real locations as random variables and the annotated points as noisy observations. The probability density function of the intermediate representation (a smooth heat map generated from dot annotations) is derived and the negative log likelihood is used as the loss function to naturally model the shift uncertainty in the intermediate representation. The missing and duplicate noise are further modeled by an empirical way with the assumption that the noise appears at high density region with a high probability. We apply the method to crowd counting, human pose estimation and visual tracking, propose robust loss functions for those tasks, and achieve superior performance and robustness on widely used datasets.
Jia Wan 0001, Qiangqiang Wu, Antoni B. Chan
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 A Lightweight and Detector-Free 3D Single Object Tracker on Point Clouds
abstract
Recent works on 3D single object tracking treat the task as a target-specific 3D detection task, where an off-the-shelf 3D detector is commonly employed for the tracking. However, it is non-trivial to perform accurate target-specific detection since the point cloud of objects in raw LiDAR scans is usually sparse and incomplete. In this paper, we address this issue by explicitly leveraging temporal motion cues and propose DMT, a Detector-free Motion-prediction-based 3D Tracking network that completely removes the usage of complicated 3D detectors and is lighter, faster, and more accurate than previous trackers. Specifically, the motion prediction module is first introduced to estimate a potential target center of the current frame in a point-cloud-free manner. Then, an explicit voting module is proposed to directly regress the 3D box from the estimated target center. Extensive experiments on KITTI and NuScenes datasets demonstrate that our DMT can still achieve better performance ($\sim $10% improvement over the NuScenes dataset) and a faster tracking speed (i.e., 72 FPS) than state-of-the-art approaches without applying any complicated 3D detectors. Our code is released athttps://github.com/jimmy-dq/DMT.
Yan Xia 0003, Qiangqiang Wu, Wei Li 0111, Antoni B. Chan, Uwe Stilla
IEEE Trans. Intell. Transp. Syst.2
2022 A New Framework for Multiple Deep Correlation Filters Based Object Tracking
abstract
In recent years, Correlation Filter (CF) based tracking methods using Convolutional Neural Network (CNN) features have achieved the state-of-the-art performance for object tracking. However, how to design an efficient deep CF based tracking method has not been well studied in the literature. To address this issue, we first develop a generic framework, which breaks a deep CF based tracking method into five components, including motion model, CNN feature extractor, CF model, CF updater, and location model. According to this framework, we design each component step by step. Then we propose a novel deep CF based tracking method by combining five effective components together. The proposed method outperforms several state-of-the-art tracking methods on two tracking benchmarks. Then the ablative experiments are conducted to study the influence of each component. The results show that the CF model and the CNN feature extractor play the most important roles in a deep CF based tracking method. Moreover, the CF updater, the location model, and the motion model can also improve the performance substantially.
Qiangqiang Wu, Liming Zhang 0002, Hanzi Wang
ICASSP3
2022 Deep Correlation Filter Tracking With Shepherded Instance-Aware Proposals
abstract
Visual tracking is a core component of intelligent transportation systems and it is crucial to reduce or avoid traffic accidents. Recently, deep correlation filter (DCF) based trackers have exhibited good tracking performance. However, existing DCF based trackers are still ineffective to cope with large scale variations and severe distortions (e.g., heavy occlusions, significant deformations, large rotations, etc.), leading to the inferior performance. To address these issues, we develop a novel DeepCFIAP++ tracker, which incorporates effective shepherded instance-aware proposals into DCFs. DeepCFIAP++ can not only estimate the target scale at every frame but also re-detect the target in the case of severe distortions. Firstly, we propose to exploit both color and edge cues to generate complementary detection proposals to effectively handle various challenging scenarios. Then, we propose to utilize multi-layer target-specific deep features to rank the generated detection proposals and choose the instance-aware proposals, which will result in more robust tracking performance. Finally, we propose to use the DCFs to shepherd the instance-aware proposals toward their best locations, which will result in more accurate tracking results. Experimental results on five challenging datasets (i.e., OTB2013, OTB2015, VOT2016, VOT2017 and UAV20L) demonstrate that DeepCFIAP++ performs competitively with several other state-of-the-art DCF based trackers.
Qiangqiang Wu, Yan Yan 0001, Hanzi Wang
IEEE Trans. Intell. Transp. Syst.2
2021 Progressive Unsupervised Learning for Visual Object Tracking
abstract
In this paper, we propose a progressive unsupervised learning (PUL) framework, which entirely removes the need for annotated training videos in visual tracking. Specifically, we first learn a background discrimination (BD) model that effectively distinguishes an object from back-ground in a contrastive learning way. We then employ the BD model to progressively mine temporal corresponding patches (i.e., patches connected by a track) in sequential frames. As the BD model is imperfect and thus the mined patch pairs are noisy, we propose a noise-robust loss function to more effectively learn temporal correspondences from this noisy data. We use the proposed noise robust loss to train backbone networks of Siamese trackers. Without online fine-tuning or adaptation, our unsupervised real-time Siamese trackers can outperform state-of-the-art unsupervised deep trackers and achieve competitive results to the supervised baselines.
Qiangqiang Wu, Jia Wan 0001, Antoni B. Chan
CVPR1
2021 Meta-Graph Adaptation for Visual Object Tracking
abstract
Existing deep trackers typically use offline-learned backbone networks for feature extraction across various online tracking tasks. However, for unseen objects, offline-learned representations are still limited due to the lack of adaptation. In this paper, we propose a Meta-Graph Adaptation Network (MGA-Net) to adapt backbones of deep trackers to specific online tracking tasks in a meta-learning fashion. Our MGA-Net is composed of a gradient embedding module (GEM) and a filter adaptation module (FAM). GEM takes gradients as an adaptation signal, and applies graph-message propagation to learn smoothed low-dimensional gradient embeddings. FAM utilizes both the learned gradient embeddings and the target exemplar to adapt the filter weights for the specific tracking task. MGA-Net can be end-to-end trained in an offline meta-learning way, and runs completely feed-forward for testing, thus enabling highly-efficient online tracking. We show that MGA-Net is generic and demonstrate its effectiveness in both template matching and correlation filter tracking frameworks.
Qiangqiang Wu, Antoni B. Chan
ICME1
2021 Dynamic Momentum Adaptation for Zero-Shot Cross-Domain Crowd Counting
abstract
Zero-shot cross-domain crowd counting is a challenging task where a crowd counting model is trained on a source domain (i.e., training dataset) and no additional labeled or unlabeled data is available for fine-tuning the model when testing on an unseen target domain (i.e., a different testing dataset). The generalisation performance of existing crowd counting methods is typically limited due to the large gap between source and target domains. Here, we propose a novel Crowd Counting framework built upon an external Momentum Template, termed C2MoT, which enables the encoding of domain specific information via an external template representation. Specifically, the Momentum Template (MoT) is learned in a momentum updating way during offline training, and then is dynamically updated for each test image in online cross-dataset evaluation. Thanks to the dynamically updated MoT, our C2MoT effectively generates dense target correspondences that explicitly accounts for head regions, and then effectively predicts the density map based on the normalized correspondence map. Experiments on large scale datasets show that our proposed C2MoT achieves leading zero-shot cross-domain crowd counting performance without model fine-tuning, while also outperforming domain adaptation methods that use fine-tuning on target domain data. Moreover, C2MoT also obtains state-of-the-art counting performance on the source domain.
Qiangqiang Wu, Jia Wan 0001, Antoni B. Chan
ACM Multimedia1
2020 End-to-End Learning of Object Motion Estimation from Retinal Events for Event-Based Object Tracking
abstract
Event cameras, which are asynchronous bio-inspired vision sensors, have shown great potential in computer vision and artificial intelligence. However, the application of event cameras to object-level motion estimation or tracking is still in its infancy. The main idea behind this work is to propose a novel deep neural network to learn and regress a parametric object-level motion/transform model for event-based object tracking. To achieve this goal, we propose a synchronous Time-Surface with Linear Time Decay (TSLTD) representation, which effectively encodes the spatio-temporal information of asynchronous retinal events into TSLTD frames with clear motion patterns. We feed the sequence of TSLTD frames to a novel Retinal Motion Regression Network (RMRNet) to perform an end-to-end 5-DoF object motion regression. Our method is compared with state-of-the-art object tracking methods, that are based on conventional cameras or event cameras. The experimental results show the superiority of our method in handling various challenging environments such as fast motion and low illumination conditions.
Haosheng Chen 0001, David Suter, Qiangqiang Wu, Hanzi Wang
AAAI3
2019 Asynchronous Tracking-by-Detection on Adaptive Time Surfaces for Event-based Object Tracking
abstract
Event cameras, which are asynchronous bio-inspired vision sensors, have shown great potential in a variety of situations, such as fast motion and low illumination scenes. However, most of the event-based object tracking methods are designed for scenarios with untextured objects and uncluttered backgrounds. There are few event-based object tracking methods that support bounding box-based object tracking. The main idea behind this work is to propose an asynchronous Event-based Tracking-by-Detection (ETD) method for generic bounding box-based object tracking. To achieve this goal, we present an Adaptive Time-Surface with Linear Time Decay (ATSLTD) event-to-frame conversion algorithm, which asynchronously and effectively warps the spatio-temporal information of asynchronous retinal events to a sequence of ATSLTD frames with clear object contours. We feed the sequence of ATSLTD frames to the proposed ETD method to perform accurate and efficient object tracking, which leverages the high temporal resolution property of event cameras. We compare the proposed ETD method with seven popular object tracking methods, that are based on conventional cameras or event cameras, and two variants of ETD. The experimental results show the superiority of the proposed ETD method in handling various challenging environments.
Haosheng Chen 0001, Qiangqiang Wu, Xinbo Gao 0001, Hanzi Wang
ACM Multimedia2
2018 DSNet: Deep and Shallow Feature Learning for Efficient Visual Tracking
Qiangqiang Wu, Yan Yan 0001, Hanzi Wang
ACCV (5)1
2018 Robust Correlation Filter Tracking with Shepherded Instance-Aware Proposals
abstract
In recent years, convolutional neural network (CNN) based correlation filter trackers have achieved state-of-the-art results on the benchmark datasets. However, the CNN based correlation filters cannot effectively handle large scale variation and distortion (such as fast motion, background clutter, occlusion, etc.), leading to the sub-optimal performance. In this paper, we propose a novel CNN based correlation filter tracker with shepherded instance-aware proposals, namely DeepCFIAP, which automatically estimates the target scale in each frame and re-detects the target when distortion happens. DeepCFIAP is proposed to take advantage of the merits of both instance-aware proposals and CNN based correlation filters. Compared with the CNN based correlation filter trackers, DeepCFIAP can successfully solve the problems of large scale variation and distortion via the shepherded instance-aware proposals, resulting in more robust tracking performance. Specifically, we develop a novel proposal ranking algorithm based on the similarities between proposals and instances. In contrast to the detection proposal based trackers, DeepCFIAP shepherds the instance-aware proposals towards their optimal positions via the CNN based correlation filters, resulting in more accurate tracking results. Extensive experiments on two challenging benchmark datasets demonstrate that the proposed DeepCFIAP performs favorably against state-of-the-art trackers and it is especially feasible for long-term tracking.
Qiangqiang Wu, Yan Yan 0001, Hanzi Wang
ACM Multimedia2