EDBT 2026 Demo / reviewers in the wild / expert
Ning Wang 0020
dblp:46/2005-20
· DBLP profile ↗
28ranked-venue papers
15as first author
17since 2021 · last 2026
0000-0002-4937-6784ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 11 first-author · 8 since 2021Artificial intelligence and machine learning · 16 · 10 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Refine, Control and Distill: A Text-to-Image Framework for Faithful Image GenerationabstractWhile text-to-image diffusion models exhibit outstanding results, they struggle to faithfully generate key subjects with corresponding attributes in prompts, challenges known as catastrophic neglect and attribute binding. Previous works typically utilize attention adjustments to solve the above problems, whereas we observe that they may still generate unfaithful images. In this paper, we carefully analyze the text-to-image process and pinpoint three pivotal bottlenecks that hinder image faithful generation: (1) unequal responses of neglected subjects in text embedding, (2) competition and entanglement between subjects' attention, and (3) suboptimal quality of intermediate features from U-Net. Based on the aforementioned observations, we propose a Refine, Control, and Distill (RCD) framework built upon the stable diffusion model to alleviate the negative effects raised by the bottlenecks mentioned above, respectively. Specifically, we achieve the above goals through a text embedding refinement module, three region-level attention control losses, and self-distillation of intermediate semantic features in the denoising process. Our approach exhibits promising capability in generating faithful and high-quality images and outperforms state-of-the-art methods through extensive quantitative and qualitative evaluations on recent advanced base diffusion models. Peng Xing, Ning Wang 0020, Yanpeng Sun, Jinhui Tang 0001, Zechao Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | APSam: An Aggregating-Then-Pruning Sampler for Question-Conditional DenoisingabstractVideo question answering (VideoQA) necessitates simultaneous understanding of visual and linguistic information, requiring both in-depth analysis of individual modality features and the establishment of cross-modal correlations to achieve precise reasoning. However, VideoQA models often struggle with irrelevant temporal and spatial noise due to the dense events and concepts in real-world complex video contents. Previous works reduce noise by only sampling a fixed number of visual tokens at the patch level, overlooking the variation in the required granularities of features and quantities of visual cues across different question conditions. To address these, we propose an Aggregating-then-Pruning Sampler (APSam), which diversifies feature granularities and adaptively denoises on a per-question basis. Specifically, we propose a conditional token aggregator to obtain multi-granularity visual semantics by merging similar question-relevant tokens. Then, we propose a conditional token pruner, which restricts noise tokens through a variable-capacity receptive field determined by the inputs. Experimental results show that APSam achieves significant performance on three challenging complex VideoQA datasets,i.e., AGQAv2, NExT-QA, and STAR. Further analyses reveal that the APSam also exhibits high reasoning capability and interpretability. Jiafeng Liang, Shixin Jiang, Wei Tang 0015, Ning Wang 0020, Zekun Wang 0001, Xun Mao, Ming Liu 0004, Bing Qin 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal InconsistencyabstractJiafeng Liang, Shixin Jiang, Xuan Dong, Ning Wang, Zheng Chu, Hui Su, Jinlan Fu, Ming Liu, See-Kiong Ng, Bing Qin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jiafeng Liang, Shixin Jiang, Ning Wang 0020, Hui Su, Jinlan Fu, Ming Liu 0004, See-Kiong Ng, Bing Qin 0001 |
ACL (1) | 4 |
| 2025 | Inv-Adapter: ID Customization Generation via Image Inversion and Lightweight Parameter AdapterabstractThe remarkable advancement in text-to-image generation models significantly boosts the research in ID customization generation. However, existing personalization methods cannot simultaneously satisfy high-fidelity and low-costs requirements. Their main bottleneck lies in the additional prompt image encoder (i.e., CLIP vision encoder), which produces weak alignment signals with the text-to-image model that may lose face information and is not well 'absorbed' by the text-to-image model. Towards this end, we propose Inv-Adapter, which first introduces a more reasonable and efficient token representation of ID image features and introduces a lightweight parameter adaptor to inject ID features. Specifically, our Inv-Adapter extracts diffusion-domain representations of ID images utilizing a pre-trained text-to-image model via DDIM image inversion, without an additional image encoder. Benefiting from the high alignment of the extracted ID prompt features and the intermediate features of the text-to-image model, we then introduce a lightweight attention adapter to embed them efficiently into the base text-to-image model. We conduct extensive experiments on different text-to-image models to assess ID fidelity, generation loyalty, speed, training costs, model scale and generalization ability in scenarios of general object, all of which show that the proposed Inv-Adapter is highly competitive in ID customization generation and model scale. Peng Xing, Ning Wang 0020, Jianbo Ouyang, Zechao Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | CAT: A Simple yet Effective Cross-Attention Transformer for One-Shot Object Detection
Yuyan Deng, Yang Gao 0001, Ning Wang 0020, Lingqiao Liu, Lei Zhang 0054, Peng Wang 0015 |
J. Comput. Sci. Technol. | 4 |
| 2023 | Searching sharing relationship for instance segmentation decoder
Yuling Xi, Ning Wang 0020, Shaohua Wan 0001, Xiaoming Wang 0010, Peng Wang 0015, Yanning Zhang 0001 |
Appl. Intell. | 2 |
| 2023 | A Dynamic Feature Interaction Framework for Multi-task Visual Perception
Yuling Xi, Hao Chen 0041, Ning Wang 0020, Peng Wang 0015, Yanning Zhang 0001, Chunhua Shen, Yifan Liu 0001 |
Int. J. Comput. Vis. | 3 |
| 2023 | Coloring anime line art videos with transformation region enhancement networkabstractAutomatic colorization of anime line art videos aims to produce color frames given line art frames and reference color images, which is challenging due to various motions and geometric transformations across frame sequences. Existing methods usually utilize the feature maps of reference images directly and treat all the regions in an image equally. However, this may overlook the details of the regions undergoing geometric transformations . To emphasize the regions with significant transformations between the reference and target frames, we propose a Transformation Region Enhancement Network (TRE-Net) to exploit useful reference information and enhance the colorization of key transformation regions with Region Localization Module (RLM) and Feature Enhancement Module (FEM). Specifically, we propose Multi-scale Euclidean Distance Difference (Multi-scale EDD) Maps in RLM which effectively locate geometric transformation regions by contrasting the Euclidean Distance Maps of two line arts and aggregating representations at multiple scales of the network. In addition, FEM is devised to enhance feature learning in the regions with geometric transformation and to ensure proper color alignment. FEM learns locally enhanced features through an attention-gating operation at a low computational cost. With the well-represented key geometric transformation regions, our method exploits the multi-scale reference information well for color alignment, thus produces perceptually pleasing frames. Comprehensive experimental results show that our proposed method is superior to existing methods in terms of the overall quality of colorized anime line art videos. Ning Wang 0020, Muyao Niu, Zhi Dou, Zhihui Wang 0001, Zhiyong Wang 0001, Zhaoyan Ming, Bin Liu 0040 |
Pattern Recognit. | 1 |
| 2022 | Improving Image Captioning via Enhancing Dual-Side Context AwarenessabstractRecent work on visual question answering demonstrate that grid features can work as well as region feature on vision language tasks. In the meantime, transformer-based model and its variants have shown remarkable performance on image captioning. However, the object-contextual information missing caused by the single granularity nature of grid feature on the encoder side, as well as the future contextual information missing due to the left2right decoding paradigm of transformer decoder, remains unexplored. In this work, we tackle these two problems by enhancing contextual information at dual-side:(i) at encoder side, we propose Context-Aware Self-Attention module, in which the key/value is expanded with adjacent rectangle region where each region contains two or more aggregated grid features; this enables grid feature with varying granularity, storing adequate contextual information for object with different scale. (ii) at decoder side, we incorporate a dual-way decoding strategy, in which left2right and right2left decoding are conducted simultaneously and interactively. It utilizes both past and future contextual information when generates current word. Combining these two modules with a vanilla transformer, our Context-Aware Transformer(CATNet) achieves a new state-of-the-art on MSCOCO benchmark. Yiqi Gao, Ning Wang 0020, Wei Suo, Mengyang Sun, Peng Wang 0015 |
ICMR | 2 |
| 2022 | Learning Temporal-Correlated and Channel- Decorrelated Siamese Networks for Visual TrackingabstractRecently, Siamese network based trackers have attracted growing popularity in visual tracking, which tackle the tracking by template matching between the initial template and successive search regions. The initial template patch is generally encoded into a convolutional feature for matching. However, the limited representational capability of the template feature limits the tracking accuracy. Besides, this fixed representation also fails to adapt to the target appearance changes. To alleviate these issues, we improve the Siamese trackers by introducing temporal correlation and channel decorrelation mechanisms. On the one hand, we consider the channel-wise correlations between the initial and historical template features to adaptively aggregate informative channel-wise representations for template update. On the other hand, we propose a decorrelation regularization to weaken the channel-wise correlations of individual template features. By end-to-end training, we learn a more complete and adaptive template for accurate object tracking. We demonstrate the generality of our approach by applying it to two prevalent Siamese trackers, i.e., SiamFC and SiamRPN. Extensive experiments on seven benchmark datasets verify the effectiveness of our method. Mao Xi, Wengang Zhou 0001, Ning Wang 0020, Houqiang Li |
IEEE Trans. Multim. | 3 |
| 2021 | Contrastive Transformation for Self-supervised Correspondence LearningabstractIn this paper, we focus on the self-supervised learning of visual correspondence using unlabeled videos in the wild. Our method simultaneously considers intra- and inter-video representation associations for reliable correspondence estimation. The intra-video learning transforms the image contents across frames within a single video via the frame pair-wise affinity. To obtain the discriminative representation for instance-level separation, we go beyond the intra-video analysis and construct the inter-video affinity to facilitate the contrastive transformation across different videos. By forcing the transformation consistency between intra- and inter-video levels, the fine-grained correspondence associations are well preserved and the instance-level feature discrimination is effectively reinforced. Our simple framework outperforms the recent self-supervised correspondence methods on a range of visual tasks including video object tracking (VOT), video object segmentation (VOS), pose keypoint tracking, etc. It is worth mentioning that our method also surpasses the fully-supervised affinity representation (e.g., ResNet) and performs competitively against the recent fully-supervised algorithms designed for the specific tasks (e.g., VOT and VOS). Ning Wang 0020, Wengang Zhou 0001, Houqiang Li |
AAAI | 1 |
| 2021 | Transformer Meets Tracker: Exploiting Temporal Context for Robust Visual TrackingabstractIn video object tracking, there exist rich temporal contexts among successive frames, which have been largely overlooked in existing trackers. In this work, we bridge the individual video frames and explore the temporal contexts across them via a transformer architecture for robust object tracking. Different from classic usage of the transformer in natural language processing tasks, we separate its encoder and decoder into two parallel branches and carefully design them within the Siamese-like tracking pipelines. The transformer encoder promotes the target templates via attention-based feature reinforcement, which benefits the high-quality tracking model generation. The transformer decoder propagates the tracking cues from previous templates to the current frame, which facilitates the object searching process. Our transformer-assisted tracking framework is neat and trained in an end-to-end manner. With the proposed transformer, a simple Siamese matching approach is able to outperform the current top-performing trackers. By combining our transformer with the recent discriminative tracking pipeline, our method sets several new state-of-the-art records on prevalent tracking benchmarks. Ning Wang 0020, Wengang Zhou 0001, Jie Wang 0005, Houqiang Li |
CVPR | 1 |
| 2021 | Joint Inductive and Transductive Learning for Video Object SegmentationabstractSemi-supervised video object segmentation is a task of segmenting the target object in a video sequence given only a mask annotation in the first frame. The limited information available makes it an extremely challenging task. Most previous best-performing methods adopt matching-based transductive reasoning or online inductive learning. Nevertheless, they are either less discriminative for similar instances or insufficient in the utilization of spatio-temporal information. In this work, we propose to integrate transductive and inductive learning into a unified framework to exploit the complementarity between them for accurate and robust video object segmentation. The proposed approach consists of two functional branches. The transduction branch adopts a lightweight transformer architecture to aggregate rich spatio-temporal cues while the induction branch performs online inductive learning to obtain discriminative target information. To bridge these two diverse branches, a two-head label encoder is introduced to learn the suitable target prior for each of them. The generated mask encodings are further forced to be disentangled to better retain their complementarity. Extensive experiments on several prevalent benchmarks show that, without the need of synthetic training data, the proposed approach sets a series of new state-of-the-art records. Code is available at https://github.com/maoyunyao/JOINT. Yunyao Mao, Ning Wang 0020, Wengang Zhou 0001, Houqiang Li |
ICCV | 2 |
| 2021 | NAS-FCOS: Efficient Search for Object Detection Architectures
Ning Wang 0020, Yang Gao 0001, Hao Chen 0041, Peng Wang 0015, Zhi Tian, Chunhua Shen, Yanning Zhang 0001 |
Int. J. Comput. Vis. | 1 |
| 2021 | Unsupervised Deep Representation Learning for Real-Time Tracking
Ning Wang 0020, Wengang Zhou 0001, Yibing Song, Chao Ma 0004, Wei Liu 0005, Houqiang Li |
Int. J. Comput. Vis. | 1 |
| 2021 | Cascaded Regression Tracking: Towards Online Hard Distractor DiscriminationabstractVisual can be easily disturbed by similar surrounding objects. Such objects as hard distractors, even though being the minority among negative samples, increase the risk of target drift and model corruption, which deserve additional attention in online tracking and model update. To enhance the tracking robustness, in this paper, we propose a cascaded regression tracker with two sequential stages. In the first stage, we filter out abundant easily-identified negative candidates via an efficient convolutional regression. In the second stage, a discrete sampling based ridge regression is designed to double-check the remaining ambiguous hard samples, which serves as an alternative of fully-connected layers and benefits from the closed-form solver for efficient learning. During the model update, we utilize the hard negative mining technique and an adaptive ridge regression scheme to improve the discrimination capability of the second-stage regressor. Extensive experiments are conducted on 11 challenging tracking benchmarks including OTB-2013, OTB-2015, VOT2018, VOT2019, UAV123, Temple-Color, NfS, TrackingNet, LaSOT, UAV20L, and OxUvA. The proposed method achieves state-of-the-art performance on prevalent benchmarks, while running in a real-time speed. Ning Wang 0020, Wengang Zhou 0001, Qi Tian 0001, Houqiang Li |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Learning Diverse Models for End-to-End Ensemble TrackingabstractIn visual tracking, how to effectively model the target appearance using limited prior information remains an open problem. In this paper, we leverage an ensemble of diverse models to learn manifold representations for robust object tracking. The proposed ensemble framework includes a shared backbone network for efficient feature extraction and multiple head networks for independent predictions. Trained by the shared data within an identical structure, the mutually correlated head models heavily hinder the potential of ensemble learning. To shrink the representational overlaps among multiple models while encouraging the diversity of individual predictions, we propose the model diversity and response diversity regularization terms during training. By fusing these distinctive prediction results via a fusion module, the tracking variance caused by the distractor objects can be largely restrained. Our whole framework is end-to-end trained in a data-driven manner, avoiding the heuristic designs of multiple base models and fusion strategies. The proposed method achieves state-of-the-art results on seven challenging benchmarks while operating in real-time. Ning Wang 0020, Wengang Zhou 0001, Houqiang Li |
IEEE Trans. Image Process. | 1 |
| 2020 | POST: POlicy-Based Switch TrackingabstractIn visual object tracking, by reasonably fusing multiple experts, ensemble framework typically achieves superior performance compared to the individual experts. However, the necessity of parallelly running all the experts in most existing ensemble frameworks heavily limits their efficiency. In this paper, we propose POST, a POlicy-based Switch Tracker for robust and efficient visual tracking. The proposed POST tracker consists of multiple weak but complementary experts (trackers) and adaptively assigns one suitable expert for tracking in each frame. By formulating this expert switch in consecutive frames as a decision-making problem, we learn an agent via reinforcement learning to directly decide which expert to handle the current frame without running others. In this way, the proposed POST tracker maintains the performance merit of multiple diverse models while favorably ensuring the tracking efficiency. Extensive ablation studies and experimental comparisons against state-of-the-art trackers on 5 prevalent benchmarks verify the effectiveness of the proposed method. Ning Wang 0020, Wengang Zhou 0001, Guo-Jun Qi, Houqiang Li |
AAAI | 1 |
| 2020 | Graph Few-shot Learning with Attribute MatchingabstractDue to the expensive cost of data annotation, few-shot learning has attracted increasing research interests in recent years. Various meta-learning approaches have been proposed to tackle this problem and have become the de facto practice. However, most of the existing approaches along this line mainly focus on image and text data in the Euclidean domain. However, in many real-world scenarios, a vast amount of data can be represented as attributed networks defined in the non-Euclidean domain, and the few-shot learning studies in such structured data have largely remained nascent. Although some recent studies have tried to combine meta-learning with graph neural networks to enable few-shot learning on attributed networks, they fail to account for the unique properties of attributed networks when creating diverse tasks in the meta-training phase---the feature distributions of different tasks could be quite different as instances (i.e., nodes) do not follow the data i.i.d. assumption on attributed networks. Hence, it may inevitably result in suboptimal performance in the meta-testing phase. To tackle the aforementioned problem, we propose a novel graph meta-learning framework--Attribute Matching Meta-learning Graph Neural Networks (AMM-GNN). Specifically, the proposed AMM-GNN leverages an attribute-level attention mechanism to capture the distinct information of each task and thus learns more effective transferable knowledge for meta-learning. We conduct extensive experiments on real-world datasets under a wide range of settings and the experimental results demonstrate the effectiveness of the proposed AMM-GNN framework. Ning Wang 0020, Minnan Luo, Kaize Ding, Lingling Zhang 0005, Jundong Li |
CIKM | 1 |
| 2020 | NAS-FCOS: Fast Neural Architecture Search for Object DetectionabstractThe success of deep neural networks relies on significant architecture engineering. Recently neural architecture search (NAS) has emerged as a promise to greatly reduce manual effort in network design by automatically searching for optimal architectures, although typically such algorithms need an excessive amount of computational resources, e.g., a few thousand GPU-days. To date, on challenging vision tasks such as object detection, NAS, especially fast versions of NAS, is less studied. Here we propose to search for the decoder structure of object detectors with search efficiency being taken into consideration. To be more specific, we aim to efficiently search for the feature pyramid network (FPN) as well as the prediction head of a simple anchor-free object detector, namely FCOS, using a tailored reinforcement learning paradigm. With carefully designed search space, search algorithms and strategies for evaluating network quality, we are able to efficiently search a top-performing detection architecture within 4 days using 8 V100 GPUs. The discovered architecture surpasses state-of-the-art object detection models (such as Faster R-CNN, RetinaNet and FCOS) by 1.5 to 3.5 points in AP on the COCO dataset, with comparable computation complexity and memory footprint, demonstrating the efficacy of the proposed NAS for object detection. Ning Wang 0020, Yang Gao 0001, Hao Chen 0041, Peng Wang 0015, Zhi Tian, Chunhua Shen, Yanning Zhang 0001 |
CVPR | 1 |
| 2020 | Hierarchical Representations with Discriminative Meta-filters in Dual Path Network for Tracking
Ning Wang 0020, Yuncong Yao, Wankou Yang, Kaihua Zhang 0001, Bo Liu 0005 |
PRCV (2) | 2 |
| 2020 | Real-Time Correlation Tracking Via Joint Model Compression and TransferabstractCorrelation filters (CF) have received considerable attention in visual tracking because of their computational efficiency. Leveraging deep features via off-the-shelf CNN models (e.g., VGG), CF trackers achieve state-of-the-art performance while consuming a large number of computing resources. This limits deep CF trackers to be deployed to many mobile platforms on which only a single-core CPU is available. In this paper, we propose to jointly compress and transfer off-the-shelf CNN models within a knowledge distillation framework. We formulate a CNN model pretrained from the image classification task as a teacher network, and distill this teacher network into a lightweight student network as the feature extractor to speed up CF trackers. In the distillation process, we propose a fidelity loss to enable the student network to maintain the representation capability of the teacher network. Meanwhile, we design a tracking loss to adapt the objective of the student network from object recognition to visual tracking. The distillation process is performed offline on multiple layers and adaptively updates the student network using a background-aware online learning scheme. The online adaptation stage exploits the background contents to improve the feature discrimination of the student network. Extensive experiments on six standard datasets demonstrate that the lightweight student network accelerates the speed of state-of-the-art deep CF trackers to real-time on a single-core CPU while maintaining almost the same tracking accuracy. Ning Wang 0020, Wengang Zhou 0001, Yibing Song, Chao Ma 0004, Houqiang Li |
IEEE Trans. Image Process. | 1 |
| 2019 | Unsupervised Deep TrackingabstractWe propose an unsupervised visual tracking method in this paper. Different from existing approaches using extensive annotated data for supervised learning, our CNN model is trained on large-scale unlabeled videos in an unsupervised manner. Our motivation is that a robust tracker should be effective in both the forward and backward predictions (i.e., the tracker can forward localize the target object in successive frames and backtrace to its initial position in the first frame). We build our framework on a Siamese correlation filter network, which is trained using unlabeled raw videos. Meanwhile, we propose a multiple-frame validation method and a cost-sensitive loss to facilitate unsupervised learning. Without bells and whistles, the proposed unsupervised tracker achieves the baseline accuracy of fully supervised trackers, which require complete and accurate labels during training. Furthermore, unsupervised framework exhibits a potential in leveraging unlabeled or weakly labeled data to further improve the tracking accuracy. Ning Wang 0020, Yibing Song, Chao Ma 0004, Wengang Zhou 0001, Wei Liu 0005, Houqiang Li |
CVPR | 1 |
| 2019 | Learning Motion-Aware Policies for Robust Visual TrackingabstractVisual object tracking aims to locate a moving target specified at the initial frame. Although this task is closely related to the temporal motion information, the motion model typically draws limited attention. In this paper, we propose a motion-aware multi-domain network for robust visual tracking. In our approach, a motion-aware agent is trained via reinforcement learning, which can infer the parameters of the particle filter in a continuous action space. Different from existing tracking-by-detection frameworks that the particle filter merely relies on the previous target state, our motion-aware agent, after receiving the current state, can adaptively change the parameters of the particle filter (e.g., particle location and scale range). As a result, our approach samples high-quality candidates for further classification/tracking, thus can better handle challenges such as fast motion and scale variation. Extensive experiments on large-scale benchmarks verify the effectiveness of our method. Liansheng Zhuang, Ning Wang 0020, Wengang Zhou 0001, Houqiang Li |
ICME | 3 |
| 2019 | Multi-tracker fusion via adaptive outlier detection
Ning Wang 0020, Wengang Zhou 0001, Weiping Li 0003, Houqiang Li |
Multim. Tools Appl. | 2 |
| 2019 | Reliable Re-Detection for Long-Term TrackingabstractIn long-term object tracking, severe occlusion and deformation could happen to the targets. Due to the accumulation and propagation of estimation errors, even a few frames of full occlusion in a video sequence could lead to the failure of the tracking. Recently, correlation filter-based trackers have received lots of attention and gained great success in real-time tracking. However, most of them ignore the reliability of the tracked results and lack an effective mechanism to refine the unreliable results. To cope with these issues, in this paper, we propose a long-term tracking framework composed of both tracking-by-detection and re-detection modules. The tracking-by-detection part is built on the discriminative correlation filter (DCF) integrated with a color-based model. The re-detection module filters a large number of detection candidates and refines the tracking results. With the proposed re-detection refinement, detected results in each frame were re-evaluated and re-detection is carried out when necessary. Besides, the reliability estimation in the re-detection module also helps adaptively update the object detector and keep it from corruption. The proposed re-detection module can be integrated into correlation filter-based trackers to consistently boost the performance. Extensive experiments on the OTB-2015, Temple-Color, and VOT-2015 benchmarks show that the proposed method performs favorably against the state-of-the-art methods while still running faster than 40 f/s. Ning Wang 0020, Wengang Zhou 0001, Houqiang Li |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Multi-Cue Correlation Filters for Robust Visual TrackingabstractIn recent years, many tracking algorithms achieve impressive performance via fusing multiple types of features, however, most of them fail to fully explore the context among the adopted multiple features and the strength of them. In this paper, we propose an efficient multi-cue analysis framework for robust visual tracking. By combining different types of features, our approach constructs multiple experts through Discriminative Correlation Filter (DCF) and each of them tracks the target independently. With the proposed robustness evaluation strategy, the suitable expert is selected for tracking in each frame. Furthermore, the divergence of multiple experts reveals the reliability of the current tracking, which is quantified to update the experts adaptively to keep them from corruption. Through the proposed multi-cue analysis, our tracker with standard DCF and deep features achieves outstanding results on several challenging benchmarks: OTB-2013, OTB-2015, Temple-Color and VOT 2016. On the other hand, when evaluated with only simple hand-crafted features, our method demonstrates comparable performance amongst complex non-realtime trackers, but exhibits much better efficiency, with a speed of 45 FPS on a CPU. Ning Wang 0020, Wengang Zhou 0001, Qi Tian 0001, Richang Hong, Meng Wang 0001, Houqiang Li |
CVPR | 1 |
| 2018 | Robust Object Tracking Via Part-Based Correlation Particle FilterabstractIn this paper, a part-based correlation particle filter framework is proposed for robust visual tracking. Through managing target parts by correlation filters in a particle filter framework, we comprehensively model the target appearance using plentiful overlapped local parts with different positions and sizes. Further, we propose a particle re-sampling mechanism with appearance and geometry reliability consideration to resam-ple the redundant particles, which guides our tracker to focus more on the discriminative and reliable local parts. Finally, to cope with the limited search range of local tracker and model corruption caused by unreliable samples, we introduce the top-down coarse-to-fine localization and bottom-up adaptive update strategies to further boost the performance. Extensive experimental results on three challenging datasets demonstrate that our tracking algorithm performs favorably against state-of-the-art methods. Specifically, our approach exhibits superior performance on tracking nonrigid objects with rotation and large deformation. Ning Wang 0020, Wengang Zhou 0001, Houqiang Li |
ICME | 1 |