Nana Fan

dblp:146/8400 · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
10since 2021 · last 2023
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2023 Siamese residual network for efficient visual tracking
Nana Fan, Qiao Liu 0001, Xin Li 0034, Zikun Zhou, Zhenyu He 0001
Inf. Sci.1
2023 Learning Dual-Level Deep Representation for Thermal Infrared Tracking
abstract
The feature models used by existing Thermal InfraRed (TIR) tracking methods are usually learned from RGB images due to the lack of a large-scale TIR image training dataset. However, these feature models are less effective in representing TIR objects and they are difficult to effectively distinguish distractors because they do not contain fine-grained discriminative information. To this end, we propose a dual-level feature model containing the TIR-specific discriminative feature and fine-grained correlation feature for robust TIR object tracking. Specifically, to distinguish inter-class TIR objects, we first design an auxiliary multi-classification network to learn the TIR-specific discriminative feature. Then, to recognize intra-class TIR objects, we propose a fine-grained aware module to learn the fine-grained correlation feature. These two kinds of features complement each other and represent TIR objects in the levels of inter-class and intra-class respectively. These two feature models are constructed using a multi-task matching framework and are jointly optimized on the TIR object tracking task. In addition, we develop a large-scale TIR image dataset to train the network for learning TIR-specific feature patterns. To the best of our knowledge, this is the largest TIR tracking training dataset with the richest object class and scenario. To verify the effectiveness of the proposed dual-level feature model, we propose an offline TIR tracker (MMNet) and an online TIR tracker (ECO-MM) based on the feature model and evaluate them on three TIR tracking benchmarks. Extensive experimental results on these benchmarks demonstrate that the proposed algorithms perform favorably against the state-of-the-art methods.
Qiao Liu 0001, Di Yuan 0002, Nana Fan, Peng Gao 0005, Xin Li 0034, Zhenyu He 0001
IEEE Trans. Multim.3
2022 Accurate bounding-box regression with distance-IoU loss for visual tracking
Di Yuan 0002, Xiu Shu, Nana Fan, Xiaojun Chang, Qiao Liu 0001, Zhenyu He 0001
J. Vis. Commun. Image Represent.3
2022 Noise-Suppressing Deep Tracking
abstract
In visual tracking, it is challenging to distinguish the target from similar objects called noises in the background. As deep trackers use convolutional neural networks for image classification as feature extractors, the extracted features are insensitive to different instances in the same class, which is prone to make prediction models confuse the target and the similar noises in the background. To this end, we propose a noise-suppressing algorithm to learn the discriminative representation for distinguishing the target from the noises in the background. First, we learn polynomial kernels for a search patch under the semantic guidance to increase the difference between representations of the target and the noises in the background. Second, we formulate the online foreground-background functions for the target and the noises in the background to learn an adaptive kernel, which suppresses the features positive for the noises and promotes the features positive for the target. We evaluate the proposed method on seven public datasets including OTB-2013, OTB-2015, VOT-2018, LaSOT, TrackingNet, GOT10k, and NFS. The comprehensive experimental results show that the proposed algorithm performs favorably against state-of-the-art methods, while running at real-time speed.
Nana Fan, Xin Li 0034, Zikun Zhou, Qiao Liu 0001, Zhenyu He 0001
IEEE Trans. Circuits Syst. Video Technol.1
2022 TCDesc: Learning Topology Consistent Descriptors for Image Matching
abstract
The triplet loss is widely used in learning the local descriptors for image matching. However, existing triplet loss-based methods, like HardNet and DSM, employ the point-to-point distance metric, which neglects the neighborhood information of descriptors. Considering the fact that local neighborhood structures of matching descriptors would be similar under the ideal condition, this paper aims to learn the neighborhood topology-consistent descriptors (TCDesc). To this end, we first propose the linear combination weight as the topology weight to depict the neighborhood topology for each descriptor, where the difference between the center descriptor and the linear combination of its neighbors is minimized. For the global comparison, we then define a global topology vector by using the local topology weights. Next, beyond the Euclidean distance, we define a topology distance with the topology vectors to indicate the topological difference between the matching descriptors. Furthermore, we propose an adaptive weighting strategy to jointly minimize the topology distance and Euclidean distance in triplet loss. Experimental results on four widely-used datasets, i.e., UBC PhotoTourism, HPatches, W1BS and Oxford, demonstrate that our method can effectively improve the performance of both HardNet and DSM.
Honghu Pan, Yongyong Chen, Zhenyu He 0001, Fanyang Meng, Nana Fan
IEEE Trans. Circuits Syst. Video Technol.5
2022 Target-Aware State Estimation for Visual Tracking
abstract
Trackers based on the IoU prediction network (IoU-Net) have shown superior performance, which refines a coarse bounding box to an accurate one by maximizing the IoU between the target and the coarse box. However, the traditional IoU-Net is less effective in exploiting the limited but crucial supervision information contained in the initial frame, including the discriminative information between the target and backgrounds and the structure information of the initial target. Missing such information makes the IoU-Net less robust to background distractors and diverse variations of the target appearance. To address this issue, we propose a target-aware state estimation network for visual tracking. A gradient-guided feature adjustment module is built on an online discriminative model to generate target-aware features for constructing the state estimation network; it conveys the online learned discriminative information into the offline trained state estimation network. In addition, we propose a structure-aware integration module and embed it into the state estimation network, enabling the tracker to explicitly model the structure information of the initial target. Extensive experimental results on the VOT2018, OTB2015, UAV123, NFS30, TC128, TrackingNet, LaSOT, and VOT2018-LT datasets demonstrate that the proposed approach performs favorably against state-of-the-art trackers.
Zikun Zhou, Xin Li 0034, Nana Fan, Hongpeng Wang 0002, Zhenyu He 0001
IEEE Trans. Circuits Syst. Video Technol.3
2021 Interactive convolutional learning for visual tracking
Nana Fan, Qiao Liu 0001, Xin Li 0034, Zikun Zhou, Zhenyu He 0001
Knowl. Based Syst.1
2021 Learning dual-margin model for visual tracking
Nana Fan, Xin Li 0034, Zikun Zhou, Qiao Liu 0001, Zhenyu He 0001
Neural Networks1
2021 Adaptive ensemble perception tracking
Zikun Zhou, Nana Fan, Kai Yang 0018, Hongpeng Wang 0002, Zhenyu He 0001
Neural Networks2
2021 Learning Deep Multi-Level Similarity for Thermal Infrared Object Tracking
abstract
Existing deep Thermal InfraRed (TIR) trackers only use semantic features to represent the TIR object, which lack the sufficient discriminative capacity for handling distractors. This becomes worse when the feature extraction network is only trained on RGB images. To address this issue, we propose a multi-level similarity model under a Siamese framework for robust TIR object tracking. Specifically, we compute different pattern similarities using the proposed multi-level similarity network. One of them focuses on the global semantic similarity and the other computes the local structural similarity of the TIR object. These two similarities complement each other and hence enhance the discriminative capacity of the network for handling distractors. In addition, we design a simple while effective relative entropy based ensemble subnetwork to integrate the semantic and structural similarities. This subnetwork can adaptive learn the weights of the semantic and structural similarities at the training stage. To further enhance the discriminative capacity of the tracker, we propose a large-scale TIR video sequence dataset for training the proposed model. To the best of our knowledge, this is the first and the largest TIR object tracking training dataset to date. The proposed TIR dataset not only benefits the training for TIR object tracking but also can be applied to numerous TIR visual tasks. Extensive experimental results on three benchmarks demonstrate that the proposed algorithm performs favorably against the state-of-the-art methods.
Qiao Liu 0001, Xin Li 0034, Zhenyu He 0001, Nana Fan, Di Yuan 0002, Hongpeng Wang 0002
IEEE Trans. Multim.4
2020 Multi-Task Driven Feature Models for Thermal Infrared Tracking
abstract
Existing deep Thermal InfraRed (TIR) trackers usually use the feature models of RGB trackers for representation. However, these feature models learned on RGB images are neither effective in representing TIR objects nor taking fine-grained TIR information into consideration. To this end, we develop a multi-task framework to learn the TIR-specific discriminative features and fine-grained correlation features for TIR tracking. Specifically, we first use an auxiliary classification network to guide the generation of TIR-specific discriminative features for distinguishing the TIR objects belonging to different classes. Second, we design a fine-grained aware module to capture more subtle information for distinguishing the TIR objects belonging to the same class. These two kinds of features complement each other and recognize TIR objects in the levels of inter-class and intra-class respectively. These two feature models are learned using a multi-task matching framework and are jointly optimized on the TIR tracking task. In addition, we develop a large-scale TIR training dataset to train the network for adapting the model to the TIR domain. Extensive experimental results on three benchmarks show that the proposed algorithm achieves a relative gain of 10% over the baseline and performs favorably against the state-of-the-art methods. Codes and the proposed TIR dataset are available at https://github.com/QiaoLiuHit/MMNet.
Qiao Liu 0001, Xin Li 0034, Zhenyu He 0001, Nana Fan, Di Yuan 0002, Wei Liu 0065, Yongsheng Liang 0001
AAAI4
2020 LSOTB-TIR: A Large-Scale High-Diversity Thermal Infrared Object Tracking Benchmark
abstract
In this paper, we present a Large-Scale and high-diversity general Thermal InfraRed (TIR) Object Tracking Benchmark, called LSOTB-TIR, which consists of an evaluation dataset and a training dataset with a total of 1,400 TIR sequences and more than 600K frames. We annotate the bounding box of objects in every frame of all sequences and generate over 730K bounding boxes in total. To the best of our knowledge, LSOTB-TIR is the largest and most diverse TIR object tracking benchmark to date. To evaluate a tracker on different attributes, we define 4 scenario attributes and 12 challenge attributes in the evaluation dataset. By releasing LSOTB-TIR, we encourage the community to develop deep learning based TIR trackers and evaluate them fairly and comprehensively. We evaluate and analyze more than 30 trackers on LSOTB-TIR to provide a series of baselines, and the results show that deep trackers achieve promising performance. Furthermore, we re-train several representative deep trackers on LSOTB-TIR, and their results demonstrate that the proposed training dataset significantly improves the performance of deep TIR trackers. Codes and dataset are available at https://github.com/QiaoLiuHit/LSOTB-TIR.
Qiao Liu 0001, Xin Li 0034, Zhenyu He 0001, Chenglong Li 0002, Zikun Zhou, Di Yuan 0002, Jing Li 0071, Kai Yang 0018, Nana Fan, Feng Zheng 0001
ACM Multimedia10
2020 SiamAtt: Siamese attention network for visual tracking
Kai Yang 0018, Zhenyu He 0001, Zikun Zhou, Nana Fan
Knowl. Based Syst.4
2020 Learning target-focusing convolutional regression model for visual object tracking
Di Yuan 0002, Nana Fan, Zhenyu He 0001
Knowl. Based Syst.2
2020 Dual-regression model for visual tracking
Xin Li 0034, Qiao Liu 0001, Nana Fan, Zikun Zhou, Zhenyu He 0001, Xiaoyuan Jing
Neural Networks3
2019 Region-filtering correlation tracking
Nana Fan, Jing Li 0071, Zhenyu He 0001, Chunkai Zhang, Xin Li 0034
Knowl. Based Syst.1
2019 Hierarchical spatial-aware Siamese network for thermal infrared object tracking
Xin Li 0034, Qiao Liu 0001, Nana Fan, Zhenyu He 0001, Hongzhi Wang 0001
Knowl. Based Syst.3
2014 Finding Vacant Taxis Using Large Scale GPS Traces
Hongyan Li 0002, Shenda Hong, Yiyong Lin 0003, Nana Fan, Gaoyan Ou, Tengjiao Wang 0003, Lilue Fan
WAIM5