VLDB 2026 Research / reviewers in the wild / expert
Zhu Teng
dblp:132/2247
· DBLP profile ↗
37ranked-venue papers
6as first author
29since 2021 · last 2026
0000-0002-1754-4878ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 3 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 3 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring representation learning from vision-language models for open-vocabulary HOI detection
Chunyan Chen, Zhu Teng, Baopeng Zhang |
Comput. Vis. Image Underst. | 2 |
| 2026 | Non-target information also matters: InverseFormer tracker for single object tracking
Qiuhang Gu, Baopeng Zhang, Zhu Teng |
Image Vis. Comput. | 3 |
| 2026 | Affinity-aware uncertainty quantification for learning with noisy labels
Rui Li 0059, Wenjie Ai, Zhu Teng, Baopeng Zhang, Junwei Du |
Pattern Recognit. | 5 |
| 2026 | Boosting Few-Shot Human-Object Recognition With Hierarchical Correlation LearningabstractIdentifying novel human-object interaction (HOI) classes with scarce data is a challenging and crucial task in computer vision. Existing methods mainly use coarse global visual information to build class prototypes in meta-learning. Despite their promising results, these methods often fail to capture fine-grained interaction semantics and effectively learn from data with low inter-class variance, leading to suboptimal performance in distinguishing similar categories. To overcome these issues, we propose a new model called hierarchical relation network for few-shot HOI recognition (FS-HOI). This model integrates multi-level interaction clues, spanning from coarse to fine-grained, to enhance HOI features. It employs a unified graph network to capture intra- and inter-relationships among human parts with contextual information, augmented by language-guided attention for semantic mining within each interactive sub-graph. In contrast to conventional methods that depend on global class prototype comparisons, our approach advances metric learning by integrating contrastive mechanisms, utilizing rich instance pairs as comparative references to effectively address inter-class variance. Furthermore, a graph relation network leverages prior knowledge of unknown HOIs, embedding task-specific features into contrastive instances. Our method establishes a new state-of-the-art on three few-shot HOI datasets, with substantial performance gains and ablation studies confirming the efficacy of each component. Jiale Yu, Baopeng Zhang, Zhu Teng, Jianping Fan 0007 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Less Attention is More: Prompt Transformer for Generalized Category DiscoveryabstractGeneralized Category Discovery (GCD) typically relies on the pre-trained Vision Transformer (ViT) to extract features from a global receptive field, followed by contrastive learning to simultaneously classify unlabeled known classes and unknown classes without priors. Owing to the deficiency in the modeling capacity for inner-patch local information within ViT, current methods primarily focus on discriminative features at global level. This results in a model with more yet scattered attention, where neither excessive nor insufficient focus can grasp subtle differences to classify fine-grained unknown and known categories. To address this issue, we propose the AptGCD to deliver apt attention for GCD. It mimics the human brain how leveraging visual perception to refine local attention and comprehend global context by proposing a Meta Visual Prompt (MVP) and Prompt Transformer (PT). MVP is introduced into GCD for the first time, refining channel-level attention, while adaptively self-learning unique inner-patch features as prompts to achieve local visual modeling for our prompt transformer. Yet, relying solely on detailed features can lead to skewed judgments. Hence, PT harmonizes local and global representations, guiding the model's interpretation of features through broader contexts, thereby capturing more useful details with less attention. Extensive experiments on seven datasets demonstrate that AptGCD outperforms current methods, it achieves an average proportional ‘New’ accuracy improvement of approximately 9.2% over SOTA method on the all four fine-grained datasets, establishing a new standard in the field. The code is available at https://github.com/wendy26zhang/AptGCD. Wei Zhang 0016, Baopeng Zhang, Zhu Teng, Wenxin Luo, Junnan Zou, Jianping Fan 0007 |
CVPR | 3 |
| 2025 | Open-Unfairness Adversarial Mitigation for Generalized Deepfake Detection
Zhaoyang Li 0014, Zhu Teng, Baopeng Zhang, Jianping Fan 0001 |
ICCV | 2 |
| 2025 | OV-DAVEL: Towards Open-Vocabulary Dense Audio-Visual Event Localization in Untrimmed VideosabstractThe Dense Audio-Visual Event Localization (DAVEL) task aims to temporally localize all audio-visual events within untrimmed videos. Existing methods typically operate under a closed-set assumption, which limits their ability to generalize to test videos containing previously unseen event categories-an essential capability for open-world scenarios. To this end, we propose a novel task setting, Open-Vocabulary Dense Audio-Visual Event Localization (OV-DAVEL), along with a one-stage method Open-DAVTR designed to detect events that were not observed during training. Open-DAVTR consists of two core components: a class-agnostic foreground-aware generator and a multi-modal semantic-aware classifier. Specifically, the generator is a Detection Transformer-based module that produces event proposals while adaptively attending to discriminative foreground snippets for downstream classification. The classifier leverages rich temporal representations and context-aware textual semantics to effectively recognize events, regardless of whether they were seen during training. In addition, we establish comprehensive OV-DAVEL benchmarks across various settings. Experimental results show that our model significantly outperforms existing baselines in detecting both seen and unseen events, highlighting its effectiveness in open-vocabulary event localization. Code and data are available at: https://github.com/yujialele/OV-DEAVEL. Jiale Yu, Baopeng Zhang, Zhu Teng, Jianping Fan 0007 |
ACM Multimedia | 3 |
| 2025 | Semantic consistency learning for unsupervised multi-modal person re-identification
Zhu Teng, Baopeng Zhang |
Image Vis. Comput. | 2 |
| 2025 | Leveraging facial landmarks improves generalization ability for deepfake detection
Baopeng Zhang, Jianghao Wu 0004, Wenxin Luo, Zhu Teng, Jianping Fan 0007 |
Pattern Recognit. | 5 |
| 2025 | Bi-Stream Coteaching Network for Weakly-Supervised Deepfake Localization in VideosabstractWith the rapid evolution of deepfake technologies, attackers can arbitrarily alter the intended message of a video by modifying just a few frames. To this extent, simplistic binary judgments of entire videos increasingly seem less convincing and interpretable. Although numerous efforts have been made to develop fine-grained interpretations, these typically depend on elaborate annotations, which are both costly and challenging to obtain in real-world scenarios. To push the related frontier research, we introduce a novel task called Weakly-Supervised Deepfake Localization (WSDL), which aims to identify manipulated frames only with cushy video-level labels. Meanwhile, we propose a new framework named Bi-stream coteaching Deepfake Localization (CoDL), which advances the WSDL task through a progressive mutual refinement strategy across complementary spatial and temporal modalities. The CoDL framework incorporates an inconsistency perception module that discerns subtle forgeries by assessing spatial and temporal incoherence, and a prototype-based enhancement module that mitigates frame noise and amplifies discrepancies to create a robust feature space. Additionally, a progressive coteaching mechanism is implemented to facilitate the exchange of valuable knowledge between modalities, enhancing the detection of subtle frame-level forgery features and thereby improving the model’s generalization capabilities. Extensive experiments are conducted to demonstrate the superiority of our approach, particularly achieving an impressive 8.83% improvement in AUC on highly compressed datasets when learning from weak supervision. Zhaoyang Li 0014, Zhu Teng, Baopeng Zhang, Jianping Fan 0007 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Boosting Non-causal Semantic Elimination: An Unconventional Harnessing of LVM for Open-World Deepfake Interpretation
Zhaoyang Li 0014, Zhu Teng, Baopeng Zhang, Jianping Fan 0007 |
ACM Multimedia | 2 |
| 2024 | OpenAVE: Moving towards Open Set Audio-Visual Event LocalizationabstractAudio-Visual Event (AVE) Localization aims to identify and classify video segments that are both audible and visible, a field that has seen substantial progress in recent years. Existing methods operate under a closed-set assumption and struggle to recognize unknown events in open-world scenarios. To better adapt to real-life applications, we introduce the Open Set Audio-Visual Event Localization task and propose a novel and effective network called OpenAVE based on evidential deep learning. To the best of our knowledge, this is the first effort to address this challenge. Our approach encompasses deep evidential AVE classification and event-relevant prediction, targeting the nuanced demands of open-set environments. The deep evidential AVE classification manages event classification uncertainty by extracting class evidence from segment-specific representations enriched with multi-scale context. To effectively distinguish between unknown events and background segments, event-relevant prediction utilizes positive-unlabeled learning. Futhermore, a learnable Gaussian-prior prediction branch is adopted to enhance the performance of event-relevant prediction. Experimental results demonstrate that OpenAVE significantly outperforms state-of-the-art models on the Audio-Visual Event dataset, confirming the effectiveness of our proposed method. Jiale Yu, Baopeng Zhang, Zhu Teng, Jianping Fan 0007 |
ACM Multimedia | 3 |
| 2024 | MINet: Modality interaction network for unified multi-modal tracking
Shuang Gong, Zhu Teng, Rui Li 0059, Jack Fan, Baopeng Zhang, Jianping Fan 0007 |
Image Vis. Comput. | 2 |
| 2024 | HCgNet: A hierarchical context-guided network for multi-object tracking
Rui Li 0059, Baopeng Zhang, Wei Liu 0165, Zhaoyang Li 0014, Jack Fan, Zhu Teng, Jianping Fan 0001 |
Knowl. Based Syst. | 6 |
| 2024 | Token-word mixer meets object-aware transformer for referring image segmentation
Zhu Teng, Jack Fan, Baopeng Zhang, Jianping Fan 0001 |
Pattern Recognit. | 2 |
| 2023 | Learning Disentangled Representations for Environment Inference in Out-of-distribution Generalization
Dongqi Li, Zhu Teng, Ziyin Wang, Baopeng Zhang, Jianping Fan 0007 |
BMVC | 2 |
| 2023 | Heterogeneous Diversity Driven Active Learning for Multi-Object TrackingabstractThe existing one-stage multi-object tracking (MOT) algorithms have achieved satisfactory performance benefiting from a large amount of labeled data. However, acquiring plenty of laborious annotated frames is not practical in real applications. To reduce the cost of human annotations, we propose Heterogeneous Diversity driven Active Multi-Object Tracking (HD-AMOT), to infer the most informative frames for any MOT tracker by observing the heterogeneous cues of samples. HD-AMOT defines the diversified informative representation by encoding the geometric and semantic information, and formulates the frame inference strategy as a Markov decision process to learn an optimal sampling policy based on the designed informative representation. Specifically, HD-AMOT consists of a diversified informative representation module as well as an informative frame selection network. The former produces the signal characterizing the diversity and distribution of frames, and the latter receives the signal and conducts multi-frame cooperation to enable batch frame sampling. Extensive experiments conducted on the MOT15, MOT17, MOT20, and Dancetrack datasets demonstrate the efficacy and effectiveness of HD-AMOT. Experiments show that under 50% budget our HD-AMOT can achieve similar or even higher performance as fully-supervised learning. Rui Li 0059, Baopeng Zhang, Jun Liu 0036, Wei Liu 0165, Zhu Teng |
ICCV | 6 |
| 2023 | Sharpness-Aware Minimization for Out-of-Distribution Generalization
Dongqi Li, Zhu Teng, Ziyin Wang |
ICONIP (7) | 2 |
| 2023 | Hierarchical Reasoning Network with Contrastive Learning for Few-Shot Human-Object Interaction RecognitionabstractFew-shot learning (FSL) for human-object interaction aims at classifying samples of new unseen HOI classes with only a few labeled samples available. Although progress has been made in few-shot human-object interaction, most of the existing methods encounter two issues in handling fine-grained interactions: the inability to capture more subtle interactive clues and the inadequacy in learning from data with low inter-class variance. To tackle the first issue, we propose a hierarchical reasoning network to integrate multi-level interactive clues (from coarse to fine-grained) for strengthening HOI representations. The hierarchical relation module mainly captures and aggregates more discriminative relation information among human parts at multiple levels (including the human instance, action region, and body part levels) and objects via a unified graph and exploits a language-guided attentive fusion way to highlight informative features of each interaction level. To address the second issue, we introduce a contrastive learning mechanism to alleviate the inter-class variance. Compared with the previous ProtoNet-based methods, our model generates more discriminative representations for low inter-class variance data, since it makes full use of potential contrastive pairs in each training episode. Extensive experimental results on two standard benchmarks demonstrate that the proposed model performs favorably against state-of-the-art FS-HOI methods. Jiale Yu, Baopeng Zhang, Zhu Teng |
ACM Multimedia | 5 |
| 2023 | MRE-Net: Multi-Rate Excitation Network for Deepfake Video DetectionabstractThe current social media is flooded with hyper realistic face-synthetic videos due to the explosion of DeepFake technology that has brought a serious impact on human society security, which calls for further exploring on deepfake video detection methods. Existing methods attempt to isolated capture spatial artifacts or extract the homogeneous temporal inconsistency to detect deepfake video, but little attention has been paid to the exploitation of dynamic spatial-temporal inconsistency. To mitigate this issue, in this paper, we propose a novel Multi-Rate Excitation Network (MRE-Net) to effectively excite dynamic spatial-temporal inconsistency from the perspective of multiple rates for deepfake video detection. The proposed MRE-Net is composed of two components: Bipartite Group Sampling (BGS) and multiple rate branches. The BGS draws the entire video into multiple bipartite groups with different rates to cover various face motion dynamic evolution. We further design multiple rate branches to capture both short-term and long-term spatial-temporal inconsistency from corresponding bipartite groups of BGS. Concretely, for the early stages of the multi-rate branches, Momentary Inconsistency Excitation (MIE) module is developed to encode the spatial artifacts and intra-group short-term temporal inconsistency. Meanwhile, for the last stages of the multi-rate branches, Longstanding Inconsistency Excitation (LIE) module is constructed to perceive inter-group long-term temporal dynamics. Extensive experiments and visualizations conducted on four popular datasets demonstrate the effectiveness of the proposed method against state-of-the-art deepfake detection methods. Guilin Pang, Baopeng Zhang, Zhu Teng, Zige Qi, Jianping Fan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Interactive Two-Stream Network Across Modalities for Deepfake DetectionabstractAs face forgery techniques have become more mature, the proliferation of deepfakes may threaten the security of human society. Although existing deepfake detection methods achieve good performance for in-dataset evaluation, it remains to be improved in the generalization ability, where the representation of the imperceptible artifacts plays a significant role. In this paper, we propose an Interactive Two-Stream Network (ITSNet) to explore the discriminant inconsistency representation from the perspective of cross-modality. In particular, the patch-wise Decomposable Discrete Cosine Transform (DDCT) is adopted to extract fine-grained high-frequency clues, and information from different modalities communicates with each other via a designed interaction module. To perceive the temporal inconsistency, we first develop a Short-term Embedding Module (SEM) to refine subtle local inconsistency representation between adjacent frames, and then a Long-term Embedding Module (LEM) is designed to further refine the erratic temporal inconsistency representation from the long-range perspective. Extensive experimental results conducted on three public datasets show that ITSNet outperforms the state-of-the-art methods both in terms of in-dataset and cross-dataset evaluations. Jianghao Wu 0004, Baopeng Zhang, Zhaoyang Li 0014, Guilin Pang, Zhu Teng, Jianping Fan 0007 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Inference-Domain Network Evolution: A New Perspective for One-Shot Multi-Object TrackingabstractThe supervised one-shot multi-object tracking (MOT) algorithms have achieved satisfactory performance benefiting from a large amount of labeled data. However, in real applications, acquiring plenty of laborious manual annotations is not practical. It is necessary to adapt the one-shot MOT model trained on a labeled domain to an unlabeled domain, yet such domain adaptation is a challenging problem. The main reason is that it has to detect and associate multiple moving objects distributed in various spatial locations, but there are obvious discrepancies in style, object identity, quantity, and scale among different domains. Motivated by this, we propose a novel inference-domain network evolution to enhance the generalization ability of the one-shot MOT model. Specifically, we design a spatial topology-based one-shot network (STONet) to perform the one-shot MOT task, where a self-supervision mechanism is employed to stimulate the feature extractor to learn the spatial contexts without any annotated information. Furthermore, a temporal identity aggregation (TIA) module is proposed to assist STONet to weaken the adverse effects of noisy labels in the network evolution. This designed TIA aggregates historical embeddings with the same identity to learn cleaner and more reliable pseudo labels. In the inference domain, the proposed STONet with TIA performs pseudo label collection and parameter update progressively to realize the network evolution from the labeled source domain to an unlabeled inference domain. Extensive experiments and ablation studies conducted on MOT15, MOT17, and MOT20, demonstrate the effectiveness of our proposed model. Rui Li 0059, Baopeng Zhang, Jun Liu 0036, Wei Liu 0165, Zhu Teng |
IEEE Trans. Image Process. | 5 |
| 2023 | PANet: An End-to-end Network Based on Relative Motion for Online Multi-object TrackingabstractThe popular tracking-by-detection paradigm of multi-object tracking (MOT) takes detections of each frame as the input and associates detections from one frame to another. Existing association methods based on the relative motion have attracted attention, because they restrain the effect of noisy detections and improve the performance of MOT. However, these methods depend only on the immediately previous frame, which may easily lead to inaccurate matches and even large accumulated errors. Furthermore, multiple objects involved in occlusions are not fully exploited in these existing methods, which leads to the aggravation of inaccurate matches. Motivated by these issues, we design the pivot to represent each object and propose a novel pivot association network (PANet) for the MOT task. Specifically, pivots are learned from spatial semantic and historical contextual clues, which alleviates the dependency on the immediately previous frame. Our online tracker PANet employs pivots and a lightweight associator to localize tracklets of objects, which can inhibit noise detections and improve the accuracy of tracklet prediction by learning the correlation responses between pivots and spatial search areas. Extensive experiments conducted on two-dimensional MOT15, MOT16, MOT17, and MOT20 demonstrate the effectiveness of the proposed method against numerous state-of-the-art MOT trackers. Rui Li 0059, Baopeng Zhang, Wei Liu 0165, Zhu Teng, Jianping Fan 0007 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | An end-to-end identity association network based on geometry refinement for multi-object tracking
Rui Li 0059, Baopeng Zhang, Zhu Teng, Jianping Fan 0001 |
Pattern Recognit. | 3 |
| 2022 | Global to Local: Clip-LSTM-Based Object Detection From Remote Sensing ImagesabstractObject detection from remote sensing images (RSIs) is a basic topic in the area of aviation and satellite image processing, which has great effects on geological disaster detection, agricultural planning, and land utilization. However, it is always faced with several severe difficulties. For instance, the scale of the target spans over a very wide range, and the difference between the target size and the image size is huge, as some targets only account for a dozen pixels compared with the RSI of the megapixel level. In this work, an innovative object detection network (GLNet) is proposed for remote sensing imagery. Our approach integrates global context clues extracted by the multiscale perception (MSP) module and local spatial contextual correlations encoded by Clip long short-term memory (Clip-LSTM). These rich semantic features are further exploited to design a self-adapted anchor (SA) module to alleviate the scale variations in RSIs. Extensive experiments are executed on several public easily accessed benchmarks, including DOTA, NWPU VHR-10, and DIOR. Experimental results have demonstrated that our GLNet outperforms numerous latest methods. Zhu Teng, Yani Duan, Baopeng Zhang, Jianping Fan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Fast-HBNet: Hybrid Branch Network for Fast Lane DetectionabstractAs one of the fundamental visual tasks in the unmanned driving area, lane detection attracts increasing attention. In practical applications, lane points are very difficult to localize because they usually appear to be sparse and incomplete due to the influence of illumination and environment. Conventional lane detection methods rely on coarse features and carefully designed postprocessing to detect the lane lines. However, these methods are usually slow, and the stability and generalization ability are unsatisfactory. In this paper, we propose a novel lane detection network – Fast-HBNet(Fast-Hybrid Branch Network), which exploits both global semantic information and spatial contexts. To enlarge receptive fields and encode more detailed information, the compound transformation is employed and the proposed hybrid branch network extracts four diverse feature maps with different receptive fields and spatial contexts. Besides, we design a Hierarchical Feature Learning (HFL) module to learn lane features from the scale, channel, and spatial levels to enhance the generalization ability of our detector. These features are further selectively coalesced to generate unified lane feature maps with large receptive fields and rich detailed information. In other words, our network can encode the global semantic information from the high-resolution feature maps and the fine-grained details in the low-resolution feature maps. Experimental results conducted on the TuSimple (2017) and CULane [Panet al.(2018)] datasets demonstrate that the proposed Fast-HBNet outperforms numerous state-of-the-art lane detectors in both speed and accuracy. Particularly, Fast-HBNet achieves an accuracy of 96.88% on the TuSimple dataset at a speed of 76 FPS. Guilin Pang, Baopeng Zhang, Zhu Teng, Nan Ma 0012, Jianping Fan 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Self-supervised Network Evolution for Few-shot ClassificationabstractFew-shot classification aims to recognize new classes by learning reliable models from very few available samples. It could be very challenging when there is no intersection between the alreadyknown classes (base set) and the novel set (new classes). To alleviate this problem, we propose to evolve the network (for the base set) via label propagation and self-supervision to shrink the distribution difference between the base set and the novel set. Our network evolution approach transfers the latent distribution from the already-known classes to the unknown (novel) classes by: (a) label propagation of the novel/new classes (novel set); and (b) design of dual-task to exploit a discriminative representation to effectively diminish the overfitting on the base set and enhance the generalization ability on the novel set. We conduct comprehensive experiments to examine our network evolution approach against numerous state-of-the-art ones, especially in a higher way setup and cross-dataset scenarios. Notably, our approach outperforms the second best state-of-the-art method by a large margin of 3.25% for one-shot evaluation over miniImageNet. Xuwen Tang, Zhu Teng, Baopeng Zhang, Jianping Fan 0001 |
IJCAI | 2 |
| 2021 | A divide-and-unite deep network for person re-identification
Rui Li 0059, Baopeng Zhang, Zhu Teng, Jianping Fan 0001 |
Appl. Intell. | 3 |
| 2021 | A Camera Identity-guided Distribution Consistency Method for Unsupervised Multi-target Domain Person Re-identificationabstractUnsupervised domain adaptation (UDA) for person re-identification (re-ID) is a challenging task due to large variations in human classes, illuminations, camera views, and so on. Currently, existing UDA methods focus on two-domain adaptation and are generally trained on one labeled source set and adapted on the other unlabeled target set. In this article, we put forward a new issue on person re-ID, namely, unsupervised multi-target domain adaptation (UMDA). It involves one labeled source set and multiple unlabeled target sets, which is more reasonable for practical real-world applications. Enabling UMDA has to learn the consistency for multiple domains, which is significantly different from the UDA problem. To ensure distribution consistency and learn the discriminative embedding, we further propose the Camera Identity-guided Distribution Consistency method that performs an alignment operation for multiple domains. The camera identities are encoded into the image semantic information to facilitate the adaptation of features. According to our knowledge, this is the first attempt on the unsupervised multi-target domain adaptation learning. Extensive experiments are executed on Market-1501, DukeMTMC-reID, MSMT17, PersonX, and CUHK03, and our method has achieved very competitive re-ID accuracy in multi-target domains against numerous state-of-the-art methods. Jiajie Tian, Qihao Tang, Rui Li 0059, Zhu Teng, Baopeng Zhang, Jianping Fan 0001 |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2020 | Integrating Long-Short Term Network for Efficient Video Object Segmentation
Zhu Teng, Baopeng Zhang, Jianping Fan 0001 |
BMVC | 2 |
| 2020 | An End-to-End Scalable Object Detection Network for Remote Sensing ImagesabstractObject detection on remote sensing images plays an important role in image analysis and understanding. The generation of anchors is a significant factor due to the diverse variations in object size and clustered small objects in remote sensing images. In this paper, we focus on this issue, and propose an anchor generator with continuously variable scales leading by semantic information and build a Scalable Object Detection Network (SODNet) to detect remote sensing targets. The proposed network exploits semantic information to guide the generation of anchors and improves the target scale variation to adapt to different remote sensing images. We execute extensive experiments on widely employed NWPU VHR-10 dataset and DOTA dataset against numerous state-of-the-art methods, and the experimental results demonstrate the performance of the proposed method, especially on small targets. Yani Duan, Zhu Teng, Baopeng Zhang, Jianping Fan 0001 |
IGARSS | 2 |
| 2020 | Three-step action search networks with deep Q-learning for real-time object tracking
Zhu Teng, Baopeng Zhang, Jianping Fan 0001 |
Pattern Recognit. | 1 |
| 2020 | Deep Spatial and Temporal Network for Robust Visual Object TrackingabstractThere are two key components that can be leveraged for visual tracking: (a) object appearances; and (b) object motions. Many existing techniques have recently employed deep learning to enhance visual tracking due to its superior representation power and strong learning ability, where most of them employed object appearances but few of them exploited object motions. In this work, a deep spatial and temporal network (DSTN) is developed for visual tracking by explicitly exploiting both the object representations from each frame and their dynamics along multiple frames in a video, such that it can seamlessly integrate the object appearances with their motions to produce compact object appearances and capture their temporal variations effectively. Our DSTN method, which is deployed into a tracking pipeline in a coarse-to-fine form, can perceive the subtle differences on spatial and temporal variations of the target (object being tracked), and thus it benefits from both off-line training and online fine-tuning. We have also conducted our experiments over four largest tracking benchmarks, including OTB-2013, OTB-2015, VOT2015, and VOT2017, and our experimental results have demonstrated that our DSTN method can achieve competitive performance as compared with the state-of-the-art techniques. The source code, trained models, and all the experimental results of this work will be made public available to facilitate further studies on this problem. Zhu Teng, Junliang Xing, Qiang Wang 0051, Baopeng Zhang, Jianping Fan 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Learning Attentions: Residual Attentional Siamese Network for High Performance Online Visual TrackingabstractOffline training for object tracking has recently shown great potentials in balancing tracking accuracy and speed. However, it is still difficult to adapt an offline trained model to a target tracked online. This work presents a Residual Attentional Siamese Network (RASNet) for high performance object tracking. The RASNet model reformulates the correlation filter within a Siamese tracking framework, and introduces different kinds of the attention mechanisms to adapt the model without updating the model online. In particular, by exploiting the offline trained general attention, the target adapted residual attention, and the channel favored feature attention, the RASNet not only mitigates the over-fitting problem in deep network training, but also enhances its discriminative capacity and adaptability due to the separation of representation learning and discriminator learning. The proposed deep architecture is trained from end to end and takes full advantage of the rich spatial temporal information to achieve robust visual tracking. Experimental results on two latest benchmarks, OTB-2015 and VOT2017, show that the RASNet tracker has the state-of-the-art tracking accuracy while runs at more than 80 frames per second. Qiang Wang 0051, Zhu Teng, Junliang Xing, Weiming Hu 0004, Stephen J. Maybank |
CVPR | 2 |
| 2017 | Robust Object Tracking Based on Temporal and Spatial Deep NetworksabstractRecently deep neural networks have been widely employed to deal with the visual tracking problem. In this work, we present a new deep architecture which incorporates the temporal and spatial information to boost the tracking performance. Our deep architecture contains three networks, a Feature Net, a Temporal Net, and a Spatial Net. The Feature Net extracts general feature representations of the target. With these feature representations, the Temporal Net encodes the trajectory of the target and directly learns temporal correspondences to estimate the object state from a global perspective. Based on the learning results of the Temporal Net, the Spatial Net further refines the object tracking state using local spatial object information. Extensive experiments on four of the largest tracking benchmarks, including VOT2014, VOT2016, OTB50, and OTB100, demonstrate competing performance of the proposed tracker over a number of state-of-the-art algorithms. Zhu Teng, Junliang Xing, Qiang Wang 0051, Congyan Lang, Songhe Feng, Yi Jin 0001 |
ICCV | 1 |
| 2016 | From sample selection to model update: A robust online visual tracking algorithm against drifting
Zhu Teng, Tao Wang 0011, Feng Liu 0061, Dong-Joong Kang, Congyan Lang, Songhe Feng |
Neurocomputing | 1 |
| 2016 | Visual railway detection by superpixel based intracellular decisions
Zhu Teng, Feng Liu 0061, Baopeng Zhang |
Multim. Tools Appl. | 1 |