EDBT 2026 Demo / reviewers in the wild / expert
Jianlin Zhang 0001
dblp:37/1545-1
· DBLP profile ↗
29ranked-venue papers
0as first author
28since 2021 · last 2026
0000-0002-5284-2942ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tracking the Unstable: Appearance-Guided Motion Modeling for Robust Multi-Object Tracking in UAV-Captured VideosabstractMulti-object tracking (MOT) aims to track multiple objects while maintaining consistent identities across frames of a given video. In unmanned aerial vehicle (UAV) recorded videos, frequent viewpoint changes and complex UAV-ground relative motion dynamics pose significant challenges, which often lead to unstable affinity measurement and ambiguous association. Existing methods typically model motion and appearance cues separately, overlooking their spatio-temporal interplay and resulting in suboptimal tracking performance. In this work, we propose AMOT, which jointly exploits appearance and motion cues through two key components: an Appearance-Motion Consistency (AMC) matrix and a Motion-aware Track Continuation (MTC) module. Specifically, the AMC matrix computes bi-directional spatial consistency under the guidance of appearance features, enabling more reliable and context-aware identity association. The MTC module complements AMC by reactivating unmatched tracks through appearance-guided predictions that align with Kalman-based predictions, thereby reducing broken trajectories caused by missed detections. Extensive experiments on three UAV benchmarks, including VisDrone2019, UAVDT, and VT-MOT-UAV, demonstrate that our AMOT outperforms current state-of-the-art methods and generalizes well in a plug-and-play and training-free manner. Jianbo Ma 0001, Hui Luo 0002, Qi Chen 0014, Yuankai Qi, Yumei Sun, Amin Beheshti, Jianlin Zhang 0001, Ming-Hsuan Yang 0001 |
AAAI | 7 |
| 2026 | RaC-Stitch: Reliability-Aware Collaborative Smoothing for Unsupervised Video StitchingabstractRecent advances in deep learning have propelled unsupervised video stitching by integrating spatial alignment and temporal stability into a unified framework. Despite this progress, state-of-the-art methods still struggle in extreme scenarios, such as rapid camera motion in low-light environments, often exhibiting severe stitching artifacts and structural distortions. These phenomena expose the inherent defects of existing frameworks when dealing with complex disturbances. More precisely, we attribute these defects to two fundamental limitations: indiscriminate feature utilization and inadequate inter-view interactions. To address these challenges, we propose RaC-Stitch, an unsupervised video stitching architecture. First, we construct a generic Reliability Prior Module (RPM). Spatially, it integrates attention to target rigid textures by suppressing interference; temporally, it filters outliers to enhance trajectory robustness. Second, we design a Cross-view Collaborative Smoothing Network. Leveraging spatiotemporal cross-attention, it facilitates view interaction to suppress distortion peaks and preserve topology. Experiments on the StabStitch-D benchmark demonstrate that RaC-Stitch significantly outperforms existing baselines. Quantitatively, our method reduces average geometric distortion by 21.0%, while simultaneously improving video stability by 2.6% and alignment accuracy by 1.0%. The code are available at https://github.com/EX-Iris/RaC-Stitch. Zhengwei Miao, Jianlin Zhang 0001, Haorui Zuo |
ICMR | 3 |
| 2026 | Revisiting color-event based tracking: A unified network, dataset, and metric
Chuanming Tang, Xiao Wang 0014, Ju Huang, Bo Jiang 0002, Lin Zhu 0012, Shifeng Chen, Jianlin Zhang 0001, Yaowei Wang 0001, Yonghong Tian 0001 |
Pattern Recognit. | 7 |
| 2026 | Multi-Stage Cross-Modality Feature Interaction for RGB-Thermal Multi-Object TrackingabstractRGB-Thermal multi-object tracking (RGB-T MOT) focuses on tracking multiple objects in complex scenarios, such as nighttime and low-light conditions, which is crucial for various applications, including video surveillance and unmanned aerial vehicle monitoring. Existing RGB-T MOT studies typically merge multi-source features in a single stage before the backbone network. These works fail to preserve fine-grained details and struggle with context-dependent fusion, resulting in suboptimal tracking performance. In this paper, we propose a novel multi-stage cross-modality spatial-temporal feature interaction network (MCTrack), which emphasizes two main aspects: temporal-aware learning and cross-modality information interaction. For temporal-aware learning, we introduce the temporal salient feature interaction (TSFI) module, which ensures trajectory continuity by capturing dynamic spatial variation of objects across RGB-T video pairs, enabling consistent object tracking over time. For cross-modality information interaction, we propose the bidirectional modality interaction (Bi-MI) module, which employs cross-modality transformers to extract complementary features from both RGB and thermal modalities at multiple stages of feature extraction, thereby improving tracking adaptability in diverse tracking scenarios. Additionally, we propose the cross-modality complementary mask (CCM) strategy, which applies non-overlapping random masks to feature maps to improve the robustness of cross-modality feature interaction. Extensive experiments and comparative analyses demonstrate that our MCTrack surpasses state-of-the-art trackers on both publicly available VT-MOT and UniRTL-MOT RGB-T datasets. The code is available at https://github.com/ydhcg-BoBo/RGB-T-MOT. Jianbo Ma 0001, Hui Luo 0002, Shuaicheng Niu, Peilin Zhao, Jianlin Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Adaptive representation-aligned modeling for visual tracking
Yumei Sun, Xiaoming Peng, Meihui Li, Jianlin Zhang 0001 |
Knowl. Based Syst. | 8 |
| 2025 | COLAFormer: Communicating local-global features with linear computational complexity
Zhengwei Miao, Hui Luo 0002, Meihui Li, Jianlin Zhang 0001 |
Pattern Recognit. | 4 |
| 2024 | STCMOT: Spatio-Temporal Cohesion Learning for UAV-Based Multiple Object TrackingabstractMultiple object tracking (MOT) in Unmanned Aerial Vehicle (UAV) videos is important for diverse applications in computer vision. Current MOT trackers rely on accurate object detection results and precise matching of target re-identification (ReID). These methods focus on optimizing target spatial attributes while overlooking temporal cues in modelling object relationships, especially for challenging tracking conditions such as object deformation and blurring, etc. To address the above-mentioned issues, we propose a novel Spatio-Temporal Cohesion Multiple Object Tracking framework (STCMOT), which utilizes historical embedding features to model the representation of ReID and detection features in a sequential order. Concretely, a temporal embedding boosting module is introduced to enhance the discriminability of individual embedding based on adjacent frame cooperation. While the trajectory embedding is then propagated by a temporal detection refinement module to mine salient target locations in the temporal field. Extensive experiments on the VisDrone2019 and UAVDT datasets demonstrate our STCMOT sets a new state-of-the-art performance in MOTA and IDF1 metrics. The source codes are released at https://github.com/ydhcg-BoBo/STCMOT. Jianbo Ma 0001, Chuanming Tang, Fei Wu 0025, Jianlin Zhang 0001, Zhiyong Xu 0007 |
ICME | 5 |
| 2024 | Tracking in tracking: An efficient method to solve the tracking distortion
Jinzhen Yao, Jianlin Zhang 0001, Qintao Hu, Chuanming Tang, Qiliang Bao, Zhenming Peng |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | Learning general features to bridge the cross-domain gaps in few-shot learning
Xiang Li 0212, Hui Luo 0002, Gaofan Zhou, Xiaoming Peng, Jianlin Zhang 0001, Meihui Li |
Knowl. Based Syst. | 6 |
| 2024 | Joint spatio-temporal modeling for visual tracking
Yumei Sun, Chuanming Tang, Hui Luo 0002, Xiaoming Peng, Jianlin Zhang 0001, Meihui Li |
Knowl. Based Syst. | 6 |
| 2024 | Diffusion-based network for unsupervised landmark detection
Kai Wang 0060, Chuanming Tang, Jianlin Zhang 0001 |
Knowl. Based Syst. | 4 |
| 2024 | Improving Visual Representations of Masked Autoencoders With Artifacts SuppressionabstractRecently, Masked Autoencoders (MAE) have gained attention for their abilities to generate visual representations efficiently through pretext tasks. However, there has been little research evaluating the visual representations obtained by pre-trained MAE during the fine-tuning process. In this study, we address the gap by examining the attention maps within each block of the pre-trained MAE during the fine-tuning process. We observed artifacts in pre-trained models, which appear as significant responses in the attention maps of shallow blocks. These artifacts may negatively impact the transfer ability performance of MAE. To address this issue, we localize the cause of these artifacts to the asymmetry between the pre-training and fine-tuning processes. To suppress these artifacts, we propose a novel semantic masking strategy. This strategy aims to preserve complete and continuous semantic information within visible patches while maintaining randomness to facilitate robust representation learning. Experimental results demonstrate that the proposed masking strategy improves the performance of various downstream tasks while reducing artifacts. Specifically, we observed a 3.2% improvement in linear probing, a 0.5% enhancement in fine-tuning on Imagenet1K, and a 0.6% increase in semantic segmentation on ADE20K. Zhengwei Miao, Hui Luo 0002, Jianlin Zhang 0001 |
IEEE Signal Process. Lett. | 4 |
| 2024 | Toward Compact and Robust Model Learning Under Dynamically Perturbed EnvironmentsabstractNetwork pruning has been widely studied to reduce the complexity of deep neural networks (DNNs) and hence speed up their inference. Unfortunately, most existing pruning methods ignore the changes in the model’s robustness before and after pruning, which makes pruned models vulnerable under dynamically perturbed environments (e.g., autonomous driving). Only a few works have explored the robustness of pruned models against adversarial attacks that significantly differ from perturbations in real-world scenarios. To bridge the gap between real-world applications and existing studies, in this work, we propose an adversarial pruning scheme, which automatically identifies and preserves robust channels to obtain robust pruned models that are suitable for practical deployment in dynamically perturbed environments. Specifically, to simulate real-world perturbations, we first employ multi-type adversarial attack samples and adversarial perturbation samples generated by an adversarial perturbation generator to create mixed noise samples. Then, we propose a plug-and-play feature scoring module and a novel contribution difference loss to evaluate the robustness of intermediate features dynamically. Next, to leverage robust intermediate features to identify robust channels, we have developed a simple but effective gating mechanism that evaluates the robustness of channels and preserves robust channels during training. Lastly, we compress the model in a layer-wise or block-wise manner. Compared to existing methods, our scheme enhances the robustness of the pruned model in a broader sense, making it better able to against dynamic perturbations in the real world. Extensive experimental results on well-known dataset benchmarks and popular network architectures demonstrate the effectiveness of our method. Hui Luo 0002, Zhuangwei Zhuang, Yuanqing Li 0001, Mingkui Tan, Cen Chen 0002, Jianlin Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Learning to Generate Diverse Data From a Temporal Perspective for Data-Free QuantizationabstractModel quantization is a prevalent method to compress and accelerate neural networks. Most existing quantization methods usually require access to real data to improve the performance of quantized models, which is often infeasible in some scenarios with privacy and security concerns. Recently, data-free quantization has been widely studied to solve the challenge of not having access to real data by generating synthetic data, among which generator-based data-free quantization is an important type. Previous generator-based methods focus on improving the performance of quantized models by optimizing the spatial distribution of synthetic data, while ignoring the study of changes in synthetic data from a temporal perspective. In this work, we reveal that generator-based data-free quantization methods usually suffer from the issue that synthetic data show homogeneity in the mid-to-late stages of the generation process due to the stagnation of the generator update, which hinders further improvement of the performance of the quantization model. To solve the above issue, we propose introducing the discrepancy between the full-precision and quantized models as new supervision information to update the generator. Specifically, we propose a simple yet effective adversarial Gaussian-margin loss, which promotes continuous updating of the generator by adding more supervision information to the generator when the discrepancy between the full-precision and quantized models is small, thereby generating heterogeneous synthetic data. Moreover, to mitigate the homogeneity of the synthetic data further, we augment the synthetic data with linear interpolation. Our proposed method can also promote the performance of other generator-based data-free quantization methods. Extensive experimental results show that our proposed method achieves superior performances for various settings on data-free quantization, especially in ultra-low-bit settings, such as 3-bit. Hui Luo 0002, Shuhai Zhang, Zhuangwei Zhuang, Jiajie Mai, Mingkui Tan, Jianlin Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Contextual HookFormer for Glacier Calving Front SegmentationabstractPosition changes of glacier calving fronts are important indicators for evaluating the health of ice sheet outlet glaciers and changes in ice dynamics. However, manual delineation of calving fronts in remote sensing imagery is a time-consuming task, resulting in potential large costs. Deep learning-based methods have made remarkable progress in automatically segmenting and delineating glacier calving fronts from remote sensing imagery. The relatively few remote sensing images and the limited geometric changes for glacier observations both reduce the diversity of the data and exacerbate the difficulty of accurate segmentation. Here we describe a novel automatic method for detecting glacier calving fronts in synthetic aperture radar (SAR) images, termed HookFormer. Our approach processes high-resolution (target) and low-resolution (context) inputs with a unified Transformer architecture. The global-local tokens from the context and the target branches are integrated purely by the proposed cross-attention mechanism and cross-interaction module to complement and enhance each other. Moreover, we redesign the HookFormer architecture based on the CNN model AMD-HookNet aiming to improve computational efficiency while achieving significant performance gains with only half of the model parameters/FLOPs. We conduct an in-depth analysis and make extensive comparisons based on the challenging glacier segmentation benchmark dataset CaFFe. As the first pure Transformer approach, HookFormer sets a new state of the art with a mean distance error of 353 m to the ground truth, outperforming the baseline, Swin-Unet, and AMD-HookNet by 53 %, 39 %, and 19 %, respectively. Fei Wu 0025, Nora Gourmelon, Thorsten Seehaus, Jianlin Zhang 0001, Matthias H. Braun, Andreas K. Maier, Vincent Christlein |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Updating Siamese trackers using peculiar mixup
Fei Wu 0025, Jianlin Zhang 0001, Zhiyong Xu 0007, Andreas K. Maier, Vincent Christlein |
Appl. Intell. | 2 |
| 2023 | HybridNet: Integrating GCN and CNN for skeleton-based action recognition
Jianlin Zhang 0001, Jingju Cai, Zhiyong Xu 0007 |
Appl. Intell. | 2 |
| 2023 | Unsupervised detection of contrast enhanced highlight landmarksabstractAbstract In the field of landmark detection based on deep learning, most of the research utilise convolutional neural networks to represent landmarks, and rarely adopt Transformer to represent and encode landmarks. Meanwhile, many works focus on modifying the network structure to improve network performance, and there is little research on the distribution of landmarks. In this article,the authors propose an unsupervised model to extract landmarks of objects in images. First, Transformer structure is combined with the convolutional neural network structure to represent and encode the landmarks; next, positive and negative sample pairs between landmarks are constructed, so that the semantically consistent landmarks on the image are pulled closer in the feature space and the semantically inconsistent landmarks are pushed farther in the feature space; then the authors concentrate their attention on the most active points to distinguish the landmarks of an object from the background; finally, based on the new contrastive loss, the network reconstructs the image by the landmarks of the object that are continuously learnt during training. Experiments show that the proposed model achieves better performance than other unsupervised methods on the CelebA, Annotated Facial Landmarks in the Wild, 300W datasets. Wenzhuo Fan, Jianlin Zhang 0001, Meihui Li |
IET Comput. Vis. | 5 |
| 2023 | Deep-Learning-Based Precision Visual TrackingabstractIn this paper, we target the “precision visual tracking” problem, which is the precise tracking of a target point on a moving object. Compared with the abundant efforts in object-level tracking, which aims at accurately predicting the bounding box (or the mask) of the object, much less attention has been drawn to precision visual tracking. To this end, we present a template tracking framework that consists of two parts, template matching and template updating. For template matching, we trained a deep architecture to directly estimate the projective transformation that deforms the template to the search image. For template updating, we came up with a systematic strategy to update the initial and new templates. To avoid drift build-ups, we incorporate a fast-running dense correspondence matching module into the template update step. The proposed method was extensively tested on both synthetic and real data. To generate the synthetic data, we created a dataset of 480 image sequences. Each image sequence comes with the trajectory ground truth of a target point moving across all the frames. We compared the proposed method with six comparative approaches on this dataset. The tracking accuracy achieved by the proposed method was as follows: around 44% of the tracking errors were less than one pixel, about 78% of them were less than three pixels, and only about 14% of them exceeded five pixels. The proposed method is fast, running at 14 frames-per-second (fps) on [Formula: see text] image sequences on a workstation equipped with a Nvidia RTX 2080Ti graphic card. Qualitative results also show that the proposed method is applicable to real-world image sequences. The related code, pre-trained models, and the test data will be made publicly available at https://github.com/XM-Peng/Precision-Visual-Tracking/ . Xiaoming Peng, Zhiyong Xu 0007, Yufan Peng, Jianlin Zhang 0001, Haorui Zuo |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2023 | Infrared Small Target Detection Based on Multidirectional GradientabstractInfrared (IR) small target detection has always been a challenge in IR search and track systems, which require excellent detection of ultraweak target, robustness, and real-time speed. To achieve these requirements, this letter proposes a new small target detection method that can remarkably suppress the background clutters and sharply enhance the weak target. First, the method locally models the image with facet model and multidirectionally derives the model to separate the target from background clutters; furthermore, via filtering, the multidirectional derivative subbands are obtained. Then, the candidate target points are screened based on the assumption that IR small targets are isotropic. Third, for candidate target points, all the derivative subbands are fused to construct the contrast measure map. Finally, the small target is extracted by threshold segmentation. Experimental results demonstrate that the proposed method can suppress the background better than several state-of-the-art methods while enhancing the target in real time. Jia Liu 0058, Jianlin Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | Information-diffused graph tracking with linear complexity
Jinzhen Yao, Chuanming Tang, Jianlin Zhang 0001, Qiliang Bao, Zhenming Peng |
Pattern Recognit. | 4 |
| 2023 | Learning Spatial-Frequency Transformer for Visual Object TrackingabstractRecently, some researchers have begun to adopt the Transformer to combine or replace the widely used ResNet as their new backbone network. As the Transformer captures the long-range relations between pixels well using the self-attention scheme, which complements the issues caused by the limited receptive field of CNN. Although their trackers work well in regular scenarios, they simply flatten the 2D features into a sequence to better match the Transformer. We believe these operations ignore the spatial prior of the target object, which may lead to sub-optimal results only. In addition, many works demonstrate that self-attention is actually a low-pass filter, which is independent of input features or keys/queries. That is to say, it may suppress the high-frequency component of the input features and preserve or even amplify the low-frequency information. To handle these issues, in this paper, we propose a unified Spatial-Frequency Transformer that models the Gaussian spatial Prior and High-frequency emphasis Attention (GPHA) simultaneously. To be specific, Gaussian spatial prior is generated using dual Multi-Layer Perceptrons (MLPs) and injected into the similarity matrix produced by multiplying Query and Key features in self-attention. The output will be fed into a softmax layer and then decomposed into two components, i.e., the direct and high-frequency signal. The low- and high-pass branches are rescaled and combined to achieve all-pass, therefore, the high-frequency features will be protected well in stacked self-attention layers. We further integrate the Spatial-Frequency Transformer into the Siamese tracking framework and propose a novel tracking algorithm termed SFTransT. The cross-scale fusion based SwinTransformer is adopted as the backbone, and also a multi-head cross-attention module is used to boost the interaction between search and template features. The output will be fed into the tracking head for target localization. Extensive experiments on short-term and long-term tracking benchmarks all demonstrate the effectiveness of our proposed framework. Source code will be released athttps://github.com/Tchuanm/SFTransT.git. Chuanming Tang, Xiao Wang 0014, Yuanchao Bai, Zhe Wu 0006, Jianlin Zhang 0001, Yongmei Huang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | AMD-HookNet for Glacier Front SegmentationabstractKnowledge on changes in glacier calving front positions is important for assessing the status of glaciers. Remote sensing imagery provides the ideal database for monitoring calving front positions; however, it is not feasible to perform this task manually for all calving glaciers globally due to time constraints. Deep-learning-based methods have shown great potential for glacier calving front delineation from optical and radar satellite imagery. The calving front is represented as a single thin line between the ocean and the glacier, which makes the task vulnerable to inaccurate predictions. The limited availability of annotated glacier imagery leads to a lack of data diversity (not all possible combinations of different weather conditions, terminus shapes, sensors, etc. are present in the data), which exacerbates the difficulty of accurate segmentation. In this article, we propose attention-multihooking-deep-supervision HookNet (AMD-HookNet), a novel glacier calving front segmentation framework for synthetic aperture radar (SAR) images. The proposed method aims to enhance the feature representation capability through multiple information interactions between low-resolution and high-resolution inputs based on a two-branch U-Net. The attention mechanism, integrated into the two branch U-Net, aims to interact between the corresponding coarse and fine-grained feature maps. This allows the network to automatically adjust feature relationships, resulting in accurate pixel classification predictions. Extensive experiments and comparisons on the challenging glacier segmentation benchmark dataset CaFFe show that our AMD-HookNet achieves a mean distance error (MDE) of 438 m to the ground truth outperforming the current state of the art by 42%, which validates its effectiveness. Fei Wu 0025, Nora Gourmelon, Thorsten Seehaus, Jianlin Zhang 0001, Matthias H. Braun, Andreas K. Maier, Vincent Christlein |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Transformer Sub-Patch Matching for High-Performance Visual Object TrackingabstractVisual tracking is a core component of intelligent transportation systems, especially for unmanned driving and road surveillance. Numerous convolutional neural network (CNN) trackers have achieved unprecedented performance. However, CNN features with regular spatial context relationships experience difficulty matching the rigid target templates when dramatic deformation and occlusion occur. In this paper, we propose a novel full Transformer Sub-patch Matching network for tracking (TSMtrack), which decomposes the tracked object into sub-patches, and interlaced matches the extracted sub-patches by leveraging the attention mechanism born with the Transformer. Roots in Transformer architecture, TSMtrack consists of image patch decomposition, sub-patch matching, and position prediction. Specifically, TSMtrack converts the whole frame into sub-patches and extracts the sub-patch features independently. By sub-patch matching and FFN-like prediction, TSMtrack enables independent similarity measurement between sub-patch features in an interlaced and iterative fashion. With a full Transformer pipeline implemented, we achieve a high-quality trade-off between tracking speed performance. Experiments on nine benchmarks demonstrate the effectiveness of our Transformer sub-patch matching framework. In particular, it realizes an AO of 75.6 on GOT-10K and SR of 57.9 on WebUAV-3M with 48 FPS on GPU RTX-2060s. Chuanming Tang, Qintao Hu, Gaofan Zhou, Jinzhen Yao, Jianlin Zhang 0001, Yongmei Huang, Qixiang Ye |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | Semantic and context features integration for robust object trackingabstractAbstract Siamese network‐based object tracking learns features of a target object marked in the first frame and that of the object in subsequent frames simultaneously and then measures similarity between two features to recognize and locate the object. Owing to their efficiency and high accuracy, Siamese networks have attracted much attention recently. However, tracking accuracy decreases significantly when there are scale changes, occlusion, and pose variations due to the way that Siamese networks estimate feature similarity. To address this issue, the authors propose a tracking algorithm, named Semantic and context features integration for robust object tracking that integrates local and global features of the object. Local features provide context information for tracking parts of the object, while global features contain semantic information for tracking the object. The authors meticulously design local and global classification and regression heads and integrate them into one uniform framework to achieve integration tracking. This method effectively alleviates low accuracy in complex scenes such as scale changes, deformation, and occlusion. Numerous experiments demonstrate that this method achieves state‐of‐art (SOTA) performance with 45 FPS on a single RTX2060 Super GPU on public tracking datasets, including VOT2016, VOT2019, OTB100, GOT‐10k, and LaSOT, and its effectiveness and efficiency is confirmed. Jinzhen Yao, Jianlin Zhang 0001, Linsong Shao |
IET Image Process. | 2 |
| 2022 | Infrared Small Target Detection Based on a Group Image-Patch Tensor ModelabstractDespite many years of research, the accuracy of small target detection is still restricted by the complex background. In addition, the contradiction between the computation complexity reduction and detection performance improvement makes it hard to apply some high-performance methods toward practical application. In this letter, a novel group image-patch tensor (GIPT) model is proposed to solve the above-mentioned problems. First, based on the structure tensor theory, a newly designed measurement criterion is employed to divide the image pixels into three types: point area, flat area, and line area. Second, image pixels with line priors and point priors are assigned to different weights to reduce the effects of strong background edges. Third, a GIPT model is proposed to better explore the low-rank of the background, in which patches with the same feature type are rearranged into the same group for tensor decomposition. Moreover, the proposed GIPT model can make the optimization process run in parallel, which can greatly reduce the complexity of the algorithms. The experiment results show the superiority of the proposed method compared with other state-of-the-art algorithms. Lanlan Yang, Meihui Li, Jianlin Zhang 0001, Zhiyong Xu 0007 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2021 | Model Uncertainty Guides Visual Object TrackingabstractModel object trackers largely rely on the online learning of a discriminative classifier from potentially diverse sample frames. However, noisy or insufficient amounts of samples can deteriorate the classifiers' performance and cause tracking drift. Furthermore, alterations such as occlusion and blurring can cause the target to be lost. In this paper, we make several improvements aimed at tackling uncertainty and improving robustness in object tracking. Our first and most important contribution is to propose a sampling method for the online learning of object trackers based on uncertainty adjustment: our method effectively selects representative sample frames to feed the discriminative branch of the tracker, while filtering out noise samples. Furthermore, to improve the robustness of the tracker to various challenging scenarios, we propose a novel data augmentation procedure, together with a specific improved backbone architecture. All our improvements fit together in one model, which we refer to as the Uncertainty Adjusted Tracker (UATracker), and can be trained in a joint and end-to-end fashion. Experiments on the LaSOT, UAV123, OTB100 and VOT2018 benchmarks demonstrate that our UATracker outperforms state-of-the-art real-time trackers by significant margins. Antoine Ledent, Qintao Hu, Jianlin Zhang 0001, Marius Kloft |
AAAI | 5 |
| 2021 | An efficient semantic segmentation method based on transfer learning from object detectionabstractAbstract Nowadays, numerous semantic segmentation techniques were used to complex scenes such as urban streets. However, speed issues are not considered in most of these methods, and real‐time methods do not mainly include enough accuracy. In this paper, an efficient semantic segmentation method is proposed, using the feature extractor of a real‐time object detection model, Darknet53, as the backbone of DeepLabv3+. By the high accuracy of DeepLabv3+ structure and great efficiency of Darknet53, a mean intersection was obtained over union of 76.3% in Cityscapes test set, and fast inference speed simultaneously (0.178 s per frame on one GTX 1080Ti GPU). A huge imbalance of objects was noticed on Cityscapes dataset. To solve this problem, a Focal Loss like loss function was proposed to concentrate more on the hard difficult pixels. Moreover, an atrous convolution block was proposed to extract more high‐level features. Based on the experimental results, it is proved that these changes contribute to a better result on the Cityscapes test set (77.8% mean Intersection over Union) and faster inference speed (0.171 s per frame). Authors' model achieves state‐of‐art results on Cityscapes test set (79.1% mean Intersection over Union) after fine‐tuning on Cityscapes coarsely annotated dataset. Jianlin Zhang 0001, Zhongbi Chen, Zhiyong Xu 0007 |
IET Image Process. | 2 |
| 2020 | SPSTracker: Sub-Peak Suppression of Response Map for Robust Object TrackingabstractModern visual trackers usually construct online learning models under the assumption that the feature response has a Gaussian distribution with target-centered peak response. Nevertheless, such an assumption is implausible when there is progressive interference from other targets and/or background noise, which produce sub-peaks on the tracking response map and cause model drift. In this paper, we propose a rectified online learning approach for sub-peak response suppression and peak response enforcement and target at handling progressive interference in a systematic way. Our approach, referred to as SPSTracker, applies simple-yet-efficient Peak Response Pooling (PRP) to aggregate and align discriminative features, as well as leveraging a Boundary Response Truncation (BRT) to reduce the variance of feature response. By fusing with multi-scale features, SPSTracker aggregates the response distribution of multiple sub-peaks to a single maximum peak, which enforces the discriminative capability of features for robust object tracking. Experiments on the OTB, NFS and VOT2018 benchmarks demonstrate that SPSTrack outperforms the state-of-the-art real-time trackers with significant margins1 Qintao Hu, Yao Mao, Jianlin Zhang 0001, Qixiang Ye |
AAAI | 5 |