Shunli Zhang 0005

dblp:18/7951-5 · DBLP profile ↗
← Back
56ranked-venue papers
8as first author
32since 2021 · last 2026
0000-0002-8186-8949ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 36 · 4 first-author · 22 since 2021Artificial intelligence and machine learning · 27 · 4 first-author · 17 since 2021Security and privacy · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 SSPFusion: A semantic structure-preserving approach for multi-modality image fusion
Yu Zhang 0026, Jian Zhang 0121, Shunli Zhang 0005
Expert Syst. Appl.5
2026 MTFusion: A dual-task-driven mean teacher framework for infrared and visible image fusion
Yu Zhang 0026, Junfu Chen, Jian Zhang 0121, Shunli Zhang 0005
Knowl. Based Syst.6
2026 GaitAdapt: Continual learning for evolving gait recognition
Shunli Zhang 0005, Senmao Tian
Pattern Recognit.2
2026 SAT-UIR: Self-Assessment Training for Semi-Supervised Underwater Image Restoration
abstract
Underwater images, often affected by light attenuation and particle scattering, pose a challenge for restoration, aggravated by the difficulty in obtaining a substantial amount of annotated data. Existing methods have tackled this issue through the development of semi-supervised frameworks; however, they commonly lack a suitable strategy or rely on additional models trained on extra data to ensure the quality of pseudo-labels. To address this, we propose a self-assessment training framework for semi-supervised underwater image restoration (SAT-UIR). SAT-UIR employs a dual-task network (DT-Net) incorporating an auxiliary assessment task to align the restored image with a target structure similarity index measure score. This enables accurate restoration completeness estimation at a feature level and effective pseudo-label filtering during self-training. Leveraging multi-scale features, the assessment task also encourages the model to learn advantageous features for image restoration. Moreover, we integrate a soft ranking loss to further refine the training process of the auxiliary assessment task. Comprehensive experiments on various underwater benchmarks demonstrate that SAT-UIR outperforms state-of-the-art methods quantitatively and qualitatively. The code is available at https://github.com/aroid721/SAT-UIR.
Qianying Tang, Xiaoyu Guo 0001, Wei Xiang 0007, Dongjin Wang, Shunli Zhang 0005
IEEE Trans. Circuits Syst. Video Technol.6
2025 Human Identification at a Distance: Challenges, Methods and Results on the Competition HID 2025
abstract
Human identification at a distance (HID) faces challenges due to the difficulty of acquiring traditional biometric modalities like face and fingerprints. Gait recognition offers a viable solution since it can be captured at a distance. To promote progress in gait recognition and provide a fair evaluation platform, the International Competition on Human Identification at a Distance (HID) has been organized annually since 2020. Since 2023, the competition has adopted the challenging SUSTech-Competition dataset, which includes significant variations in clothing, carried objects, and view angles. No training data is provided, requiring participants to train their models using external datasets. Each year, the competition applies a different random seed to generate distinct evaluation splits, reducing the risk of overfitting and ensuring fair evaluation of cross-domain generalization. Although the previous two competitions (HID 2023 and HID 2024) already utilized this dataset, HID 2025 aimed explicitly to explore whether algorithmic improvements could surpass the accuracy limits observed previously. Despite these heightened challenges, participants again demonstrated significant advancements, with the highest accuracy reaching 94.2%, setting a new benchmark for this dataset. We also analyze key technical trends and outline potential directions for future research on gait recognition.
Jingzhe Ma, Jianlong Yu, Zunxiao Xu, Xue Cheng, Zepeng Wang 0002, Kazuki Osamura, Rujie Liu, Narishige Abe, Shunli Zhang 0005, Haojun Xie, Weiming Wu, Wenxiong Kang, Qingshuo Gao, Jiaming Xiong, Xianye Ben, Lei Chen 0095, Lichen Song, Junjian Cui, Haijun Xiong, Junhao Lu, Bin Feng 0001, Baoquan Zhao, Ke Xu 0001, Yongzhen Huang, Liang Wang 0001, Manuel J. Marín-Jiménez, Md. Atiqur Rahman Ahad, Shiqi Yu 0001
IJCB15
2025 QGait: Toward Accurate Quantization for Gait Recognition
abstract
Existing deep learning methods have made significant progress in gait recognition. Quantization can facilitate the application of gait models as a model-agnostic general compression technique. Typically, appearance-based models binarize inputs into silhouette sequences. However, mainstream quantization methods prioritize minimizing task loss over quantization error, which is detrimental to gait recognition with binarized inputs. To address this, we propose a differentiable soft quantizer, which better simulates the gradient of the round function during backpropagation. This enables the network to learn from subtle input perturbations. However, our theoretical analysis and empirical studies reveal that directly applying the soft quantizer can hinder network convergence. We addressed this issue by adopting a two-stage training strategy, introducing a soft quantizer during the fine-tuning phase. However, in the first stage of training, we observed a significant change in the output distribution of different samples in the feature space compared to the full-precision network. It is this change that led to a loss in performance. Based on this, we propose an Inter-class Distance-guided Calibration (IDC) strategy to preserve the relative distance between the embeddings of samples with different labels. Extensive experiments validate the effectiveness of our approach, demonstrating state-of-the-art accuracy across various settings and datasets. The code will be made publicly available.
Senmao Tian, Gangyi Hong, JingJie Wang, Xin Yu 0002, Shunli Zhang 0005
IJCB7
2024 NightRain: Nighttime Video Deraining via Adaptive-Rain-Removal and Adaptive-Correction
abstract
Existing deep-learning-based methods for nighttime video deraining rely on synthetic data due to the absence of real-world paired data. However, the intricacies of the real world, particularly with the presence of light effects and low-light regions affected by noise, create significant domain gaps, hampering synthetic-trained models in removing rain streaks properly and leading to over-saturation and color shifts. Motivated by this, we introduce NightRain, a novel nighttime video deraining method with adaptive-rain-removal and adaptive-correction. Our adaptive-rain-removal uses unlabeled rain videos to enable our model to derain real-world rain videos, particularly in regions affected by complex light effects. The idea is to allow our model to obtain rain-free regions based on the confidence scores. Once rain-free regions and the corresponding regions from our input are obtained, we can have region-based paired real data. These paired data are used to train our model using a teacher-student framework, allowing the model to iteratively learn from less challenging regions to more challenging regions. Our adaptive-correction aims to rectify errors in our model's predictions, such as over-saturation and color shifts. The idea is to learn from clear night input training videos based on the differences or distance between those input videos and their corresponding predictions. Our model learns from these differences, compelling our model to correct the errors. From extensive experiments, our method demonstrates state-of-the-art performance. It achieves a PSNR of 26.73dB, surpassing existing nighttime video deraining methods by a substantial margin of 13.7%.
Beibei Lin, Yeying Jin, Wending Yan, Wei Ye 0005, Yuan Yuan 0039, Shunli Zhang 0005, Robby T. Tan
AAAI6
2024 Unveiling the Significance of Width Dimension in Bird's-Eye View Segmentation
abstract
Recent advancements in autonomous driving have prominently utilized the bird’s-eye view (BEV) as an intermediary representation of the world. However, with escalating precision requirements, the dimensionality of the BEV representation has expanded, leading to increased complexity. In this paper, we propose a novel real-time BEV segmentation model, which can perform height dimension reduction while view transformation. It also utilizes a novel decoder structure aimed at improving the fusion of overall situations and details. Consequently, we prove that the height dimension for BEV feature extraction and processing tasks is not a necessary component. Experimental results on the nuScenes dataset substantiate that our model surpasses previous approaches in terms of speed and achieves the state-of-the-art performance.
Shunli Zhang 0005
ICME5
2024 Data-driven Multi-stage Vehicle Trajectory Prediction for Complex Scenarios
abstract
Owing to the advancements in deep learning methods, predicting vehicle trajectories based on vast amounts of traffic trajectory data is no longer an elusive objective. However, traditional time-series prediction methods struggle to effectively address issues such as V-V interactions, road guidance, and the inherent uncertainty in vehicle trajectories within road traffic. To tackle these issues, this paper proposes a data-driven multistage prediction method. Firstly, in terms of map representation, a slide-window-driven lane node aggregation is employed to compensate for the absence of scale information in lane nodes. Secondly, distance-wise attributes for complex areas are propagated and encoded to obtain enhanced features. Thirdly, by defining the hist features, the perception range of vehicle interactions is expanded both temporally and spatially. Finally, a multi-stage scheme is utilized to achieve the final aggregation of trajectory features, and a feature decoder integrated with a focal loss is employed to obtain multi-modal prediction results. Both qualitative and quantitative analyses of the experimental results demonstrate the effectiveness of our proposed method.
Jian Zhang 0121, Zikun Feng, Shunli Zhang 0005
ISPA4
2024 ABAE: Auxiliary Balanced AutoEncoder for class-imbalanced semi-supervised learning
Qianying Tang, Wei Xiang 0007, Shunli Zhang 0005
Pattern Recognit. Lett.4
2024 IAIFNet: An Illumination-Aware Infrared and Visible Image Fusion Network
abstract
Infrared and visible image fusion (IVIF) aims to create fused images that encompass the comprehensive features of both input images, thereby facilitating downstream vision tasks. However, existing methods often overlook illumination conditions in low-light environments, resulting in fused images where targets lack prominence. To address these shortcomings, we introduce the Illumination-Aware Infrared and Visible Image Fusion Network, abbreviated by IAIFNet. Within our framework, an illumination enhancement network initially estimates the incident illumination maps of input images, based on which the textural details of input images under low-light conditions are enhanced specifically. Subsequently, an image fusion network adeptly merges the salient features of illumination-enhanced infrared and visible images to produce a fusion image of superior visual quality. Our network incorporates a Salient Target Aware Module (STAM) and an Adaptive Differential Fusion Module (ADFM) to respectively enhance gradient and contrast with sensitivity to brightness. Extensive experimental results validate the superiority of our method over seven state-of-the-art approaches for fusing infrared and visible images on the public LLVIP dataset. Additionally, the lightweight design of our framework enables highly efficient fusion of infrared and visible images. Finally, evaluation results on the downstream multi-object detection task demonstrate the significant performance boost our method provides for detecting objects in low-light environments.
Yu Zhang 0026, Zijing Zhao 0002, Jian Zhang 0121, Shunli Zhang 0005
IEEE Signal Process. Lett.5
2024 Exploring the Applicability of Spectral Recovery in Semantic Segmentation of RGB Images
abstract
Compared with RGB images, hyperspectral images (HSIs) offer a distinct advantage in that they can record continuous spectral bands of light reflectance in each pixel, reflecting the physical and chemical characteristics of materials. This capability enables differentiation between objects that may have similar textures but different spectral characteristics. It is desirable to recover spectral information from RGB images to improve semantic segmentation accuracy. Additionally, semantic information can serve as a guide for spectral information recovery, thereby ensuring the quality of the recovered spectral information. The two tasks are mutually beneficial in this regard. In light of these considerations, we propose a multi-task framework that exploits the complementary relationship between spectral recovery and semantic segmentation tasks, comprising a complementary spectral-semantic attentive fusion model (CSSF) that enables the two tasks to mutually facilitate each other by fusing information from both branches. Specifically, the proposed CSSF incorporates a window-based spectral-semantic attentive fusion (WSSAF) module to incorporate recovered spectral information into the segmentation process effectively, and a pixel-shuffle-based fusion (PSF) module to provide semantic guidance for spectral recovery. To evaluate the effectiveness of our approach, we built the first flower hyperspectral image dataset (FHRS) with corresponding segmentation annotations and RGB images. By doing so, we have made the first attempt to explore the complementary relationship between semantic segmentation and spectral recovery. Experimental results on both the FHRS dataset and the publicly available LIB-HSI dataset demonstrate that our proposed method has the ability to enhance both tasks by utilizing their complementary relationship, indicating the generalization ability of our method.
Zhuoran Du, Shikui Wei, Ting Liu 0012, Shunli Zhang 0005, Shiyin Zhang, Yao Zhao 0001
IEEE Trans. Multim.4
2024 DCRP: Class-Aware Feature Diffusion Constraint and Reliable Pseudo-Labeling for Imbalanced Semi-Supervised Learning
abstract
Despite the astounding progress made in semi-supervised learning (SSL) and imbalanced supervised learning (ISL), there has been little attention devoted to the research of imbalanced semi-supervised learning (ISSL). The ‘Matthew effect’, a phenomenon where a disparity in data representation becomes more severe in a class-imbalanced dataset during training, could be amplified in a semi-supervised setting. In this study, we addressed two key challenges in ISSL: maintaining the reliability of pseudo-labels and ensuring a balanced representation of features. Specifically, we propose a class-aware feature-diffusion constraint and reliable pseudo-labeling (DCRP) framework to address these issues. In the DCRP, we counteract the overconfidence problem of softmax by adding an extra class to the typical K class problem without the need for additional parameters. Moreover, we introduced a flexible class-aware feature diffusion constraint in the feature extractor, promoting a more balanced feature diversity. Experimental validations on various datasets, such as CIFAR10-LT, CIFAR100-LT, SVHN-LT, and Small ImageNet-127, demonstrated consistent improvements in accuracy with our DCRP method. In particular, we achieved a steady improvement in accuracy of approximately 1% under the newly published ACR prototype across most settings. The code is available athttps://github.com/guoxiaoyuatbjtu/DCRP.
Xiaoyu Guo 0001, Wei Xiang 0007, Shunli Zhang 0005, Wei Lu 0010, Weiwei Xing
IEEE Trans. Multim.3
2023 CABM: Content-Aware Bit Mapping for Single Image Super-Resolution Network with Large Input
abstract
With the development of high-definition display devices, the practical scenario of Super-Resolution (SR) usually needs to super-resolve large input like 2K to higher resolution (4K/8K). To reduce the computational and memory cost, current methods first split the large input into local patches and then merge the SR patches into the output. These methods adaptively allocate a subnet for each patch. Quantization is a very important technique for network acceleration and has been used to design the subnets. Current methods train an MLP bit selector to determine the propoer bit for each layer. However, they uniformly sample subnets for training, making simple subnets overfitted and complicated subnets underfitted. Therefore, the trained bit selector fails to determine the optimal bit. Apart from this, the introduced bit selector brings additional cost to each layer of the$SR$network. In this paper, we propose a novel method named Content-Aware Bit Mapping (CABM), which can remove the bit selector without any performance loss. CABM also learns a bit selector for each layer during training. After training, we analyze the relation between the edge information of an input patch and the bit of each layer. We observe that the edge information can be an effective metric for the selected bit. Therefore, we design a strategy to build an Edge-to-Bit lookup table that maps the edge score of a patch to the bit of each layer during inference. The bit configuration of SR network can be determined by the lookup tables of all layers. Our strategy can find better bit configuration, resulting in more efficient mixed precision networks. We conduct detailed experiments to demonstrate the generalization ability of our method. The code will be released.
Senmao Tian, Ming Lu 0002, Jiaming Liu 0003, Yandong Guo, Yurong Chen 0001, Shunli Zhang 0005
CVPR6
2023 A Comprehensive Comparison of Projections in Omnidirectional Super-Resolution
abstract
Super-Resolution (SR) has gained increasing research attention over the past few years. With the development of Deep Neural Networks (DNNs), many super-resolution methods based on DNNs have been proposed. Although most of these methods are aimed at ordinary frames, there are few works on super-resolution of omnidirectional frames. In these works, omnidirectional frames are projected from the 3D sphere to a 2D plane by Equi-Rectangular Projection (ERP). Although ERP has been widely used for projection, it has severe projection distortion near poles. Current DNN-based SR methods use 2D convolution modules, which is more suitable for the regular grid. In this paper, we find that different projection methods have great impact on the performance of DNNs. To study this problem, a comprehensive comparison of projections in omnidirectional super-resolution is conducted. We compare the SR results of different projection methods. Experimental results show that Equi-Angular cube map projection (EAC), which has minimal distortion, achieves the best result in terms of WS-PSNR compared with other projections. Code and data will be released.
Huicheng Pi, Senmao Tian, Ming Lu 0002, Jiaming Liu 0003, Yandong Guo, Shunli Zhang 0005
ICASSP6
2023 Gait Recognition with Mask-based Regularization
abstract
Most gait recognition methods exploit spatial-temporal representations from static appearances and dynamic walking patterns. However, we observe that many part-based methods neglect representations at boundaries. In addition, the phenomenon of overfitting on training data is relatively common in gait recognition, which is perhaps due to insufficient data and low-informative gait silhouettes. Motivated by these observations, we propose a novel mask-based regularization method named ReverseMask. By injecting perturbation on the feature map, the proposed regularization method helps convolutional architecture learn the discriminative representations and enhances generalization. Also, we design an Inception-like ReverseMask Block, which has three branches composed of a global branch, a feature-dropping branch, and a feature scaling branch. Precisely, the dropping branch can extract fine-grained representations when partial activations are zero-outed. Meanwhile, the scaling branch randomly scales the feature map, keeping structural information of activations and preventing overfitting. The plug-and-play Inception-like ReverseMask block is simple and effective, improving the performance of many state-of-the-art methods. Extensive experiments demonstrate that the ReverseMask regularization help baseline achieves higher accuracy and better generalization. Moreover, the base-line with Inception-like Block significantly outperforms state-of-the-art methods on the two most popular datasets, CASIA-B and OUMVLP.
Chuanfu Shen, Beibei Lin, Shunli Zhang 0005, Xin Yu 0002, George Q. Huang, Shiqi Yu 0001
IJCB3
2023 DyGait: Exploiting Dynamic Representations for High-performance Gait Recognition
abstract
Gait recognition is a biometric technology that recognizes the identity of humans through their walking patterns. Compared with other biometric technologies, gait recognition is more difficult to disguise and can be applied to the condition of long-distance without the cooperation of subjects. Thus, it has unique potential and wide application for crime prevention and social security. At present, most gait recognition methods directly extract features from the video frames to establish representations. However, these architectures learn representations from different features equally but do not pay enough attention to dynamic features, which refers to a representation of dynamic parts of silhouettes over time (e.g. legs). Since dynamic parts of the human body are more informative than other parts (e.g. bags) during walking, in this paper, we propose a novel and high-performance framework named DyGait. This is the first framework on gait recognition that is designed to focus on the extraction of dynamic features. Specifically, to take full advantage of the dynamic information, we propose a Dynamic Augmentation Module (DAM), which can automatically establish spatial-temporal feature representations of the dynamic parts of the human body. The experimental results show that our DyGait network outperforms other state-of-the-art gait recognition methods. It achieves an average Rank-1 accuracy of 71.4% on the GREW dataset, 66.3% on the Gait3D dataset, 98.4% on the CAS1A-B dataset and 98.3% on the OU-MVLP dataset.
Xianda Guo, Beibei Lin, Lincheng Li, Shunli Zhang 0005, Xin Yu 0002
ICCV7
2023 HQRetouch: Learning Professional Face Retouching Via Masked Feature Fusion and Semantic-Aware Modulation
abstract
Face retouching is a crucial technique for many consumer-level products. The goal of face retouching is to remove skin imperfections and preserve facial details simultaneously. However, it usually requires tedious manual work to achieve professional retouching effect. With the advent of Deep Neural Networks (DNNs), some methods were recently proposed to complete the task of face retouching automatically by using DNNs. They divide a portrait photo into local patches and train a DNN for face retouching. Although they can produce professional results automatically, there are still some limitations. Firstly, the network architecture fails to preserve sufficient facial details. Secondly, the facial semantic information is ignored when dividing a photo into some local patches. In this paper, we propose a novel method to solve these limitations. We first introduce the Masked Feature Fusion (MFF) module to a UNet, enabling the network to better preserve details in facial regions. Then, we exploit the semantic information by the Semantic-Aware Modulation (SAM) module, further boosting the retouching performance. Experiments on the recent public dataset Flickr-Faces-HQ-Retouched (FFHQR) demonstrate the effectiveness of our method. The code will be released.
Gangyi Hong, Fangshi Wang, Senmao Tian, Ming Lu 0002, Jiaming Liu 0003, Shunli Zhang 0005
ICIP6
2023 A coarse-to-fine parallelizable surface defect detection approach for railway trackside equipment
abstract
Surface defects of railway trackside equipment pose a serious risk on the safety of railway transportation systems. Image-based surface defect detection methods have made significant progress. However, the image background of trackside equipment is complex, and there is a large amount of noise, which makes existing methods inadequate in accurately detecting small surface defect regions. To tackle with this issue, we propose a coarse-to-fine parallelizable surface defect detection approach to hierarchically detect the defects of trackside equipment. Firstly, a detection network is designed to locate and extract trackside equipment, which aims at roughly focusing the detection field from the original image to the region of interest of individual trackside equipment. Then, a novel semantic segmentation network is proposed to segment the major components of trackside equipment, so as to further finely focus on the defect regions. We apply multiple segmentation networks to parallelly segment various trackside equipment. In the segmentation network, a dense feature enhancement method is introduced to strengthen the high-level semantic information, and a feature partitioning enhancement strategy is designed to improve the segmentation performance for small defect regions. Finally, according to the visual characteristics of the segmentation output, we propose a defect recognizer to discriminate the defects. Extensive experimental results demonstrate that the proposed surface defect detection approach achieves higher accuracy for trackside equipment.
Guanjia Zhang, Weiwei Xing, Shuzhong Yang, Weibin Liu, Wei Xiang 0007, Jian Zhang 0121, Shunli Zhang 0005
ICPADS7
2023 Using Mask-Based Enhancement and Feature Aggregation for Single Image Deraining
abstract
Image rain removal is an essential and challenging low-level task in computer vision. In recent years, although great progress has been achieved in the field of image and video deraining, the restored details may be incomplete in many cases, making it difficult to recover the clear images from diverse rain forms. Moreover, the features of layers may not be fully exploited, which may reduce the rain removal performance to some degree. To address the above issues, we propose a novel Mask-based Enhancement and Feature Aggregation Network (MEFA-Net). First, we develop a novel structure with enhancement branch, which can greatly improve the feature representation capability. The enhancement branch is implemented based on the random masking technique, which can generate complementary features and improve the robustness. Specially, we design a novel adaptive weighting block in the MEFA-Net, which can adaptively assign weights to the enhanced features for each specific image, effectively improving the generalization ability of MEFA-Net. Second, we propose a feature aggregation sub-module (FAS) for comprehensive representation by integrating the features from different layers. Experimental results demonstrate that the proposed MEFA-Net can achieve better performance than most state-of-the-art image deraining methods on four synthetic and one real-world datasets.
Shengdi Qin, Shunli Zhang 0005, Yu Zhang 0026
IEEE Signal Process. Lett.2
2023 Contrastive JS: A Novel Scheme for Enhancing the Accuracy and Robustness of Deep Models
abstract
Deep learning technologies have been applied in various computer vision tasks in recent years. However, deep models suffer performance decay when some unforeseen data are contained in the testing dataset. Although data enhancement techniques can alleviate this dilemma, the diversity of real data is too tremendous to simulate. To tackle this challenge, we study a scheme for improving the robustness and efficiency of the deep network training process in visual tasks. Specifically, first, we build positive and negative sample pairs based on a class-sensitive strategy. Then, we construct a feature-consistent learning strategy based on contrastive learning to constrain the representations of interclass features while paying attention to the intraclass features. To extend the effect of the consistent strategy, we propose a novel contrastive Jensen-Shannon divergence consistency loss (JS loss) to restrict the probability distributions of different sample pairs. The proposed scheme successfully enhances the robustness and accuracy of the utilized model. We validated our approach by conducting extensive experiments in the domains of model robustness and few-shot object detection (FSOD). The results showed that the proposed method achieved remarkable gains over state-of-the-art (SOTA) methods. We obtained a 3.2% average improvement over the best-performing FSOD method.
Weiwei Xing, Zixia Liu, Weibin Liu, Shunli Zhang 0005, Liqiang Wang 0001
IEEE Trans. Multim.5
2022 GaitStrip: Gait Recognition via Effective Strip-Based Feature Representations and Multi-level Framework
Beibei Lin, Xianda Guo, Lincheng Li, Jiande Sun 0001, Shunli Zhang 0005, Xin Yu 0002
ACCV (4)7
2022 Detecting prohibited objects with physical size constraint from cluttered X-ray baggage images
An Chang, Yu Zhang 0026, Shunli Zhang 0005, Leisheng Zhong, Li Zhang 0023
Knowl. Based Syst.3
2022 Using Segmentation With Multi-Scale Selective Kernel for Visual Object Tracking
abstract
Generic visual object tracking is challenging due to various difficulties, e.g. scale variations and deformations. To solve those problems, we propose a novel multi-scale selective kernel module for tracking, which contains small-scale and large-scale branches to model the target at different scales and attention mechanism to capture the more effective appearance information of the target. In our module, we cascade multiple small-scale convolutional blocks as an equivalent large-scale branch to extract large-scale features of the target effectively. Besides, we present a hybrid strategy for feature selection to extract significant information from features of different scales. Based on the current excellent segmentation tracking framework, we propose a novel tracking network that leverages our module at multiple places in the up-sample phase to construct a more accurate and robust appearance model. Extensive experimental results show that our tracker outperforms other state-of-the-art trackers on multiple challenging benchmarks including VOT2018, TrackingNet, DAVIS-2017, and YouTube-VOS-2018 while achieves real-time tracking.
Yifei Cao, Shunli Zhang 0005, Beibei Lin, Sicong Zhao
IEEE Signal Process. Lett.3
2022 NoisyOTNet: A Robust Real-Time Vehicle Tracking Model for Traffic Surveillance
abstract
With the rapid development of intelligent transportation, automated traffic surveillance is considered as an important component. In the field of traffic surveillance, it is particularly important to achieve robust and real-time tracking of vehicles in complex scenes. In this paper, a robust real-time vehicle tracking model namedNoisyOTNetis proposed, which formulates tracking as reinforcement learning with parameter space noise. In this formulation, the exploration ability of the model is enhanced to improve the robustness of tracking. Specifically, we develop a new implementation for noisy network based on deep deterministic policy gradients (DDPGs) with parameter noise, which can better cope with the tracking task and directly predict the tracking result. To improve the tracking accuracy in complex conditions, e.g. fast motion and large deformation, this paper presents an adaptive update strategy that can exploit the vehicle spatial-temporal information based on Upper Confidence Bound (UCB) algorithm by exploiting. Moreover, as for the recovery of the lost target, a relocation algorithm based on incremental learning is developed. The results of extensive experiments demonstrate that the proposed NoisyOTNet can effectively track vehicles in complex scenes and achieve competitive performance compared to the state-of-the-art methods.
Weiwei Xing, Yuxiang Yang 0002, Shunli Zhang 0005, Liqiang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 Single-Image Deraining via Recurrent Residual Multiscale Networks
abstract
Existing deraining approaches represent rain streaks with different rain layers and then separate the layers from the background image. However, because of the complexity of real-world rain, such as various densities, shapes, and directions of rain streaks, it is very difficult to decompose a rain image into clean background and rain layers. In this article, we develop a novel single-image deraining method based on residual multiscale pyramid to mitigate the difficulty of rain image decomposition. To be specific, we progressively remove rain streaks in a coarse-to-fine fashion, where heavy rain is first removed in coarse-resolution levels and then light rain is eliminated in fine-resolution levels. Furthermore, based on the observation that residuals between a restored image and its corresponding rain image give critical clues of rain streaks, we regard the residuals as an attention map to remove rains in the consecutive finer level image. To achieve a powerful yet compact deraining framework, we construct our network by recurrent layers and remove rain with the same network in different pyramid levels. In addition, we design a multiscale kernel selection network (MSKSN) to facilitate our single network to remove rain streaks at different levels. In this manner, we reduce 81% of the model parameters without decreasing deraining performance compared with our prior work. Extensive experimental results on widely used benchmarks show that our approach achieves superior deraining performance compared with the state of the art.
Yupei Zheng, Xin Yu 0002, Miaomiao Liu 0001, Shunli Zhang 0005
IEEE Trans. Neural Networks Learn. Syst.4
2021 GaitMask: Mask-based Model for Gait Recognition
Beibei Lin, Shunli Zhang 0005
BMVC3
2021 Gait Recognition via Effective Global-Local Feature Representation and Local Temporal Aggregation
abstract
Gait recognition is one of the most important biometric technologies and has been applied in many fields. Recent gait recognition frameworks represent each gait frame by descriptors extracted from either global appearances or local regions of humans. However, the representations based on global information often neglect the details of the gait frame, while local region based descriptors cannot capture the relations among neighboring regions, thus reducing their discriminativeness. In this paper, we propose a novel feature extraction and fusion framework to achieve discriminative feature representations for gait recognition. Towards this goal, we take advantage of both global visual information and local region details and develop a Global and Local Feature Extractor (GLFE). Specifically, our GLFE module is composed of our newly designed multiple global and local convolutional layers (GLConv) to ensemble global and local features in a principle manner. Furthermore, we present a novel operation, namely Local Temporal Aggregation (LTA), to further preserve the spatial information by reducing the temporal resolution to obtain higher spatial resolution. With the help of our GLFE and LTA, our method significantly improves the discriminativeness of our visual features, thus improving the gait recognition performance. Extensive experiments demonstrate that our proposed method outperforms state-of-the-art gait recognition methods on two popular datasets.
Beibei Lin, Shunli Zhang 0005, Xin Yu 0002
ICCV2
2021 Multi-Scale Temporal Information Extractor For Gait Recognition
abstract
Gait recognition is one of the most important biometric technologies and 3D convolutional neural networks (CNNs) has achieved great success in this field. However, most existing gait recognition frameworks based on 3D CNNs only extract gait features from a single temporal scale, which may not pays enough attention to the gait information in different scales. To solve this problem, we propose a novel multi-scale temporal information extractor to aggregate temporal information from different scales and then represent gait features comprehensively. The small-temporal-scale branch extracts the temporal features from the adjacent frames, which contains the information of slow changes, while the larger-temporal-scale one is used to capture the rapid gait changes. Experiments demonstrate that the proposed method outperforms most existing gait recognition methods on CASIA-B and OutdoorGait datasets.
Beibei Lin, Shunli Zhang 0005, Shengdi Qin
ICIP2
2021 Blind Image Deblurring Based on Dual Attention Network and 2D Blur Kernel Estimation
abstract
In the problem of image deblurring, the restoration of details in severely blurred images has always been difficult. In this paper, we focus on effectively eliminating the ringing artifact and wrinkles that appear after deburring, and propose a novel blind debluring method based on dual attention deep image prior (DADIP) network and 2-dimensional (2D) blur kernel estimation with convolutional neural network (CNN). In the DADIP network, the dual attention mechanism is firstly combined with squeeze and excitation network (SENet), which greatly improves the restoration effect of image details. More importantly, the 2D blur kernel estimation approach via CNN is developed to suppress the ringing artifact of the image, which significantly outperforms previous fully connected network based methods. Experiments show that our deblurring approach achieves superior performance compared with most existing methods.
Senmao Tian, Shunli Zhang 0005, Beibei Lin
ICIP2
2021 AEVRNet: Adaptive exploration network with variance reduced optimization for visual tracking
Yuxiang Yang 0002, Weiwei Xing, Dongdong Wang 0011, Shunli Zhang 0005, Liqiang Wang 0001
Neurocomputing4
2021 Correlation filters with adaptive convolution response fusion for object tracking
Lanlan Yang, Chuihan Kong, Xiaojun Chang, Sicong Zhao, Shunli Zhang 0005
Knowl. Based Syst.6
2020 Gait Recognition with Multiple-Temporal-Scale 3D Convolutional Neural Network
abstract
Gait recognition which is one of the most important and effective biometric technologies has a significant advantage in long-distance recognition systems. For existing gait recognition methods, the template-based approaches may lose temporal information, while the sequence-based methods cannot fully exploit the temporal relations among the sequence. To address the above issues, we propose a novel multiple-temporal-scale gait recognition framework which integrates the temporal information in multiple temporal scales, making use of both the frame and interval fusion information. Moreover, the interval-level representation is realized by a local transformation module. Concretely, 3D convolution neural network (3D CNN) is applied in both the small and the large temporal scales to extract the spatial-temporal information. Moreover, a frame pooling method is developed to address the mismatch of the input of 3D network and video frames, and a novel 3D basic network block is designed to improve efficiency. Experiments demonstrate that the multiple-temporal-scale 3D CNN based gait recognition method can achieve better performance than most recent state-of-the-art methods in CASIA-B dataset. The proposed method obtains the rank-1 accuracy with 96.7% under normal condition, and outperforms other methods on average accuracy by at least 5.8% and 11.1%, respectively, in complex scenarios.
Beibei Lin, Shunli Zhang 0005
ACM Multimedia2
2020 Attention shake siamese network with auxiliary relocation branch for visual object tracking
Jun Wang 0114, Weibin Liu, Weiwei Xing, Liqiang Wang 0001, Shunli Zhang 0005
Neurocomputing5
2020 Learning Scale-Adaptive Tight Correlation Filter for Object Tracking
abstract
In this paper, we propose a novel tracking method by formulating tracking as a correlation filtering as well as a ridge regression problem. First, we develop a tight correlation filter-based tracking framework from the signal detection perspective. In this formulation, the correlation filter is set as the same size as the target, which can make full use of the relations of the adjacent image patches and effectively exclude the influence of the background. Specifically, we point out that the novel correlation filter model can be regarded as the ridge regression model which takes into account the different importance of the samples and has the consistent objective with tracking. Second, we focus on the scale variation problem in tracking. By making use of the spatial structure of the correlation filter, the multiscale filter banks can be generated via interpolation to handle the scale estimation problem easily. Third, we present a novel distance importance-based confidence calculation model to determine the final tracking result, which not only makes use of the fine discriminability of the correlation filter but also takes the distance importance of the candidate samples into account to alleviate the impact of similar distractors. Experimental results demonstrate that our method is superior to several state-of-the-art trackers and many other correlation filter-based methods in the benchmark datasets.
Shunli Zhang 0005, Wei Lu 0010, Weiwei Xing, Li Zhang 0023
IEEE Trans. Cybern.1
2020 Occlusion-Aware Region-Based 3D Pose Tracking of Objects With Temporally Consistent Polar-Based Local Partitioning
abstract
Region-based methods have become the state-of-art solution for monocular 6-DOF object pose tracking in recent years. However, two main challenges still remain: the robustness to heterogeneous configurations (both foreground and background), and the robustness to partial occlusions. In this paper, we propose a novel region-based monocular 3D object pose tracking method to tackle these problems. Firstly, we design a new strategy to define local regions, which is simple yet efficient in constructing discriminative local color histograms. Contrary to previous methods which define multiple circular regions around the object contour, we propose to define multiple overlapped, fan-shaped regions according to polar coordinates. This local region partitioning strategy produces much less number of local regions that need to be maintained and updated, while still being temporally consistent. Secondly, we propose to detect occluded pixels using edge distance and color cues. The proposed occlusion detection strategy is seamlessly integrated into the region-based pose optimization pipeline via a pixel-wise weight function, which significantly alleviates the interferences caused by partial occlusions. We demonstrate the effectiveness of the proposed two new strategies with a careful ablation study. Furthermore, we compare the performance of our method with the most recent state-of-art region-based methods in a recently released large dataset, in which the proposed method achieves competitive results with a higher average tracking success rate. Evaluations on two real-world datasets also show that our method is capable of handling realistic tracking scenarios.
Leisheng Zhong, Yu Zhang 0026, Shunli Zhang 0005, Li Zhang 0023
IEEE Trans. Image Process.4
2020 Toward Precise Osteotomies: A Coarse-to-Fine 3D Cut Plane Planning Method for Image-Guided Pelvis Tumor Resection Surgery
abstract
Surgical resection is the main clinical method for the treatment of bone tumors. A critical procedure for bone tumor resection is to plan a set of cut planes that enable resecting the bone tumor with a safe margin while preserving the maximum amount of healthy bone. Currently, the surgeons rely on manual methods to plan the cut planes, which highly depend on the surgeons' experiences and have been demonstrated to be error-prone, and in turn, increase the recurrence rate or resect much healthy bone. This study targets on improving the precision of cut plane planning for the image guided pelvis tumor resection surgeries. A semi-automatic approach to cut plane planning was proposed via a coarse-to-fine strategy. It can efficiently identify a dangerous region in the 3D space, which contains the bone tumor and its surrounding normal tissue with a safe margin. By projecting the dangerous region into an appropriate 2D space, a segmented boundary-constrained linear regression method was leveraged to plan a set of 3D cut planes that ensure the minimum area of the resected specimen in the 2D space while having the dangerous region cleared. Further, a coarse-to-fine 3D cut plane planning method was developed by incorporating a 3D cut plane refinement scheme with our 2D planning method. Extensive experiments, on the surgical data from nine previous pelvis tumor resection surgeries, demonstrated that our proposed approach substantially improved the localization precision of cut planes ( ) and decreased the amount of resected specimen ( ), as compared to the manual method.
Yu Zhang 0026, Fengzan Li, Lei Qiu 0004, Lihui Xu, Xiaohui Niu, Yao Sui, Shunli Zhang 0005, Li Zhang 0023
IEEE Trans. Medical Imaging7
2020 Fuzzy Least Squares Support Vector Machine With Adaptive Membership for Object Tracking
abstract
Fuzzy learning has been introduced into tracking and achieved great success. However, the membership in the existing fuzzy learning based tracking algorithm is fixed, which lacks the adaptivity to measure the importance of the samples. To improve the tracking adaptivity and flexibility, in this paper, we propose a novel tracking method based on fuzzy least squares support vector machine with adaptive membership (FLS-SVM-AM). First, we formulate tracking as an adaptive membership based fuzzy learning problem, which addresses the issue of fixed membership in existing methods and can better measure the importance of the training samples. Second, we present the FLS-SVM-AM method to build the appearance model, and develop an iterative optimization process to solve the FLS-SVM-AM problem. Third, we define a new membership based on the PASCAL VOC overlap rate and exponential function, which is used to measure the importance of different samples more accurately. Experimental results in the benchmark datasets demonstrate that the proposed method not only outperforms the existing fuzzy learning based tracking methods, but also is comparable to many state-of-the-art methods.
Shunli Zhang 0005, Li Zhang 0023, Alex Hauptmann 0001
IEEE Trans. Multim.1
2019 Residual Multiscale Based Single Image Deraining
Yupei Zheng, Xin Yu 0002, Miaomiao Liu 0001, Shunli Zhang 0005
BMVC4
2019 A framework of tracking by multi-trackers with multi-features in a hybrid cascade way
Jun Wang 0114, Weibin Liu, Weiwei Xing, Shunli Zhang 0005
Signal Process. Image Commun.4
2019 Single Image Depth Estimation With Normal Guided Scale Invariant Deep Convolutional Fields
abstract
Estimating scene depth from a single image can be widely applied to understand 3D environments due to the easy access of the images captured by consumer-level cameras. Previous works exploit conditional random fields (CRFs) to estimate image depth, where neighboring pixels (superpixels) with similar appearances are constrained to share the same depth. However, the depth may vary significantly in the slanted surface, thus leading to severe estimation errors. In order to eliminate those errors, we propose a superpixel-based normal guided scale invariant deep convolutional field by encouraging the neighboring superpixels with similar appearance to lie on the same 3D plane of the scene. In doing so, a depth-normal multitask CNN is introduced to produce the superpixel-wise depth and surface normal predictions simultaneously. To correct the errors of the roughly estimated superpiexl-wise depth, we develop a normal guided scale invariant CRF (NGSI-CRF). NGSI-CRF consists of a scale invariant unary potential that is able to measure the relative depth between superpixels as well as the absolute depth of superpixels, and a normal guided pairwise potential that constrains spatial relationships between superpixels in accordance with the 3D layout of the scene. In other words, the normal guided pairwise potential is designed to smooth the depth prediction without deteriorating the 3D structure of the depth prediction. The superpixel-wise depth maps estimated by NGSI-CRF will be fed into a pixel-wise refinement module to produce a smooth fine-grained depth prediction. Furthermore, we derive a closed-form solution for the maximum a posteriori (MAP) inference of NGSI-CRF. Thus, our proposed network can be efficiently trained in an end-to-end manner. We conduct our experiments on various datasets, such as NYU-D2, KITTI, and Make 3D. As demonstrated in the experimental results, our method achieves superior performance in both indoor and outdoor scenes.
Han Yan 0007, Xin Yu 0002, Yu Zhang 0026, Shunli Zhang 0005, Li Zhang 0023
IEEE Trans. Circuits Syst. Video Technol.4
2018 Monocular depth estimation with guidance of surface normal map
Han Yan 0007, Shunli Zhang 0005, Yu Zhang 0026, Li Zhang 0023
Neurocomputing2
2018 Using fuzzy least squares support vector machine with metric learning for object tracking
Shunli Zhang 0005, Wei Lu 0010, Weiwei Xing, Li Zhang 0023
Pattern Recognit.1
2018 Visual object tracking with multi-scale superpixels and color-feature guided kernelized correlation filters
Jun Wang 0114, Weibin Liu, Weiwei Xing, Shunli Zhang 0005
Signal Process. Image Commun.4
2018 PMSC: PatchMatch-Based Superpixel Cut for Accurate Stereo Matching
abstract
Estimating the disparity and normal direction of one pixel simultaneously, instead of only disparity, also known as 3D label methods, can achieve much higher subpixel accuracy in the stereo matching problem. However, it is extremely difficult to assign an appropriate 3D label to each pixel from the continuous label space R3 while maintaining global consistency because of the infinite parameter space. In this paper, we propose a novel algorithm called PatchMatch-based superpixel cut to assign 3D labels of an image more accurately. In order to achieve robust and precise stereo matching between local windows, we develop a bilayer matching cost, where a bottom-up scheme is exploited to design the two layers. The bottom layer is employed to measure the similarity between small square patches locally by exploiting a pretrained convolutional neural network, and then, the top layer is developed to assemble the local matching costs in large irregular windows induced by the tangent planes of object surfaces. To optimize the spatial smoothness of local assignments, we propose a novel strategy to update 3D labels. In the procedure of optimization, both segmentation information and random refinement of PatchMatch are exploited to update candidate 3D label set for each pixel with high probability of achieving lower loss. Since pairwise energy of general candidate label sets violates the submodular property of graph cut, we propose a novel multilayer superpixel structure to group candidate label sets into candidate assignments, which thereby can be efficiently fused by α-expansion graph cut. Extensive experiments demonstrate that our method can achieve higher subpixel accuracy in different data sets, and currently ranks first on the new challenging Middlebury 3.0 benchmark among all the existing methods.
Lincheng Li, Shunli Zhang 0005, Xin Yu 0002, Li Zhang 0023
IEEE Trans. Circuits Syst. Video Technol.2
2018 Towards Occlusion Handling: Object Tracking With Background Estimation
abstract
The appearance model of the target needs to be updated for online single object tracking. However, the variation of the observation can be caused by active appearance change of the target, or the occlusion from the background. For the former case, we should update the appearance model and for the latter, the current model should be preserved. In this paper, we distinguish these two cases and resist the impact from heavy occlusion by estimating the background in the scene with moving cameras, while retaining the adaptivity to stationary cameras at the same time. The proposed method formulates the background as a Gaussian model and the target is determined in a coarse-to-fine manner. Experimental results demonstrate that our method achieves competitive results in the sequences with appearance changes and outperforms the state-of-the-art algorithms in dealing with complex occlusions.
Sicong Zhao, Shunli Zhang 0005, Li Zhang 0023
IEEE Trans. Cybern.2
2017 Two-level superpixel and feedback based visual object tracking
Jun Wang 0114, Weibin Liu, Weiwei Xing, Shunli Zhang 0005
Neurocomputing4
2017 Graph-Regularized Structured Support Vector Machine for Object Tracking
abstract
How to build a robust and accurate appearance model is a crucial problem in object tracking. However, in most existing tracking methods, the structures among the adjacent video frames and neighboring regions, which may be helpful to improve the representation capability of the appearance model, have not been fully exploited. In this paper, we propose a novel tracking method by taking into account these structures to represent the appearance model. First, we propose a novel graph-regularized structured support vector machine (GS-SVM) algorithm by combining manifold learning and structured learning. Then, the proposed GS-SVM algorithm is employed to build the appearance model and a novel tracking method is developed. This new tracker not only absorbs the advantage of structured learning that deals with the intermediate classification step existing in tracking-by-detection methods, but also exploits the geometry structures in the tracked results and neighboring regions. In addition, a hybrid update strategy is introduced to fit with the GS-SVM-based appearance model. The experimental results demonstrate that the proposed tracking algorithm can outperform several state-of-the-art tracking methods in the benchmark dataset.
Shunli Zhang 0005, Yao Sui, Sicong Zhao, Li Zhang 0023
IEEE Trans. Circuits Syst. Video Technol.1
2016 Object Tracking With Spatial Context Model
abstract
In object tracking, building a reliable appearance model can greatly improve the performance. In this letter, we propose a novel method that uses the spatial context to help tracking based on support vector machines (SVMs). The spatial context is decomposed into different subregions that include some parts of both the target and the background. We build appearance submodels for each group of subregions and combine them to get a robust appearance model, which can help to handle some complex problems, e.g., occlusion and deformation. Besides, we add an update strategy to retain the accuracy of the appearance model. A large number of experiments on various challenging videos demonstrate that our method outperforms many other state-of-the-art methods.
Juntao Sun, Shunli Zhang 0005, Li Zhang 0023
IEEE Signal Process. Lett.2
2015 Self-expressive tracking
Yao Sui, Shunli Zhang 0005, Xin Yu 0002, Sicong Zhao, Li Zhang 0023
Pattern Recognit.3
2015 Hybrid support vector machines for robust object tracking
Shunli Zhang 0005, Yao Sui, Xin Yu 0002, Sicong Zhao, Li Zhang 0023
Pattern Recognit.1
2015 Multi-local-task learning with global regularization for object tracking
Shunli Zhang 0005, Yao Sui, Sicong Zhao, Xin Yu 0002, Li Zhang 0023
Pattern Recognit.1
2015 Robust Visual Tracking via Sparsity-Induced Subspace Learning
abstract
Target representation is a necessary component for a robust tracker. However, during tracking, many complicated factors may make the accumulated errors in the representation significantly large, leading to tracking drift. This paper aims to improve the robustness of target representation to avoid the influence of the accumulated errors, such that the tracker only acquires the information that facilitates tracking and ignores the distractions. We observe that the locally mutual relations between the feature observations of temporally obtained targets are beneficial to the subspace representation in visual tracking. Thus, we propose a novel subspace learning algorithm for visual tracking, which imposes joint row-wise sparsity structure on the target subspace to adaptively exclude distractive information. The sparsity is induced by exploiting the locally mutual relations between the feature observations during learning. To this end, we formulate tracking as a subspace sparsity inducing problem. A large number of experiments on various challenging video sequences demonstrate that our tracker outperforms many other state-of-the-art trackers.
Yao Sui, Shunli Zhang 0005, Li Zhang 0023
IEEE Trans. Image Process.2
2015 Single Object Tracking With Fuzzy Least Squares Support Vector Machine
abstract
Single object tracking, in which a target is often initialized manually in the first frame and then is tracked and located automatically in the subsequent frames, is a hot topic in computer vision. The traditional tracking-by-detection framework, which often formulates tracking as a binary classification problem, has been widely applied and achieved great success in single object tracking. However, there are some potential issues in this formulation. For instance, the boundary between the positive and negative training samples is fuzzy, and the objectives of tracking and classification are inconsistent. In this paper, we attempt to address the above issues from the fuzzy system perspective and propose a novel tracking method by formulating tracking as a fuzzy classification problem. First, we introduce the fuzzy strategy into tracking and propose a novel fuzzy tracking framework, which can measure the importance of the training samples by assigning different memberships to them and offer more strict spatial constraints. Second, we develop a fuzzy least squares support vector machine (FLS-SVM) approach and employ it to implement a concrete tracker. In particular, the primal form, dual form, and kernel form of FLS-SVM are analyzed and the corresponding closed-form solutions are derived for efficient realizations. Besides, a least squares regression model is built to control the update adaptively, retaining the robustness of the appearance model. The experimental results demonstrate that our method can achieve comparable or superior performance to many state-of-the-art methods.
Shunli Zhang 0005, Sicong Zhao, Yao Sui, Li Zhang 0023
IEEE Trans. Image Process.1
2015 Object Tracking With Multi-View Support Vector Machines
abstract
How to build an accurate and reliable appearance model to improve the performance is a crucial problem in object tracking. Since the multi-view learning can lead to more accurate and robust representation of the object, in this paper, we propose a novel tracking method via multi-view learning framework by using multiple support vector machines (SVM). The multi-view SVMs tracking method is constructed based on multiple views of features and a novel combination strategy. To realize a comprehensive representation, we select three different types of features, i.e., gray scale value, histogram of oriented gradients (HOG), and local binary pattern (LBP), to train the corresponding SVMs. These features represent the object from the perspectives of description, detection, and recognition, respectively . In order to realize the combination of the SVMs under the multi-view learning framework, we present a novel collaborative strategy with entropy criterion, which is acquired by the confidence distribution of the candidate samples. In addition, to learn the changes of the object and the scenario, we propose a novel update scheme based on subspace evolution strategy. The new scheme can control the model update adaptively and help to address the occlusion problems . We conduct our approach on several public video sequences and the experimental results demonstrate that our method is robust and accurate, and can achieve the state-of-the-art tracking performance.
Shunli Zhang 0005, Xin Yu 0002, Yao Sui, Sicong Zhao, Li Zhang 0023
IEEE Trans. Multim.1
2014 Efficient Patch-Wise Non-Uniform Deblurring for a Single Image
abstract
In this paper, we address the problem of estimating a latent sharp image from a single spatially variant blurred image. Non-uniform deblurring methods based on projective motion path models formulate the blur as a linear combination of homographic projections of a clear image. But they are computationally expensive and require large memory due to the calculation and storage of a large number of the projections. Patch-wise non-uniform deblurring algorithms have been proposed to estimate each kernel locally by a uniform deblurring algorithm, which does not require to calculate and store the projections. The key issues of these methods are the accuracy of kernel estimation and the identification of erroneous kernels. To perform accurate kernel estimation, we employ the total variation (TV) regularization to recover a latent image, in which the edges are better enhanced and the ringing artifacts are reduced, rather than Tikhonov regularization that previous algorithms adopt. Thus blur kernels can be estimated more accurately from the latent image and estimated in a closed form while previous methods cannot estimate kernels in closed forms. To identify the erroneous kernels, we develop a novel metric, which is able to measure the similarity between the neighboring kernels. After replacing the erroneous kernels with the well-estimated ones, a clear image is obtained. The experiments show that our approach can achieve better results on the real-world blurry images while using less computation and memory.
Xin Yu 0002, Feng Xu 0005, Shunli Zhang 0005, Li Zhang 0023
IEEE Trans. Multim.3