Jifeng Ning

dblp:83/212 · DBLP profile ↗
← Back
43ranked-venue papers
10as first author
26since 2021 · last 2026
0000-0002-7751-0936ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 8 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 4 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Exploring Modality-Aware Fusion and Decoupled Temporal Propagation for Multi-Modal Object Tracking
Shilei Wang 0001, Pujian Lai, Jifeng Ning, Gong Cheng 0003
AAAI4
2026 Precise weed identification and differentiated laser weeding strategies for Salvia miltiorrhiza fields based on an enhanced object detection network
abstract
Effective weed control is crucial for Salvia miltiorrhiza cultivation, yet traditional methods are often inefficient, costly, or polluting. To address this, this study developed a laser weeding robot based on an improved object detection model capable of identifying weeds and implementing targeted strategies. First, a self-propelled laser weeding robot was constructed for Salvia miltiorrhiza fields to meet operational requirements. Second, a real-world field dataset was established for Salvia miltiorrhiza and five weed families. The detection model, optimized from the You Only Look Once (YOLO) architecture, integrates attention-based feature interaction, dynamic spatial attention, and small object feature enhancement modules. These improvements enhanced the features of small objects, improved occluded target localization, and strengthened similar object discrimination. Third, drawing on weed biological characteristics, a multi-level, differentiated laser weeding strategy was developed to precisely target growth points while ensuring crop safety. Finally, the model and strategy were deployed on the robot to perform real-time detection and intelligent laser weeding. Test results demonstrate the superior performance of the proposed model: the precision of object detection reached 78.09% (2.54% over baseline) and that of keypoint detection stood at 80.69% (8.22% over baseline). The mean average precision ( mAP50 ) metrics improved to 78.14% and 80.56%, representing increases of 2.34% and 2.88% respectively. Field tests achieved a 90.2% weed control rate alongside a low 1.9% damage rate to Salvia miltiorrhiza . These results validate the system's effectiveness and practicality, providing crucial technical support for intelligent weed management in Salvia miltiorrhiza and other high-value medicinal crops.
Xianlin Cao, Jinkai Zhang, Kaidong Liu, Yatuan Ma, Jifeng Ning, Shuqin Yang
Eng. Appl. Artif. Intell.6
2026 ReBaR: Reference-based reasoning for robust pose estimation from monocular images
Yongkang Cheng, Mingjiang Liang, Jifeng Ning, Gaoge Han, Wei Liu 0007, Shaoli Huang
Pattern Recognit.3
2026 Cross-alignment for efficient visual object tracking
Shilei Wang 0001, Mingjiang Liang, Shaoli Huang, Jifeng Ning, Gong Cheng 0003
Pattern Recognit.4
2026 LTSTrack: Visual tracking with long-term temporal sequence
Zhaochuan Zeng, Shilei Wang 0001, Yidong Song, Zhenhua Wang 0003, Jifeng Ning
Pattern Recognit.5
2026 Exploring Pruning-Based Efficient Object Tracking via Hybrid Knowledge Distillation
Yidong Song, Shilei Wang 0001, Zhaochuan Zeng, Jikai Zheng, Zhenhua Wang 0003, Jifeng Ning
IEEE Trans. Circuits Syst. Video Technol.6
2025 DIDiffGes: Decoupled Semi-Implicit Diffusion Models for Real-time Gesture Generation from Speech
abstract
Diffusion models have demonstrated remarkable synthesis quality and diversity in generating co-speech gestures. However, the computationally intensive sampling steps associated with diffusion models hinder their practicality in real-world applications. Hence, we present DIDiffGes, for a Decoupled Semi-Implicit Diffusion model-based framework, that can synthesize high-quality, expressive gestures from speech using only a few sampling steps. Our approach leverages Generative Adversarial Networks (GANs) to enable large-step sampling for diffusion model. We decouple gesture data into body and hands distributions and further decompose them into marginal and conditional distributions. GANs model the marginal distribution implicitly, while L2 reconstruction loss learns the conditional distributions exciplictly. This strategy enhances GAN training stability and ensures expressiveness of generated full-body gestures. Our framework also learns to denoise root noise conditioned on local body representation, guaranteeing stability and realism. DIDiffGes can generate gestures from speech with just 10 sampling steps, without compromising quality and expressiveness, reducing the number of sampling steps by a factor of 100 compared to existing methods. Our user study reveals that our method outperforms state-of-the-art approaches in human likeness, appropriateness, and style correctness.
Yongkang Cheng, Shaoli Huang, Xuelin Chen, Jifeng Ning, Mingming Gong
AAAI4
2025 Conditional GAN for Enhancing Diffusion Models in Efficient and Authentic Global Gesture Generation from Audios
abstract
Audio-driven simultaneous gesture generation is vital for human-computer communication, AI games, and film pro-duction. While previous research has shown promise, there are still limitations. Methods based on VAEs are accompa-nied by issues of local jitter and global instability, whereas methods based on diffusion models are hampered by low generation efficiency. This is because the denoising process of DDPM in the latter relies on the assumption that the noise added at each step is sampled from a unimodal distribution, and the noise values are small. DDIM bor-rows the idea from the Euler method for solving differential equations, disrupts the Markov chain process, and increases the noise step size to reduce the number of denoising steps, thereby accelerating generation. However, simply increasing the step size during the step-by-step denoising process causes the results to gradually deviate from the original data distribution, leading to a significant drop in the quality of the generated actions and the emergence of unnatural artifacts. In this paper, we break the assumptions of DDPM and achieves breakthrough progress in denoising speed and fidelity. Specifically, we introduce a conditional GAN to capture audio control signals and implicitly match the multimodal denoising distribution between the diffusion and denoising steps within the same sampling step, aiming to sample larger noise values and apply fewer denoising steps for high-speed generation. In addition, to enable the model to generate high-fidelity global gestures and avoid artifacts, we introduce an explicit motion geometric loss to enhance the quality and global stability of the generated gestures. Numerous qualitative and quantitative experiments show that compared to contemporary diffusion-based methods, our method offers faster generation speed and higher fidelity, and compared to non-diffusion methods, it provides a more stable global effect and a more natural user experience.
Yongkang Cheng, Mingjiang Liang, Shaoli Huang, Gaoge Han, Jifeng Ning, Wei Liu 0007
WACV5
2024 Exploring the Feature Extraction and Relation Modeling For Light-Weight Transformer Tracking
Jikai Zheng, Mingjiang Liang, Shaoli Huang, Jifeng Ning
ECCV (29)4
2024 Edge-Reserved Knowledge Distillation for Image Matting
abstract
Deep learning-based methods have made significant progress in natural image matting. However, mainstream approaches like CNN and ViT mainly focus on capturing the global features of images, but they lack specialized treatment for edges. They struggle to accurately distinguish between foregrounds and backgrounds that have similar colors and textures, which results in blurred edge areas between the foreground and the background. To solve the above problems, we propose Edge-reserved Knowledge Distillation Model (ERKD), which can reserving good edge features while distill multimodal semantic information from large models. In order to acquire multi-scale edge features, we design the Edge-Reserved Module and the Multimodal Feature Fusion Module. At the same time, to enhance the capture of edge features, we introduce CLIP for feature-level knowledge distillation operations. The comprehensive evaluation on Composition-1k and Distinction-646 datasets shows that the performance of this method surpasses existing techniques.
Zhenhua Wang 0003, Jifeng Ning
ICIP3
2024 ExpGest: Expressive Speaker Generation Using Diffusion Model and Hybrid Audio-Text Guidance
abstract
Existing gesture generation methods primarily focus on upper body gestures based on audio features, neglecting speech content, emotion, and locomotion. These limitations result in stiff, mechanical gestures that fail to convey the true meaning of audio content. We introduce ExpGest, a novel framework leveraging synchronized text and audio information to generate expressive full-body gestures. Unlike AdaIN or one-hot encoding methods, we design a noise emotion classifier for optimizing adversarial direction noise, avoiding melody distortion and guiding results towards specified emotions. Moreover, aligning semantic and gestures in the latent space provides better generalization capabilities. ExpGest, a diffusion model-based gesture generation framework, is the first attempt to offer mixed generation modes, including audio-driven gestures and text-shaped motion. Experiments show that our framework effectively learns from combined text-driven motion and audio-induced gesture datasets, and preliminary results demonstrate that ExpGest achieves more expressive, natural, and controllable global motion in speakers compared to state-of-the-art models.
Yongkang Cheng, Mingjiang Liang, Shaoli Huang, Jifeng Ning, Wei Liu 0007
ICME4
2024 IoUNet++: Spatial cross-layer interaction-based bounding box regression for visual tracking
abstract
Abstract Accurate target prediction, especially bounding box estimation, is a key problem in visual tracking. Many recently proposed trackers adopt the refinement module called IoU predictor by designing a high‐level modulation vector to achieve bounding box estimation. However, due to the lack of spatial information that is important for precise box estimation, this simple one‐dimensional modulation vector has limited refinement representation capability. In this study, a novel IoU predictor (IoUNet++) is designed to achieve more accurate bounding box estimation by investigating spatial matching with a spatial cross‐layer interaction model. Rather than using a one‐dimensional modulation vector to generate representations of the candidate bounding box for overlap prediction, this paper first extracts and fuses multi‐level features of the target to generate template kernel with spatial description capability. Then, when aggregating the features of the template and the search region, the depthwise separable convolution correlation is adopted to preserve the spatial matching between the target feature and candidate feature, which makes their IoUNet++ network have better template representation and better fusion than the original network. The proposed IoUNet++ method with a plug‐and‐play style is applied to a series of strengthened trackers including DiMP++, SuperDiMP++ and SuperDIMP_AR++, which achieve consistent performance gain. Finally, experiments conducted on six popular tracking benchmarks show that their trackers outperformed the state‐of‐the‐art trackers with significantly fewer training epochs.
Shilei Wang 0001, Baozhen Sun, Jifeng Ning
IET Comput. Vis.4
2024 Transformer-based visual object tracking via fine-coarse concatenated attention and cross concatenated MLP
Langkun Chen, Jifeng Ning
Pattern Recognit.6
2024 Local-Global Self-Attention for Transformer-Based Object Tracking
abstract
Transformer-based tracking methods have been widely studied in the field of visual object tracking. The long-range information capturing ability of the transformer improves the performance of the tracking network. However, the self-attention learning procedure in the transformer module neglects the local information, the target and the background around it, which can be beneficial for trackers to handle background clutter and deformation. In this paper, the local-global self-attention (LGSA) learning is proposed for the object tracking task, which obtains the local and global information simultaneously in one attention learning block. Based on the LGSA, the encoder and the decoder are designed to fuse the features corresponding to the template and search images. Additionally, two tracking networks, LGSAT-T and LGSAT-B instantiated with the proposed encoder and decoder are introduced. Exclusive experiments on the commonly used datasets, including OTB100, GOT-10K, LaSOT, and TrackingNet, demonstrate the effectiveness of LGSA, and indicate the state-of-the-art performance of the proposed tracking network. The code will be released athttps://github.com/lgao001/LGSAT.
Langkun Chen, Yunsong Li 0001, Gang He 0002, Jifeng Ning
IEEE Trans. Circuits Syst. Video Technol.6
2024 Bidirectional Interaction of CNN and Transformer Feature for Visual Tracking
abstract
Empowered by the sophisticated long-range dependency modeling ability of Transformer, tracking performance has seen a dynamic increase in recent years. Approaches in this vein leverage the Transformer feature to integrate the information of target and search regions while neglecting the superior local representation extracted by their CNN backbone. To address this, we introduce a BIdirectional inTeraction mechanism between CNN and Transformer features for visual tracking, termed BIT-Tracker, which admits a comprehensive fusion of local and global representations, and thus boosts tracking performance. The first ingredient of BIT-Tracker is an aggregation of multi-level Transformer features to achieve a better global modeling ability. In order to combine the merits of both local and global representations, our second ingredient performs a bi-directional interaction between CNN and Transformer features, where the interaction is achieved via either querying the CNN feature from the Transformer feature or querying the Transformer feature from the CNN feature. Afterwards, the outputs from both directions are fused to predict the temporal locations of targets. Extensive experiments demonstrate the effectiveness of the proposed feature aggregation and bi-directional interaction modules. Impressively, BIT-Tracker achieves leading performance on eight tracking benchmarks and outperforms SOTA results by salient margins. Code will be made available.
Baozhen Sun, Zhenhua Wang 0003, Shilei Wang 0001, Yongkang Cheng, Jifeng Ning
IEEE Trans. Circuits Syst. Video Technol.5
2024 A Spatial Arrangement Preservation-Based Stitching Method via Geographic Coordinates of UAV for Farmland Remote Sensing Image
abstract
High-quality panoramic images captured by UAV for farmland remote sensing play a crucial role in monitoring crop growth. However, the stitching algorithms that rely on local image matching often face challenges due to high similarity in color and texture of farmland remote sensing images. In this work, for the first time we explore the spatial arrangement reflected by these geographic coordinates of each captured image during the UAV flight to obtain the high-quality panoramic images. We theoretically propose a novel algorithm for preserving the spatial arrangement structure based on the triangular similarity transformation, which ensures that the spatial arrangement of all images in the stitched image remains invariant to that before stitching. Moreover, we incorporate the proposed spatial arrangement preservation (SAP) module as an energy term into the Global Similarity Prior (GSP) stitching model and thus obtain the proposed SAP-GSP stitching method. In a wheat-dominated experimental field, we collected UAV remote sensing images at various flight altitudes and multiple growth stages over a three-year period, which formed the image stitching dataset (Wheat-UAV). We evaluated the proposed SAP-GSP algorithm from both subjective and objective perspectives on Wheat-UAV dataset. The experimental results conclusively demonstrate that the SAP-GSP algorithm outperforms representative stitching algorithms in terms of preserving spatial arrangement and achieving natural-looking global stitched images. The proposed SAP-GSP offers a novel and effective technique for robustly capturing large-scale UAV remote sensing panoramic images in agricultural fields.
Shuqin Yang, Jifeng Ning
IEEE Trans. Geosci. Remote. Sens.5
2024 Modeling of Multiple Spatial-Temporal Relations for Robust Visual Object Tracking
abstract
Recently, one-stream trackers have achieved parallel feature extraction and relation modeling through the exploitation of Transformer-based architectures. This design greatly improves the performance of trackers. However, as one-stream trackers often overlook crucial tracking cues beyond the template, they prone to give unsatisfactory results against complex tracking scenarios. To tackle these challenges, we propose a multi-cue single-stream tracker, dubbed MCTrack here, which seamlessly integrates template information, historical trajectory, historical frame, and the search region for synchronized feature extraction and relation modeling. To achieve this, we employ two types of encoders to convert the template, historical frames, search region, and historical trajectory into tokens, which are then collectively fed into a Transformer architecture. To distill temporal and spatial cues, we introduce a novel adaptive update mechanism, which incorporates a thresholding component and a local multi-peak component to filter out less accurate and overly disturbed tracking cues. Empirically, MCTrack achieves leading performance on mainstream benchmark datasets, surpassing the most advanced SeqTrack by 2.0% in terms of the AO metric on GOT-10k. The code is available at https://github.com/wsumel/MCTrack.
Shilei Wang 0001, Zhenhua Wang 0003, Qianqian Sun, Gong Cheng 0003, Jifeng Ning
IEEE Trans. Image Process.5
2023 LoTE-Animal: A Long Time-span Dataset for Endangered Animal Behavior Understanding
abstract
Understanding and analyzing animal behavior is increasingly essential to protect endangered animal species. However, the application of advanced computer vision techniques in this regard is minimal, which boils down to lacking large and diverse datasets for training deep models. To break the deadlock, we present LoTE-Animal, a large-scale endangered animal dataset collected over 12 years, to foster the application of deep learning in rare species conservation. The collected data contains vast variations such as ecological seasons, weather conditions, periods, viewpoints, and habitat scenes. So far, we retrieved at least 500K videos and 1.2 million images. Specifically, we selected and annotated 11 endangered animals for behavior understanding, including 10K video sequences for the action recognition task, 28K images for object detection, instance segmentation, and pose estimation tasks. In addition, we gathered 7K web images of the same species as source domain data for the domain adaptation task. We provide evaluation results of representative vision understanding approaches and cross-domain experiments. LoTE-Animal dataset would facilitate the community to research more advanced machine learning models and explore new tasks to aid endangered animal conservation. Our dataset will be released with the paper. Our dataset can be found at https://LoTE-Animal.github.io
Jin Hou, Shaoli Huang, Bochuan Zheng, Jifeng Ning
ICCV7
2023 Human-to-Human Interaction Detection
Zhenhua Wang 0003, Kaining Ying, Jiajun Meng, Jifeng Ning
ICONIP (4)4
2023 NCSiam: Reliable Matching via Neighborhood Consensus for Siamese-Based Object Tracking
abstract
An essential need for accurate visual object tracking is to capture better correlations between the tracking target and the search region. However, the dominant Siamese-based trackers are limited to producing dense similarity maps at once via a cross-correlations operation, ignoring to remedy the contamination caused by erroneous or ambiguous matches. In this paper, we propose a novel tracker, termed neighborhood consensus constraint-based siamese tracker (NCSiam), which takes the idea of neighborhood consensus constraint to refine the produced correlation maps. The intuition behind our approach is that we can support the nearby erroneous or ambiguous matches by analyzing a larger context of the scene that contains a unique match. Specifically, we devise a 4D convolution-based multi-level similarity refinement (MLSR) strategy. Taking the primary similarity maps obtained from a cross-correlation as input, MLSR acquires reliable matches by analyzing neighborhood consensus patterns in 4D space, thus enhancing the discriminability between the tracking target and the distractors. Besides, traditional Siamese-based trackers directly perform classification and regression on similarity response maps which discard appearance or semantic information. Therefore, an appearance affinity decoder (AAD) is developed to take full advantage of the semantic information of the search region. To further improve performance, we design a task-specific disentanglement (TSD) module to decouple the learned representations into classification-specific and regression-specific embeddings. Extensive experiments are conducted on six challenging benchmarks, including GOT-10k, TrackingNet, LaSOT, UAV123, OTB2015, and VOT2020. The results demonstrate the effectiveness of our method. The code will be available at https://github.com/laybebe/NCSiam.
Pujian Lai, Gong Cheng 0003, Meili Zhang, Jifeng Ning, Xiangtao Zheng, Junwei Han 0001
IEEE Trans. Image Process.4
2022 Geometric Structure Preserving Warp for Natural Image Stitching
abstract
Preserving geometric structures in the scene plays a vital role in image stitching. However, most of the existing methods ignore the large-scale layouts reflected by straight lines or curves, decreasing overall stitching quality. To address this issue, this work presents a structure-preserving stitching approach that produces images with natural visual effects and less distortion. Our method first employs deep learning-based edge detection to extract various types of large-scale edges. Then, the extracted edges are sampled to construct multiple groups of triangles to represent geometric structures. Meanwhile, a GEometric Structure preserving (GES) energy term is introduced to make these triangles undergo similarity transformation. Further, an optimized GES energy term is presented to reasonably determine the weights of the sampling points on the geometric structure, and the term is added into the Global Similarity Prior (GSP) stitching model called GES-GSP to achieve a smooth transition between local alignment and geometric structure preservation. The effectiveness of GES-GSP is validated through comprehensive experiments on a stitching dataset. The experimental results show that the proposed method outperforms several state-of-the-art methods in geometric structure preservation and obtains more natural stitching results. The code and dataset are available at https://github.com/flowerDuo/GES-GSP-Stitching.
Jifeng Ning, Jiguang Cui, Shaoli Huang, Xinchao Wang
CVPR2
2022 Visual object tracking via non-local correlation attention learning
Pan Liu 0009, Jifeng Ning, Yunsong Li 0001
Knowl. Based Syst.3
2022 Multi-criteria Selection of Rehearsal Samples for Continual Learning
Chen Zhuang, Shaoli Huang, Gong Cheng 0003, Jifeng Ning
Pattern Recognit.4
2021 Do We Really Need Frame-by-Frame Annotation Datasets for Object Tracking?
abstract
There has been an increasing emphasis on building large-scale datasets as the driver of deep learning-based trackers' success. However, accurately annotating tracking data is highly labor-intensive and expensive, making it infeasible in real-world applications. In this study, we investigate the necessity of large-scale training data to ensure tracking algorithms' performance. To this end, we introduce a FAT (Few-Annotation Tracking) benchmark constructed by sampling one or a few frames per video from some existing tracking datasets. The proposed dataset can be used to evaluate the effectiveness of tracking algorithms considering data efficiency and new data augmentation approaches for object tracking. We further present AMMC (Augmentation by Mimicking Motion Change), a data augmentation strategy that enables learning high-performing trackers using small-scale datasets. AMMC first cuts out the tracked targets and performs a sequence of transformations to simulate the possible change by object motion. Then the transformed targets are pasted on the inpainted background images and further conjointly augmented to mimic variability caused by camera motion. Compared with standard augmentation methods, AMMC explicitly considers tracking data characteristics, which synthesizes more valid data for object tracking. We extensively evaluate our approach with two popular trackers on the FAT datasets. Experiments show that our method allows these trackers to even trained on a dataset requiring much less annotation to achieve comparable or even better performance to those on the full-annotation dataset. The results imply complete video annotation might not be necessary for object tracking if leveraging motion-driven data augmentations during training.
Shaoli Huang, Shilei Wang 0001, Wei Liu 0007, Jifeng Ning
ACM Multimedia5
2021 Hyperspectral image denoising via global spatial-spectral total variation regularized nonconvex local low-rank tensor approximation
Haijin Zeng, Xiaozhen Xie, Jifeng Ning
Signal Process.3
2021 Hyperspectral Image Restoration via Global L1-2 Spatial-Spectral Total Variation Regularized Local Low-Rank Tensor Recovery
abstract
Hyperspectral images (HSIs) are usually corrupted by various noises, e.g., Gaussian noise, impulse noise, stripes, dead lines, and many others. In this article, motivated by the good performance of the L1-2nonconvex metric in image sparse structure exploitation, we first develop a 3-D L1-2spatial-spectral total variation ( L1-2SSTV) regularization to globally represent the sparse prior in the gradient domain of HSIs. Then, we divide HSIs into local overlapping 3-D patches, and low-rank tensor recovery (LTR) is locally used to effectively separate the low-rank clean HSI patches from complex noise. The patchwise LTR can not only adapt to the local low-rank property of HSIs well but also significantly reduce the information loss caused by the global LTR. Finally, integrating the advantages of both the global L1-2SSTV regularization and local LTR model, we propose a L1-2SSTV regularized local LTR model for hyperspectral restoration. In the framework of the alternating direction method of multipliers, the difference of convex algorithm, the split Bregman iteration method, and tensor singular value decomposition method are adopted to solve the proposed model efficiently. Simulated and real HSI experiments show that the proposed model can reduce the dependence on noise independent and identical distribution hypotheses, and simultaneously remove various types of noise, even structure-related noise.
Haijin Zeng, Xiaozhen Xie, Haojie Cui, Hanping Yin, Jifeng Ning
IEEE Trans. Geosci. Remote. Sens.5
2020 Hyperspectral Image Restoration via Global Total Variation Regularized Local Nonconvex Low-Rank Matrix Approximation
abstract
Several bandwise total variation (TV) regularized low-rank (LR)-based models have been proposed to remove mixed noise in hyperspectral images (HSIs). Conventionally, the rank of LR matrix is approximated using nuclear norm (NN). The NN is defined by adding all singular values together, which is essentially a L1-norm of the singular values. It results in non-negligible approximation errors and thus the resulting matrix estimator can be significantly biased. Moreover, these bandwise TV-based methods exploit the spatial information in a separate manner. To cope with these problems, we propose a spatial-spectral TV (SSTV) regularized non-convex local LR matrix approximation (NonLLRTV) method to remove mixed noise in HSIs. From one aspect, local LR of HSIs is formulated using a non-convex Lγ-norm, which provides a closer approximation to the matrix rank than the traditional NN. From another aspect, HSIs are assumed to be piecewisely smooth in the global spatial domain. The TV regularization is effective in preserving the smoothness and removing Gaussian noise. These facts inspire the integration of the NonLLR with TV regularization. To address the limitations of bandwise TV, we use the SSTV regularization to simultaneously consider global spatial structure and spectral correlation of neighboring bands. Experiment results indicate that the use of local non-convex penalty and global SSTV can boost the preserving of spatial piecewise smoothness and overall structural information.
Haijin Zeng, Xiaozhen Xie, Jifeng Ning
IGARSS3
2020 Hyperspectral image restoration via CNN denoiser prior regularized low-rank tensor recovery
Haijin Zeng, Xiaozhen Xie, Haojie Cui, Yuan Zhao 0016, Jifeng Ning
Comput. Vis. Image Underst.5
2019 Maximum margin object tracking with weighted circulant feature maps
abstract
Support vector machine (SVM) based tracking algorithms training with dense circulant samples have shown favourable performance due to its strong discriminative power and high efficiency. However, the challenges caused by the circulant sampling remain unaddressed. In this study, the authors give each training sample a weight based on their accuracy to reduce the influence of inaccurate samples. Moreover, they reform the SVM model with weighted circulant training samples. Secondly, they advocate an efficient solution by using the property of circulant matrices to solve the learning problem. Thirdly, a model update strategy is introduced to prevent the tracking models polluted by wrong samples. Experimental results on large benchmark datasets with 50 and 100 video sequences demonstrate that the authors’ tracking algorithms achieve state‐of‐art performance in terms of precision and accuracy. In addition, their tracker runs in real time.
Jifeng Ning
IET Comput. Vis.3
2019 Robust correlation filter tracking via context fusion and subspace constraint
Cheng Cai, Jifeng Ning, Yunsong Li 0001
J. Vis. Commun. Image Represent.3
2019 Spatially Regularized Structural Support Vector Machine for Robust Visual Tracking
abstract
Structural support vector machine (SSVM) is popular in the visual tracking field as it provides a consistent target representation for both learning and detection. However, the spatial distribution of feature is not considered in standard SSVM-based trackers, therefore leading to limited performance. To obtain a robust discriminative classifier, this paper proposes a novel tracking framework that spatially regularizes SSVM, which yields a new spatially regularized SSVM (SRSSVM). We utilize the spatial regularization prior to penalize the learning classifier with the same size as the target region. The location of classifier spatially located far from the center of region is assigned large weight and vice versa. Then, it is introduced into the SSVM model as a regularization factor to learn the robust discriminative model. Furthermore, an optimizing algorithm with dual coordination descent is presented to efficiently solve the SRSSVM tracking model. Our proposed SRSSVM tracking method has low computational cost like the traditional linear SSVM tracker while can significantly improve the robustness of the discriminative classifier. The experimental results on three popular tracking benchmark data sets show that the proposed SRSSVM tracking method performs favorably against the state-of-the-art trackers.
Yuhui Zheng, Le Sun 0002, Shunfeng Wang, Jianwei Zhang 0005, Jifeng Ning
IEEE Trans. Neural Networks Learn. Syst.5
2018 Improved kernelized correlation filter tracking by using spatial regularization
Yunsong Li 0001, Jifeng Ning
J. Vis. Commun. Image Represent.3
2016 Object Tracking via Dual Linear Structured SVM and Explicit Feature Map
abstract
Structured support vector machine (SSVM) based methods have demonstrated encouraging performance in recent object tracking benchmarks. However, the complex and expensive optimization limits their deployment in real-world applications. In this paper, we present a simple yet efficient dual linear SSVM (DLSSVM) algorithm to enable fast learning and execution during tracking. By analyzing the dual variables, we propose a primal classifier update formula where the learning step size is computed in closed form. This online learning method significantly improves the robustness of the proposed linear SSVM with lower computational cost. Second, we approximate the intersection kernel for feature representations with an explicit feature map to further improve tracking performance. Finally, we extend the proposed DLSSVM tracker with multi-scale estimation to address the "drift" problem. Experimental results on large benchmark datasets with 50 and 100 video sequences show that the proposed DLSSVM tracking algorithm achieves state-of-the-art performance.
Jifeng Ning, Jimei Yang, Shaojie Jiang, Lei Zhang 0006, Ming-Hsuan Yang 0001
CVPR1
2014 A camera calibration method based on two orthogonal vanishing points
abstract
SUMMARY A camera calibration method based on the two vanishing points, which are orthogonal to each other, is presented in this paper. On the basis of the facts that the modulus constraint can provide an equation and that two equations can be obtained from two pairs of vanishing points, which are orthogonal to each other, the self‐calibration is quasilinearly realized. The experiment results with simulate data and real image sequences show that the presented method is efficient. Copyright © 2013 John Wiley & Sons, Ltd.
Lugang Zhao, Chengke Wu 0001, Jifeng Ning
Concurr. Comput. Pract. Exp.3
2014 Improved appearance updating method in multiple instance learning tracking
abstract
Multiple instance learning (MIL) tracker becomes recently very popular because of their great success in complex scenes. Dynamically reflecting the appearance changes of the tracked object, the appearance updating plays an important role on tracking. In the original MIL tracker, the appearance model is assumed to obey normal distribution and its updating rule consists of a simple linearly weighted sum of the original and the current target distributions in the current frame. However, this updating method is not proved theoretically. In this work, the authors deduce a novel appearance updating method by estimating the mean and the variance of the sum of two normal distributions being merged in maximum likelihood estimation. The method can be naturally extended to multivariable distributions, useful to track colour object. Experimental results on some benchmark video sequences show that the method achieve higher precision and reliability than the three state‐of‐art trackers.
Jifeng Ning, Wuzhen Shi, Shuqin Yang, Paul Yanne
IET Comput. Vis.1
2013 An active contour tracking method by matching foreground and background simultaneously
abstract
This paper presents a novel active contour tracking method, which is used to estimate the non-grid deformation of the motion target and can get the accurate contour of the tracked target. In the proposed method, level set is employed to represent the target region. By using Bhattacharyya similarity as a metric, the proposed method aims at finding a best candidate region in the video frame, whose foreground distribution and background distribution match maximally those of the predefined tracked target. Based on this metric, we derive a level set based object tracking formulation, which estimates iteratively the contour change of the target. Experimental results on the representative video sequences show that the proposed methods perform better than EM-shift and SOAMST algorithm.
Jifeng Ning, Shuqin Yang
ICIP1
2013 Visual tracking based on Distribution Fields and online weighted multiple instance learning
Jifeng Ning, Wuzhen Shi, Shuqin Yang, Paul Yanne
Image Vis. Comput.1
2013 Joint Registration and Active Contour Segmentation for Object Tracking
abstract
This paper presents a novel object tracking framework by joint registration and active contour segmentation (JRACS), which can robustly deal with the non-rigid shape changes of the target. The target region, which includes both foreground and background pixels, is implicitly represented by a level set. A Bhattacharyya similarity based metric is proposed to locate the region whose foreground and background distributions best match those of the tracked target. Based on this metric, a tracking framework that consists of a registration stage and a segmentation stage is then established. The registration step roughly locates the target object by modeling its motion as an affine transformation, and the segmentation step refines the registration result and computes the true contour of the target. The robust tracking performance of the proposed JRACS method is demonstrated by real video sequences where the objects have clear non-rigid shape changes.
Jifeng Ning, Lei Zhang 0006, David Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2012 Automatic tongue image segmentation based on gradient vector flow and region merging
Jifeng Ning, David Zhang 0001, Chengke Wu 0001
Neural Comput. Appl.1
2010 Interactive image segmentation by maximal similarity based region merging
Jifeng Ning, Lei Zhang 0006, David Zhang 0001, Chengke Wu 0001
Pattern Recognit.1
2009 Robust Object Tracking Using Joint Color-Texture Histogram
abstract
A novel object tracking algorithm is presented in this paper by using the joint color-texture histogram to represent a target and then applying it to the mean shift framework. Apart from the conventional color histogram features, the texture features of the object are also extracted by using the local binary pattern (LBP) technique to represent the object. The major uniform LBP patterns are exploited to form a mask for joint color-texture feature selection. Compared with the traditional color histogram based algorithms that use the whole target region for tracking, the proposed algorithm extracts effectively the edge and corner features in the target region, which characterize better and represent more robustly the target. The experimental results validate that the proposed method improves greatly the tracking accuracy and efficiency with fewer mean shift iterations than standard mean shift tracking. It can robustly track the target under complex scenes, such as similar target and background appearance, on which the traditional color based schemes may fail to track.
Jifeng Ning, Lei Zhang 0006, David Zhang 0001, Chengke Wu 0001
Int. J. Pattern Recognit. Artif. Intell.1
2007 NGVF: An improved external force field for active contour model
Jifeng Ning, Chengke Wu 0001, Shigang Liu, Shuqin Yang
Pattern Recognit. Lett.1
2006 A New Active Contour Model: Curvature Gradient Vector Flow
Jifeng Ning, Chengke Wu 0001, Shigang Liu, Peizhi Wen
ACCV (1)1