EDBT 2026 Demo / reviewers in the wild / expert
Chengjun Xie
dblp:126/2325
· DBLP profile ↗
32ranked-venue papers
3as first author
23since 2021 · last 2026
0000-0002-0629-2038ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 14 · 2 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A vision-language network for stored-grain pest counting
Rui Li 0027, Chengjun Xie, Peng Chen 0001, Jie Zhang 0033, Jianming Du, Runsheng Qi |
Expert Syst. Appl. | 3 |
| 2026 | Shuffle Mamba: State Space Models With Random Shuffle for Multi-Modal Image FusionabstractMulti-modal image fusion integrates complementary information from different modalities to produce enhanced and informative images. Although State-Space Models, such as Mamba, are proficient in long-range modeling with linear complexity, most Mamba-based approaches use fixed scanning strategies, which can introduce biased prior information. To mitigate this issue, we propose a novel Bayesian-inspired scanning strategy called Random Shuffle, supplemented by a theoretically feasible inverse shuffle to maintain information coordination invariance, aiming to eliminate biases associated with fixed sequence scanning. Based on this transformation pair, we customized the Shuffle Mamba Framework, penetrating modality-aware information representation and cross-modality information interaction across spatial and channel axes to ensure robust interaction and an unbiased global receptive field for multi-modal image fusion. Furthermore, we develop a testing methodology based on Monte-Carlo averaging to ensure the model’s output aligns more closely with expected results. Extensive experiments across multiple multi-modal image fusion tasks demonstrate the effectiveness of our proposed method, yielding excellent fusion quality compared to state-of-the-art alternatives. The code is available at https://github.com/caoke-963/Shuffle-Mamba. Ke Cao 0001, Xuanhua He, Tao Hu 0027, Chengjun Xie, Man Zhou 0003, Jie Zhang 0033 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Freq-RWKV: Granularity-Aware Spatial-Frequency Synergy via Dual-Domain Recurrent Scanning for Pan-sharpening
Xueheng Li, Xuanhua He, Tao Hu 0027, Jie Zhang 0033, Man Zhou 0003, Chengjun Xie, Yingying Wang 0005, Bo Huang 0001 |
ACM Multimedia | 6 |
| 2025 | Frequency Decoupled Domain-Irrelevant Feature Learning for Pan-SharpeningabstractPan-sharpening aims to generate high-detail multi-spectral images (HRMS) through the fusion of panchromatic (PAN) and multi-spectral (MS) images. However, existing pan-sharpening methods often suffer from significant performance degradation when dealing with out-of-distribution data, as they assume the training and test datasets are independent and identically distributed. To overcome this challenge, we propose a novel frequency domain-irrelevant feature learning framework that exhibits exceptional generalization capabilities. Our approach involves parallel extraction and processing of domain-irrelevant information from the amplitude and phase components of the input images. Specifically, we design a frequency information separation module to extract the amplitude and phase components of the paired images. The learnable high-pass filter is then employed to eliminate domain-specific information from the amplitude spectrums. After that, we devised two specialized sub-networks (AFL-Net and PFL-Net) to perform targeted learning of the frequency domain-irrelevant information. This allows our method to effectively capture the complementary domain-irrelevant information contained in the amplitude and phase spectra of the images. Finally, the information fusion and restoration module dynamically adjusts the feature channel weights, enabling the network to output high-quality HRMS images. Through this frequency domain-irrelevant feature learning framework, our method balances generalization capability and network performance on the distribution of training dataset. Extensive experiments conducted on various satellite datasets demonstrate the effectiveness of our method for generalized pan-sharpening. Our proposed network outperforms state-of-the-art methods in terms of both quantitative metrics and visual quality, showcasing its superior ability to handle diverse, out-of-distribution data. Jie Zhang 0033, Ke Cao 0001, Yunlong Lin, Xuanhua He, Yingying Wang 0005, Rui Li 0027, Chengjun Xie, Jun Zhang 0034, Man Zhou 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2025 | Pan-Sharpening via Causal-Aware Feature Distribution CalibrationabstractIn this work, we reveal an interesting observation within the multi-spectral modality: high-frequency components exhibit a long-tailed distribution, in contrast to the Gaussian distribution of dominant low-frequency components. This dual-distribution characteristic presents a challenge for network optimization, leading to overfitting on low-frequency information while neglecting essential high-frequency details. Addressing this issue from a causal inference perspective, we identify optimizer momentum as a confounding factor that biases models towards focusing on the head part of the high-frequency distribution during training. To counteract this effect, we propose a novel optimization strategy and supplement the global-modeling network architecture to balance the frequencies learning. In the training stage, we employ the Recurrent Weighted Key-Value (RWKV) architecture, which features a global receptive field, to effectively learn the long-tailed distribution of high-frequency components and quantify the cumulative direction of feature bias. During the testing stage, we apply counterfactual reasoning to adjust feature distributions based on the quantified bias. To our knowledge, this is the first time to investigate the imbalance of frequency learning within pan-sharpening from the causal inference perspective. Extensive experiments on three benchmark datasets demonstrate that our method significantly outperforms state-of-the-art approaches, showcasing its effectiveness and robustness in pan-sharpening tasks. Xueheng Li, Tao Hu 0027, Ke Cao 0001, Jie Zhang 0033, Chengjun Xie, Man Zhou 0003, Danfeng Hong |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Exploring Text-Guided Information Fusion Through Chain-of-Reasoning for PansharpeningabstractPan-sharpening aims to enhance the spatial resolution of low-resolution multispectral (LRMS) images by integrating high-frequency information from a corresponding texture-rich panchromatic (PAN) image, while maintaining the spectral integrity of the LRMS image. Although text-guided multi-modal learning has made considerable strides in the natural image domain, its potential to pan-sharpening remains underexplored, primarily due to the limited availability of multi-modal remote sensing datasets. To this end, we construct an entirely new pan-sharpening framework by making efforts from three key aspects: (1) text-equipped multi-modal data collection through chain-of-reasoning, (2) large model prior-driven multi-modal information fusion, and (3) visual information interaction through prompt engineering, leveraging textual information to guide the pan-sharpening process within a multi-modal fusion framework. We initially utilize the generic large language model priors to generate descriptive captions for MS images, forming a multi-modal pan-sharpening dataset. By integrating super-resolved imagery and segmentation maps generated by segment anything, we apply Chain-of-Thought (CoT) prompting to generate spatially focused captions across diverse satellite datasets. These captions enhance visual features and provide high-level contextual information, improving semantic understanding for pan-sharpening. Building on the aforementioned multi-modal data, we tailor two text-guided information fusion modules: Textual Enhancement Block (TEB) standing on large model prior and Textual Modulated Block (TMB) utilizing text information to effectively guide and refine the pan-sharpening fusion process. Extensive experiments on multiple satellite datasets demonstrate that our proposed framework outperforms state-of-the-art methods, highlighting its effectiveness and superior performance in pan-sharpening. Xueheng Li, Xuanhua He, Ke Cao 0001, Jie Zhang 0033, Chengjun Xie, Man Zhou 0003, Danfeng Hong, Bo Huang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Enhancing RAW-to-sRGB with Decoupled Style Structure in Fourier DomainabstractRAW to sRGB mapping, which aims to convert RAW images from smartphones into RGB form equivalent to that of Digital Single-Lens Reflex (DSLR) cameras, has become an important area of research. However, current methods often ignore the difference between cell phone RAW images and DSLR camera RGB images, a difference that goes beyond the color matrix and extends to spatial structure due to resolution variations. Recent methods directly rebuild color mapping and spatial structure via shared deep representation, limiting optimal performance. Inspired by Image Signal Processing (ISP) pipeline, which distinguishes image restoration and enhancement, we present a novel Neural ISP framework, named FourierISP. This approach breaks the image down into style and structure within the frequency domain, allowing for independent optimization. FourierISP is comprised of three subnetworks: Phase Enhance Subnet for structural refinement, Amplitude Refine Subnet for color learning, and Color Adaptation Subnet for blending them in a smooth manner. This approach sharpens both color and structure, and extensive evaluations across varied datasets confirm that our approach realizes state-of-the-art results. Code will be available at https://github.com/alexhe101/FourierISP. Xuanhua He, Tao Hu 0027, Guoli Wang 0004, Zejin Wang, Qian Zhang 0009, Rui Li 0027, Chengjun Xie, Jie Zhang 0033, Man Zhou 0003 |
AAAI | 10 |
| 2024 | Frequency-Adaptive Pan-Sharpening with Mixture of ExpertsabstractPan-sharpening involves reconstructing missing high-frequency information in multi-spectral images with low spatial resolution, using a higher-resolution panchromatic image as guidance. Although the inborn connection with frequency domain, existing pan-sharpening research has not almost investigated the potential solution upon frequency domain. To this end, we propose a novel Frequency Adaptive Mixture of Experts (FAME) learning framework for pan-sharpening, which consists of three key components: the Adaptive Frequency Separation Prediction Module, the Sub-Frequency Learning Expert Module, and the Expert Mixture Module. In detail, the first leverages the discrete cosine transform to perform frequency separation by predicting the frequency mask. On the basis of generated mask, the second with low-frequency MOE and high-frequency MOE takes account for enabling the effective low-frequency and high-frequency information reconstruction. Followed by, the final fusion module dynamically weights high frequency and low-frequency MOE knowledge to adapt to remote sensing images with significant content variations. Quantitative and qualitative experiments over multiple datasets demonstrate that our method performs the best against other state-of-the-art ones and comprises a strong generalization ability for real-world scenes. Code will be made publicly at https://github.com/alexhe101/FAME-Net. Xuanhua He, Rui Li 0027, Chengjun Xie, Jie Zhang 0033, Man Zhou 0003 |
AAAI | 4 |
| 2024 | Frequency Decomposition-Driven Network for JPEG Artifacts RemovalabstractJPEG compression, a widely adopted image format, often introduces visual artifacts and quality degradation in image quality. Removal of these JPEG artifacts, especially under high compression rates, proves challenging and typically results in overly smoothed images. This issue primarily arises due to the prevalence of low-frequency regions in natural images. This distribution leads models towards capturing low-frequency information and generating over-smoothing results. To address this issue, we propose the Frequency Decomposition-Driven Network (FDDNet) for JPEG artifact removal. FDDNet incorporates three core modules: the Decomposition Module (DM), inspired by wavelet lifting schemes, extracts both low-frequency and high-frequency components by considering feature channel relationships. The lightweight Low-Frequency Restoration Module and the High-Frequency Refinement Module, are each adept at handling distinct frequency components effectively. By emphasizing high-frequency components, our method surpasses existing approaches in terms of both quantitative metrics and visual quality across various datasets. Ke Cao 0001, Xuanhua He, Tao Hu 0027, Rui Li 0027, Chengjun Xie, Jie Zhang 0033 |
ICME | 6 |
| 2024 | Pan-Sharpening With Wavelet-Enhanced High-Frequency InformationabstractPan-sharpening is essentially a panchromatic (PAN)-guided super-resolution process, primarily focused on enhancing multi-spectral image quality. This methodology intricately incorporates the high-frequency derived from texture-rich PAN images into the lower-resolution multi-spectral (LRMS) counterparts. However, current spatial domain techniques frequently face challenges in accurately restoring texture details, while frequency domain methods lack efficient interaction with spatial domains, thus restricting the overall model performance. In response to these challenges, we introduce a novel High-frequency Wavelet Network that capitalizes on the spatial-frequency interaction and frequency division capabilities inherent in wavelet transform. In particular, our approach consists of two fundamental modules: the Wavelet-Inspired Fusion Block and the High-Frequency Enhancement Block. The former is inspired by wavelet lifting schemes, enabling the fusion of frequencies and facilitating information exchange across various subbands. The latter harnesses wavelet’s frequency division attributes to enhance high-frequency information learning. Comprehensive experiments over multiple satellite datasets demonstrate that our approach outperforms state-of-the-art techniques in both quantitative and qualitative assessments. Moreover, our model showcases exceptional generalization capabilities in real-world scenarios. Code is available at https://github.com/alexhe101/WINet. Jie Zhang 0033, Xuanhua He, Ke Cao 0001, Rui Li 0027, Chengjun Xie, Man Zhou 0003, Danfeng Hong |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Linking Garment with Person via Semantically Associated Landmarks for Virtual Try-OnabstractIn this paper, a novel virtual try-on algorithm, dubbed SAL-VTON, is proposed, which links the garment with the person via semantically associated landmarks to alleviate misalignment. The semantically associated landmarks are a series of landmark pairs with the same local semantics on the in-shop garment image and the try-on image. Based on the semantically associated landmarks, SAL-VTON effectively models the local semantic association between garment and person, making up for the misalignment in the overall deformation of the garment. The outcome is achieved with a three-stage framework: 1) the semantically associated landmarks are estimated using the landmark localization model; 2) taking the landmarks as input, the warping model explicitly associates the corresponding parts of the garment and person for obtaining the local flow, thus refining the alignment in the global flow; 3) finally, a generator consumes the landmarks to better capture local semantics and control the try-on results. Moreover, we propose a new landmark dataset with a unified labelling rule of landmarks for diverse styles of garments. Extensive experimental results on popular datasets demonstrate that SALVTON can handle misalignment and outperform state-of-the-art methods both qualitatively and quantitatively. The dataset is available on https://modelscope.cn/datasets/damo/SAL-HG/summary. Tingwei Gao, Chengjun Xie |
CVPR | 4 |
| 2023 | Pyramid Dual Domain Injection Network for Pan-sharpeningabstractPan-sharpening, a panchromatic image guided low-spatial-resolution multi-spectral super-resolution task, aims to reconstruct the missing high-frequency information of high-resolution multi-spectral counterpart. Although the inborn connection with frequency domain, existing pan-sharpening research has almost investigated the potential solution upon frequency domain, thus limiting the model performance improvement. To this end, we first revisit the degradation process of pan-sharpening in Fourier space, and then devise a Pyramid Dual Domain Injection pan-sharpening Network upon the above observation by fully exploring and exploiting the distinguished information in both the spatial and frequency domains. Specifically, the proposed network is organized with multi-scale U-shape manner and composed by two core parts: a spatial guidance pyramid sub-network for fusing local spatial information and a frequency guidance pyramid sub-network for fusing global frequency domain information, thus encouraging dual-domain complementary learning. In this way, the model can capture multi-scale dual-domain information to enable generating high-quality pan-sharpening results. Quantitative and qualitative experiments over multiple datasets demonstrate that our method performs the best against other state-of-the-art ones and comprises a strong generalization ability for real-world scenes. Xuanhua He, Rui Li 0027, Chengjun Xie, Jie Zhang 0033, Man Zhou 0003 |
ICCV | 4 |
| 2023 | Multiscale Dual-Domain Guidance Network for Pan-SharpeningabstractThe goal of pan-sharpening is to produce a high-spatial-resolution multi-spectral (HRMS) image from a low-spatial-resolution multi-spectral (LRMS) counterpart by super-resolving the LRMS one under the guidance of a texture-rich panchromatic (PAN) image. Existing research has concentrated on using spatial information to generate HRMS images, but has neglected to investigate the frequency domain, which severely restricts the performance improvement. In this work, we propose a novel pan-sharpening approach, named Multi-Scale Dual-Domain Guidance Network (MSDDN) by fully exploring and exploiting the distinguished information in both the spatial and frequency domains. Specifically, the network is inborn with multi-scale U-shape manner and composed by two core parts: a spatial guidance sub-network for fusing local spatial information and a frequency guidance sub-network for fusing global frequency domain information and encouraging dual-domain complementary learning. In this way, the model can capture multi-scale dual-domain information to help it generate high-quality pan-sharpening results. Employing the proposed model on different datasets, the quantitative and qualitative results demonstrate that our method performs appreciatively against other state-of-the-art approaches and comprises a strong generalization ability for real-world scenes. The source code is available at https://github.com/alexhe101/MSDDN. Xuanhua He, Jie Zhang 0033, Rui Li 0027, Chengjun Xie, Man Zhou 0003, Danfeng Hong |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Memory-Augmented Model-Driven Network for Pansharpening
Man Zhou 0003, Li Zhang 0104, Chengjun Xie |
ECCV (19) | 4 |
| 2022 | Panchromatic and Multispectral Image Fusion via Alternating Reverse Filtering NetworkabstractPanchromatic (PAN) and multi-spectral (MS) image fusion, named Pan-sharpening, refers to super-resolve the low-resolution (LR) multi-spectral (MS) images in the spatial domain to generate the expected high-resolution (HR) MS images, conditioning on the corresponding high-resolution PAN images. In this paper, we present a simple yet effective alternating reverse filtering network for pan-sharpening. Inspired by the classical reverse filtering that reverses images to the status before filtering, we formulate pan-sharpening as an alternately iterative reverse filtering process, which fuses LR MS and HR MS in an interpretable manner. Different from existing model-driven methods that require well-designed priors and degradation assumptions, the reverse filtering process avoids the dependency on pre-defined exact priors. To guarantee the stability and convergence of the iterative process via contraction mapping on a metric space, we develop the learnable multi-scale Gaussian kernel module, instead of using specific filters. We demonstrate the theoretical feasibility of such formulations. Extensive experiments on diverse scenes to thoroughly verify the performance of our method, significantly outperforming the state of the arts. Man Zhou 0003, Jie Huang 0017, Feng Zhao 0004, Chengjun Xie, Chongyi Li, Danfeng Hong |
NeurIPS | 5 |
| 2022 | Detection of Pin Defects in Transmission Lines Based on Dynamic Receptive Field
Jianming Du, Shaowei Qian, Chengjun Xie |
PRCV (4) | 4 |
| 2022 | Towards densely clustered tiny pest detection in the wild environment
Jianming Du, Liu Liu 0012, Rui Li 0027, Lin Jiao, Chengjun Xie, Rujing Wang |
Neurocomputing | 5 |
| 2022 | A global activated feature pyramid network for tiny pest detection in the wild
Liu Liu 0012, Rujing Wang, Chengjun Xie, Rui Li 0027, Fangyuan Wang 0001 |
Mach. Vis. Appl. | 3 |
| 2022 | When Pansharpening Meets Graph Convolution Network and Knowledge DistillationabstractIn this article, we propose a novel graph convolutional network (GCN) for pansharpening, defined as GCPNet, which consists of three main modules: the spatial GCN module (SGCN), the spectral band GCN module (BGCN), and the atrous spatial pyramid module (ASPM). Specifically, due to the nature of GCN, the proposed SGCN and BGCN are capable of exploring the long-range relationship between the object and the global state in the spatial and spectral aspects, which benefits pansharpened results and has not been fully investigated before. In addition, the designed ASPM is equipped with multiscale atrous convolutions and learns richer local feature information, so as to cover the objects of different sizes in satellite images. To further enhance the representation of our proposed GCPNet, asynchronous knowledge distillation is introduced to provide compact features by heterogeneous task imitation in a teacher–student paradigm. In the paradigm, the teacher network acts as a variational autoencoder to extract compact features of the ground-truth MS images. The student network, devised for pansharpening, is trained with the assistance of the teacher network to transfer the important information of the expected ground-truth MS images. Extensive experimental results on different satellite datasets demonstrate that our proposed network outperforms the state-of-the-art methods both visually and quantitatively. The source code is released athttps://github.com/Keyu-Yan/GCPNet. Man Zhou 0003, Liu Liu 0012, Chengjun Xie, Danfeng Hong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Recognition and counting of wheat mites in wheat fields by a three-step deep learning method
Peng Chen 0001, Weilu Li, Sijie Yao, Chun Ma, Jun Zhang 0011, Bing Wang 0004, Chun-Hou Zheng 0001, Chengjun Xie |
Neurocomputing | 8 |
| 2021 | ReinforceNet: A reinforcement learning embedded object detection framework with region selection network
Man Zhou 0003, Rujing Wang, Chengjun Xie, Liu Liu 0012, Rui Li 0027, Fangyuan Wang 0001, Dengshan Li |
Neurocomputing | 3 |
| 2021 | Learning region-guided scale-aware feature selection for object detection
Liu Liu 0012, Rujing Wang, Chengjun Xie, Rui Li 0027, Fangyuan Wang 0001, Man Zhou 0003 |
Neural Comput. Appl. | 3 |
| 2021 | Deep Learning Based Automatic Multiclass Wild Pest Monitoring Approach Using Hybrid Global and Local Activated FeaturesabstractSpecialized control of pests and diseases have been a high-priority issue for the agriculture industry in many countries. On account of automation and cost effectiveness, image analytic pest recognition systems are widely utilized in practical crops prevention applications. But due to powerless hand-crafted features, current image analytic approaches achieve low accuracy and poor robustness in practical large-scale multiclass pest detection and recognition. To tackle this problem, this article proposes a novel deep learning based automatic approach using hybrid and local activated features for pest monitoring. In the presented method, we exploit the global information from feature maps to build our global activated feature pyramid network to extract pests' highly discriminative features across various scales over both depth and position levels. It makes changes of depth or spatial sensitive features in pest images more visible during downsampling. Next, an improved pest localization module named local activated region proposal network is proposed to find the precise pest objects positions by augmenting contextualized and attentional information for feature completion and enhancement in local level. The approach is evaluated on our seven-year large-scale pest data-set containing 88.6 K images (16 types of pests) with 582.1 K manually labeled pest objects. The experimental results show that our solution performs over 75.03% mean average precision (mAP) in industrial circumstances, which outweighs two other state-of-the-art methods: Faster R-CNN with mAP up to 70% and feature pyramid network mAP up to 72%. Liu Liu 0012, Chengjun Xie, Rujing Wang, Po Yang 0001, Sud Sudirman, Jie Zhang 0033, Rui Li 0027, Fangyuan Wang 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2020 | C-FCN: Corners-based fully convolutional network for visual object detection
Lin Jiao, Rujing Wang, Chengjun Xie |
Multim. Tools Appl. | 3 |
| 2019 | Deep Learning based Automatic Approach using Hybrid Global and Local Activated Features towards Large-scale Multi-class Pest MonitoringabstractMonitoring pest in agriculture has been a high-priority issue all over the world. Computer vision techniques are widely utilized in practical crop pest prevention applications due to the rapid development of artificial intelligence technology. However, current deep learning image analytic approaches achieve low accuracy and poor robustness in agriculture pest monitoring task. This paper targets at this challenge by proposing a novel two-stage deep learning based automatic pest monitoring system with hybrid global and local activated feature. In this approach, a Global activated Feature Pyramid Network (GaFPN) is firstly proposed for extracting highly representative features of pests over both depth and spatial position activation levels. Then, an improved Local activated Region Proposal Network (LaRPN) augmenting contextual and attentional information is represented for precisely locating pest objects. Finally, we design a fully connected neural network to estimate the severity of input image under the detected pests. The experimental results on our 88.6K images dataset (with 16 types of common pests) show that our approach outweighs the state-of-the-art methods in industrial circumstances. Liu Liu 0012, Rujing Wang, Chengjun Xie, Po Yang 0001, Sud Sudirman, Fangyuan Wang 0001, Rui Li 0027 |
INDIN | 3 |
| 2017 | A novel super-resolution image and video reconstruction approach based on Newton-Thiele's rational kernel in sparse principal component analysis
Lei He 0002, Jieqing Tan, Xing Huo, Chengjun Xie |
Multim. Tools Appl. | 4 |
| 2017 | Robust object tracking via multi-scale patch based sparse coding histogram
Zhongpei Wang, Hao Wang 0008, Jieqing Tan, Peng Chen 0001, Chengjun Xie |
Multim. Tools Appl. | 5 |
| 2015 | Super-resolution by polar Newton-Thiele's rational kernel in centralized sparsity paradigm
Lei He 0002, Jieqing Tan, Zhuo Su 0001, Chengjun Xie |
Signal Process. Image Commun. | 5 |
| 2015 | Corrigendum to "Super-resolution by polar Newton-Thiele's rational kernel in centralized sparsity paradigm" [Signal Processing: Image Communication, 31(2015), pp 86-99]
Lei He 0002, Jieqing Tan, Zhuo Su 0001, Chengjun Xie |
Signal Process. Image Commun. | 5 |
| 2014 | Collaborative object tracking model with local sparse representation
Chengjun Xie, Jieqing Tan, Peng Chen 0001, Jie Zhang 0033, Lei He 0002 |
J. Vis. Commun. Image Represent. | 1 |
| 2014 | Multi-scale patch-based sparse appearance model for robust object tracking
Chengjun Xie, Jieqing Tan, Peng Chen 0001, Jie Zhang 0033, Lei He 0002 |
Mach. Vis. Appl. | 1 |
| 2013 | Multiple instance learning tracking method with local sparse representationabstractWhen objects undergo large pose change, illumination variation or partial occlusion, most existed visual tracking algorithms tend to drift away from targets and even fail in tracking them. To address this issue, in this study, the authors propose an online algorithm by combining multiple instance learning (MIL) and local sparse representation for tracking an object in a video system. The key idea in our method is to model the appearance of an object by local sparse codes that can be formed as training data for the MIL framework. First, local image patches of a target object are represented as sparse codes with an overcomplete dictionary, where the adaptive representation can be helpful in overcoming partial occlusion in object tracking. Then MIL learns the sparse codes by a classifier to discriminate the target from the background. Finally, results from the trained classifier are input into a particle filter framework to sequentially estimate the target state over time in visual tracking. In addition, to decrease the visual drift because of the accumulative errors when updating the dictionary and classifier, a two‐step object tracking method combining a static MIL classifier with a dynamical MIL classifier is proposed. Experiments on some publicly available benchmarks of video sequences show that our proposed tracker is more robust and effective than others. Chengjun Xie, Jieqing Tan, Peng Chen 0001, Jie Zhang 0033, Lei He 0002 |
IET Comput. Vis. | 1 |