Jie Zhang 0033

dblp:84/6889-33 · DBLP profile ↗
← Back
21ranked-venue papers
2as first author
18since 2021 · last 2026
0000-0002-0853-4379ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 7 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 7 since 2021
YearPublicationVenuePosition
2026 RelaCtrl: Relevance-Guided Efficient Control for Diffusion Transformers
abstract
The Diffusion Transformer plays a pivotal role in advancing text-to-image and text-to-video generation, owing primarily to its inherent scalability. However, existing controlled diffusion transformer methods incur significant parameter and computational overheads and suffer from inefficient resource allocation due to their failure to account for the varying relevance of control information across different transformer layers. To address this, we propose the Relevance-Guided Efficient Controllable Generation framework, RelaCtrl, enabling efficient and resource-optimized integration of control signals into the Diffusion Transformer. First, we evaluate the relevance of each layer in the Diffusion Transformer to the control information by assessing the ControlNet Relevance Score, which measures the impact of skipping each control layer on both the quality of generation and the control effectiveness during inference. Based on the strength of the relevance, we then tailor the positioning, parameter scale, and modeling capacity of the control layers to reduce unnecessary parameters and redundant computations. Additionally, to further improve efficiency, we replace the self-attention and FFN in the commonly used copy block with the carefully designed Two-Dimensional Shuffle Mixer (TDSM), enabling efficient implementation of both the token mixer and channel mixer. Both qualitative and quantitative experimental results demonstrate that our approach achieves superior performance with only 15% of the parameters and computational complexity compared to PixArt-delta.
Ke Cao 0001, Jing Wang 0021, Ao Ma 0005, Jiasong Feng, Xuanhua He, Run Ling, Haozhe Wang 0002, Hongjuan Pei, Yihua Shao, Zhanjie Zhang, Jie Zhang 0033
AAAI14
2026 A vision-language network for stored-grain pest counting
Rui Li 0027, Chengjun Xie, Peng Chen 0001, Jie Zhang 0033, Jianming Du, Runsheng Qi
Expert Syst. Appl.5
2026 Shuffle Mamba: State Space Models With Random Shuffle for Multi-Modal Image Fusion
abstract
Multi-modal image fusion integrates complementary information from different modalities to produce enhanced and informative images. Although State-Space Models, such as Mamba, are proficient in long-range modeling with linear complexity, most Mamba-based approaches use fixed scanning strategies, which can introduce biased prior information. To mitigate this issue, we propose a novel Bayesian-inspired scanning strategy called Random Shuffle, supplemented by a theoretically feasible inverse shuffle to maintain information coordination invariance, aiming to eliminate biases associated with fixed sequence scanning. Based on this transformation pair, we customized the Shuffle Mamba Framework, penetrating modality-aware information representation and cross-modality information interaction across spatial and channel axes to ensure robust interaction and an unbiased global receptive field for multi-modal image fusion. Furthermore, we develop a testing methodology based on Monte-Carlo averaging to ensure the model’s output aligns more closely with expected results. Extensive experiments across multiple multi-modal image fusion tasks demonstrate the effectiveness of our proposed method, yielding excellent fusion quality compared to state-of-the-art alternatives. The code is available at https://github.com/caoke-963/Shuffle-Mamba.
Ke Cao 0001, Xuanhua He, Tao Hu 0027, Chengjun Xie, Man Zhou 0003, Jie Zhang 0033
IEEE Trans. Circuits Syst. Video Technol.6
2026 Distilling Textual Priors From LLM to Efficient Image Fusion
abstract
Multi-modality image fusion aims to synthesize a single, comprehensive image from multiple source inputs. Traditional approaches, such as CNNs and GANs, offer efficiency but struggle to handle low-quality or complex inputs. Recent advances in text-guided methods leverage large model priors to overcome these limitations, but at the cost of significant computational overhead, both in memory and inference time. To address this challenge, we propose a novel framework for distilling large model priors, eliminating the need for text guidance during inference while dramatically reducing model size. Our framework utilizes a teacher-student architecture, where the teacher network incorporates large model priors and transfers this knowledge to a smaller student network via a tailored distillation process. Crucially, our experiments demonstrate that this knowledge transfer is the primary driver of performance gains, rather than mere architectural optimization. Additionally, we introduce a spatial-channel cross-fusion module to enhance the model’s ability to leverage textual priors across both spatial and channel dimensions. Our method achieves a favorable trade-off between computational efficiency and fusion quality. The distilled network, requiring only 10% of the parameters and inference time of the teacher network, retains 90% of its performance and outperforms existing SOTA methods. Extensive experiments demonstrate the effectiveness of our approach. Codes are available at https://github.com/Zirconium233/DTPF.
Xuanhua He, Ke Cao 0001, Liu Liu 0012, Li Zhang 0104, Man Zhou 0003, Jie Zhang 0033, Dan Guo 0001, Meng Wang 0001
IEEE Trans. Circuits Syst. Video Technol.7
2025 Freq-RWKV: Granularity-Aware Spatial-Frequency Synergy via Dual-Domain Recurrent Scanning for Pan-sharpening
Xueheng Li, Xuanhua He, Tao Hu 0027, Jie Zhang 0033, Man Zhou 0003, Chengjun Xie, Yingying Wang 0005, Bo Huang 0001
ACM Multimedia4
2025 Frequency Decoupled Domain-Irrelevant Feature Learning for Pan-Sharpening
abstract
Pan-sharpening aims to generate high-detail multi-spectral images (HRMS) through the fusion of panchromatic (PAN) and multi-spectral (MS) images. However, existing pan-sharpening methods often suffer from significant performance degradation when dealing with out-of-distribution data, as they assume the training and test datasets are independent and identically distributed. To overcome this challenge, we propose a novel frequency domain-irrelevant feature learning framework that exhibits exceptional generalization capabilities. Our approach involves parallel extraction and processing of domain-irrelevant information from the amplitude and phase components of the input images. Specifically, we design a frequency information separation module to extract the amplitude and phase components of the paired images. The learnable high-pass filter is then employed to eliminate domain-specific information from the amplitude spectrums. After that, we devised two specialized sub-networks (AFL-Net and PFL-Net) to perform targeted learning of the frequency domain-irrelevant information. This allows our method to effectively capture the complementary domain-irrelevant information contained in the amplitude and phase spectra of the images. Finally, the information fusion and restoration module dynamically adjusts the feature channel weights, enabling the network to output high-quality HRMS images. Through this frequency domain-irrelevant feature learning framework, our method balances generalization capability and network performance on the distribution of training dataset. Extensive experiments conducted on various satellite datasets demonstrate the effectiveness of our method for generalized pan-sharpening. Our proposed network outperforms state-of-the-art methods in terms of both quantitative metrics and visual quality, showcasing its superior ability to handle diverse, out-of-distribution data.
Jie Zhang 0033, Ke Cao 0001, Yunlong Lin, Xuanhua He, Yingying Wang 0005, Rui Li 0027, Chengjun Xie, Jun Zhang 0034, Man Zhou 0003
IEEE Trans. Circuits Syst. Video Technol.1
2025 PanDiT: A Few-Step Diffusion Transformer for High-Fidelity and Efficient Pansharpening
abstract
Pansharpening plays a crucial role in remote sensing by fusing low-resolution multispectral (LRMS) images and high-resolution panchromatic (PAN) images to generate high-resolution multispectral (HRMS) images. While denoising diffusion models offer potential for high-fidelity image generation, their practical application in pansharpening has been severely hindered by huge computational costs from iterative sampling and naive conditioning strategies that struggle to fuse multi-modal information effectively. In this paper, we introduce PanDiT, a novel Diffusion Transformer framework designed to address these challenges. PanDiT is built on the core principle of a Decoupled Conditioning Mechanism, which explicitly disentangles and injects spatial and spectral guidance, and is engineered for practical, few-step inference. Our framework leverages a powerful Diffusion Transformer (DiT) backbone, where conditioning is achieved through two specialized injection blocks that capture spatial and time-frequency features. Crucially, by integrating an implicit sampling strategy, we accelerate the inference process to as few as two steps. Extensive experiments on multiple benchmark datasets demonstrate that PanDiT not only establishes a new state-of-the-art in fusion quality and achieves a good quality-efficiency trade-off. Code is available at https://github.com/para133/PanDIT.
Jiabin Fang, Ke Cao 0001, Xuanhua He, Jie Zhang 0033, Man Zhou 0003, Liu Liu 0012
IEEE Trans. Geosci. Remote. Sens.4
2025 Pan-Sharpening via Causal-Aware Feature Distribution Calibration
abstract
In this work, we reveal an interesting observation within the multi-spectral modality: high-frequency components exhibit a long-tailed distribution, in contrast to the Gaussian distribution of dominant low-frequency components. This dual-distribution characteristic presents a challenge for network optimization, leading to overfitting on low-frequency information while neglecting essential high-frequency details. Addressing this issue from a causal inference perspective, we identify optimizer momentum as a confounding factor that biases models towards focusing on the head part of the high-frequency distribution during training. To counteract this effect, we propose a novel optimization strategy and supplement the global-modeling network architecture to balance the frequencies learning. In the training stage, we employ the Recurrent Weighted Key-Value (RWKV) architecture, which features a global receptive field, to effectively learn the long-tailed distribution of high-frequency components and quantify the cumulative direction of feature bias. During the testing stage, we apply counterfactual reasoning to adjust feature distributions based on the quantified bias. To our knowledge, this is the first time to investigate the imbalance of frequency learning within pan-sharpening from the causal inference perspective. Extensive experiments on three benchmark datasets demonstrate that our method significantly outperforms state-of-the-art approaches, showcasing its effectiveness and robustness in pan-sharpening tasks.
Xueheng Li, Tao Hu 0027, Ke Cao 0001, Jie Zhang 0033, Chengjun Xie, Man Zhou 0003, Danfeng Hong
IEEE Trans. Geosci. Remote. Sens.4
2025 Exploring Text-Guided Information Fusion Through Chain-of-Reasoning for Pansharpening
abstract
Pan-sharpening aims to enhance the spatial resolution of low-resolution multispectral (LRMS) images by integrating high-frequency information from a corresponding texture-rich panchromatic (PAN) image, while maintaining the spectral integrity of the LRMS image. Although text-guided multi-modal learning has made considerable strides in the natural image domain, its potential to pan-sharpening remains underexplored, primarily due to the limited availability of multi-modal remote sensing datasets. To this end, we construct an entirely new pan-sharpening framework by making efforts from three key aspects: (1) text-equipped multi-modal data collection through chain-of-reasoning, (2) large model prior-driven multi-modal information fusion, and (3) visual information interaction through prompt engineering, leveraging textual information to guide the pan-sharpening process within a multi-modal fusion framework. We initially utilize the generic large language model priors to generate descriptive captions for MS images, forming a multi-modal pan-sharpening dataset. By integrating super-resolved imagery and segmentation maps generated by segment anything, we apply Chain-of-Thought (CoT) prompting to generate spatially focused captions across diverse satellite datasets. These captions enhance visual features and provide high-level contextual information, improving semantic understanding for pan-sharpening. Building on the aforementioned multi-modal data, we tailor two text-guided information fusion modules: Textual Enhancement Block (TEB) standing on large model prior and Textual Modulated Block (TMB) utilizing text information to effectively guide and refine the pan-sharpening fusion process. Extensive experiments on multiple satellite datasets demonstrate that our proposed framework outperforms state-of-the-art methods, highlighting its effectiveness and superior performance in pan-sharpening.
Xueheng Li, Xuanhua He, Ke Cao 0001, Jie Zhang 0033, Chengjun Xie, Man Zhou 0003, Danfeng Hong, Bo Huang 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Enhancing RAW-to-sRGB with Decoupled Style Structure in Fourier Domain
abstract
RAW to sRGB mapping, which aims to convert RAW images from smartphones into RGB form equivalent to that of Digital Single-Lens Reflex (DSLR) cameras, has become an important area of research. However, current methods often ignore the difference between cell phone RAW images and DSLR camera RGB images, a difference that goes beyond the color matrix and extends to spatial structure due to resolution variations. Recent methods directly rebuild color mapping and spatial structure via shared deep representation, limiting optimal performance. Inspired by Image Signal Processing (ISP) pipeline, which distinguishes image restoration and enhancement, we present a novel Neural ISP framework, named FourierISP. This approach breaks the image down into style and structure within the frequency domain, allowing for independent optimization. FourierISP is comprised of three subnetworks: Phase Enhance Subnet for structural refinement, Amplitude Refine Subnet for color learning, and Color Adaptation Subnet for blending them in a smooth manner. This approach sharpens both color and structure, and extensive evaluations across varied datasets confirm that our approach realizes state-of-the-art results. Code will be available at https://github.com/alexhe101/FourierISP.
Xuanhua He, Tao Hu 0027, Guoli Wang 0004, Zejin Wang, Qian Zhang 0009, Rui Li 0027, Chengjun Xie, Jie Zhang 0033, Man Zhou 0003
AAAI11
2024 Frequency-Adaptive Pan-Sharpening with Mixture of Experts
abstract
Pan-sharpening involves reconstructing missing high-frequency information in multi-spectral images with low spatial resolution, using a higher-resolution panchromatic image as guidance. Although the inborn connection with frequency domain, existing pan-sharpening research has not almost investigated the potential solution upon frequency domain. To this end, we propose a novel Frequency Adaptive Mixture of Experts (FAME) learning framework for pan-sharpening, which consists of three key components: the Adaptive Frequency Separation Prediction Module, the Sub-Frequency Learning Expert Module, and the Expert Mixture Module. In detail, the first leverages the discrete cosine transform to perform frequency separation by predicting the frequency mask. On the basis of generated mask, the second with low-frequency MOE and high-frequency MOE takes account for enabling the effective low-frequency and high-frequency information reconstruction. Followed by, the final fusion module dynamically weights high frequency and low-frequency MOE knowledge to adapt to remote sensing images with significant content variations. Quantitative and qualitative experiments over multiple datasets demonstrate that our method performs the best against other state-of-the-art ones and comprises a strong generalization ability for real-world scenes. Code will be made publicly at https://github.com/alexhe101/FAME-Net.
Xuanhua He, Rui Li 0027, Chengjun Xie, Jie Zhang 0033, Man Zhou 0003
AAAI5
2024 Frequency Decomposition-Driven Network for JPEG Artifacts Removal
abstract
JPEG compression, a widely adopted image format, often introduces visual artifacts and quality degradation in image quality. Removal of these JPEG artifacts, especially under high compression rates, proves challenging and typically results in overly smoothed images. This issue primarily arises due to the prevalence of low-frequency regions in natural images. This distribution leads models towards capturing low-frequency information and generating over-smoothing results. To address this issue, we propose the Frequency Decomposition-Driven Network (FDDNet) for JPEG artifact removal. FDDNet incorporates three core modules: the Decomposition Module (DM), inspired by wavelet lifting schemes, extracts both low-frequency and high-frequency components by considering feature channel relationships. The lightweight Low-Frequency Restoration Module and the High-Frequency Refinement Module, are each adept at handling distinct frequency components effectively. By emphasizing high-frequency components, our method surpasses existing approaches in terms of both quantitative metrics and visual quality across various datasets.
Ke Cao 0001, Xuanhua He, Tao Hu 0027, Rui Li 0027, Chengjun Xie, Jie Zhang 0033
ICME7
2024 Spatially-Adaptive Large-Kernel Network for Efficient Image Super-Resolution
abstract
In the realm of image super-resolution (SR), efficiency on low-power devices remains a significant challenge due to the high computational demands of current methods. In response to this issue, we propose a novel Spatially-Adaptive Large-Kernel Network(SALKN) tailored for efficient image super-resolution. Inspired by the spectrum convolution theorem, we introduce a spatially-adaptive large kernel convolution unit integrated into a vision transformer architecture. This approach implements dynamic large kernel convolution through an input-adaptive frequency domain multiplication and multi-head mechanism, significantly reducing computational overhead. Our implementation approach, which involves dynamically generating spatially-adaptive frequency filters, enables the dynamic large kernel convolution to be performed with reduced computational costs, facilitating the realization of a global receptive field and promoting scale diversity within the features. Extensive experiments validate that SALKN outperforms existing efficient super-resolution methods with reduced complexity, achieving state-of-the-art performance.
Xuanhua He, Ke Cao 0001, Tao Hu 0027, Jie Zhang 0033, Rui Li 0027
IEEE Signal Process. Lett.4
2024 RSProtoSeg: High Spatial Resolution Remote Sensing Images Segmentation Based on Non-Learnable Prototypes
abstract
Semantic segmentation of high spatial resolution remote sensing images presents unique challenges due to the imbalanced foreground-background distribution and large intra-class variance. This study proposes a novel semantic segmentation algorithm based on non-learnable prototypes, named RSProtoSeg. This approach optimizes the spatial relationship between foreground-background prototypes and intra-class prototypes. Specifically, we propose a foreground-background distance optimization loss function to enhance sparsity between these phototypes, effectively mitigating foreground-background distribution imbalances. Moreover, we introduce an online discrete clustering module that represents each class with a set of prototypes and adds an adaptive regular term penalty to promote sparse structure and reduce the variance issue. Evaluation on three remote sensing datasets (iSAID, ISPRS Potsdam, and Vaihingen) demonstrates significant accuracy improvements, aligning our approach with state-of-the-art methods. Our non-learnable prototype-based approach offers a promising solution for semantic segmentation in high spatial resolution remote sensing images.
Jie Zhang 0033, Yujie Lei, Danfeng Hong
IEEE Trans. Geosci. Remote. Sens.2
2024 Pan-Sharpening With Wavelet-Enhanced High-Frequency Information
abstract
Pan-sharpening is essentially a panchromatic (PAN)-guided super-resolution process, primarily focused on enhancing multi-spectral image quality. This methodology intricately incorporates the high-frequency derived from texture-rich PAN images into the lower-resolution multi-spectral (LRMS) counterparts. However, current spatial domain techniques frequently face challenges in accurately restoring texture details, while frequency domain methods lack efficient interaction with spatial domains, thus restricting the overall model performance. In response to these challenges, we introduce a novel High-frequency Wavelet Network that capitalizes on the spatial-frequency interaction and frequency division capabilities inherent in wavelet transform. In particular, our approach consists of two fundamental modules: the Wavelet-Inspired Fusion Block and the High-Frequency Enhancement Block. The former is inspired by wavelet lifting schemes, enabling the fusion of frequencies and facilitating information exchange across various subbands. The latter harnesses wavelet’s frequency division attributes to enhance high-frequency information learning. Comprehensive experiments over multiple satellite datasets demonstrate that our approach outperforms state-of-the-art techniques in both quantitative and qualitative assessments. Moreover, our model showcases exceptional generalization capabilities in real-world scenarios. Code is available at https://github.com/alexhe101/WINet.
Jie Zhang 0033, Xuanhua He, Ke Cao 0001, Rui Li 0027, Chengjun Xie, Man Zhou 0003, Danfeng Hong
IEEE Trans. Geosci. Remote. Sens.1
2023 Pyramid Dual Domain Injection Network for Pan-sharpening
abstract
Pan-sharpening, a panchromatic image guided low-spatial-resolution multi-spectral super-resolution task, aims to reconstruct the missing high-frequency information of high-resolution multi-spectral counterpart. Although the inborn connection with frequency domain, existing pan-sharpening research has almost investigated the potential solution upon frequency domain, thus limiting the model performance improvement. To this end, we first revisit the degradation process of pan-sharpening in Fourier space, and then devise a Pyramid Dual Domain Injection pan-sharpening Network upon the above observation by fully exploring and exploiting the distinguished information in both the spatial and frequency domains. Specifically, the proposed network is organized with multi-scale U-shape manner and composed by two core parts: a spatial guidance pyramid sub-network for fusing local spatial information and a frequency guidance pyramid sub-network for fusing global frequency domain information, thus encouraging dual-domain complementary learning. In this way, the model can capture multi-scale dual-domain information to enable generating high-quality pan-sharpening results. Quantitative and qualitative experiments over multiple datasets demonstrate that our method performs the best against other state-of-the-art ones and comprises a strong generalization ability for real-world scenes.
Xuanhua He, Rui Li 0027, Chengjun Xie, Jie Zhang 0033, Man Zhou 0003
ICCV5
2023 Multiscale Dual-Domain Guidance Network for Pan-Sharpening
abstract
The goal of pan-sharpening is to produce a high-spatial-resolution multi-spectral (HRMS) image from a low-spatial-resolution multi-spectral (LRMS) counterpart by super-resolving the LRMS one under the guidance of a texture-rich panchromatic (PAN) image. Existing research has concentrated on using spatial information to generate HRMS images, but has neglected to investigate the frequency domain, which severely restricts the performance improvement. In this work, we propose a novel pan-sharpening approach, named Multi-Scale Dual-Domain Guidance Network (MSDDN) by fully exploring and exploiting the distinguished information in both the spatial and frequency domains. Specifically, the network is inborn with multi-scale U-shape manner and composed by two core parts: a spatial guidance sub-network for fusing local spatial information and a frequency guidance sub-network for fusing global frequency domain information and encouraging dual-domain complementary learning. In this way, the model can capture multi-scale dual-domain information to help it generate high-quality pan-sharpening results. Employing the proposed model on different datasets, the quantitative and qualitative results demonstrate that our method performs appreciatively against other state-of-the-art approaches and comprises a strong generalization ability for real-world scenes. The source code is available at https://github.com/alexhe101/MSDDN.
Xuanhua He, Jie Zhang 0033, Rui Li 0027, Chengjun Xie, Man Zhou 0003, Danfeng Hong
IEEE Trans. Geosci. Remote. Sens.3
2021 Deep Learning Based Automatic Multiclass Wild Pest Monitoring Approach Using Hybrid Global and Local Activated Features
abstract
Specialized control of pests and diseases have been a high-priority issue for the agriculture industry in many countries. On account of automation and cost effectiveness, image analytic pest recognition systems are widely utilized in practical crops prevention applications. But due to powerless hand-crafted features, current image analytic approaches achieve low accuracy and poor robustness in practical large-scale multiclass pest detection and recognition. To tackle this problem, this article proposes a novel deep learning based automatic approach using hybrid and local activated features for pest monitoring. In the presented method, we exploit the global information from feature maps to build our global activated feature pyramid network to extract pests' highly discriminative features across various scales over both depth and position levels. It makes changes of depth or spatial sensitive features in pest images more visible during downsampling. Next, an improved pest localization module named local activated region proposal network is proposed to find the precise pest objects positions by augmenting contextualized and attentional information for feature completion and enhancement in local level. The approach is evaluated on our seven-year large-scale pest data-set containing 88.6 K images (16 types of pests) with 582.1 K manually labeled pest objects. The experimental results show that our solution performs over 75.03% mean average precision (mAP) in industrial circumstances, which outweighs two other state-of-the-art methods: Faster R-CNN with mAP up to 70% and feature pyramid network mAP up to 72%.
Liu Liu 0012, Chengjun Xie, Rujing Wang, Po Yang 0001, Sud Sudirman, Jie Zhang 0033, Rui Li 0027, Fangyuan Wang 0001
IEEE Trans. Ind. Informatics6
2014 Collaborative object tracking model with local sparse representation
Chengjun Xie, Jieqing Tan, Peng Chen 0001, Jie Zhang 0033, Lei He 0002
J. Vis. Commun. Image Represent.4
2014 Multi-scale patch-based sparse appearance model for robust object tracking
Chengjun Xie, Jieqing Tan, Peng Chen 0001, Jie Zhang 0033, Lei He 0002
Mach. Vis. Appl.4
2013 Multiple instance learning tracking method with local sparse representation
abstract
When objects undergo large pose change, illumination variation or partial occlusion, most existed visual tracking algorithms tend to drift away from targets and even fail in tracking them. To address this issue, in this study, the authors propose an online algorithm by combining multiple instance learning (MIL) and local sparse representation for tracking an object in a video system. The key idea in our method is to model the appearance of an object by local sparse codes that can be formed as training data for the MIL framework. First, local image patches of a target object are represented as sparse codes with an overcomplete dictionary, where the adaptive representation can be helpful in overcoming partial occlusion in object tracking. Then MIL learns the sparse codes by a classifier to discriminate the target from the background. Finally, results from the trained classifier are input into a particle filter framework to sequentially estimate the target state over time in visual tracking. In addition, to decrease the visual drift because of the accumulative errors when updating the dictionary and classifier, a two‐step object tracking method combining a static MIL classifier with a dynamical MIL classifier is proposed. Experiments on some publicly available benchmarks of video sequences show that our proposed tracker is more robust and effective than others.
Chengjun Xie, Jieqing Tan, Peng Chen 0001, Jie Zhang 0033, Lei He 0002
IET Comput. Vis.4