Xiangyong Cao

dblp:175/1407 · DBLP profile ↗
← Back
67ranked-venue papers
9as first author
55since 2021 · last 2026
0000-0001-7912-3457ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 30 · 4 first-author · 28 since 2021Artificial intelligence and machine learning · 27 · 3 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 4 first-author · 19 since 2021
YearPublicationVenuePosition
2026 DynamicEarth: How Far Are We from Open-Vocabulary Change Detection?
abstract
Monitoring Earth's evolving land covers requires methods capable of detecting changes across a wide range of categories and contexts. Existing change detection methods are hindered by their dependency on predefined classes, reducing their effectiveness in open-world applications. To address this issue, we introduce open-vocabulary change detection (OVCD), a novel task that bridges vision and language to detect changes across any category. Considering the lack of high-quality data and annotation, we propose two training-free frameworks, M-C-I and I-M-C, which leverage and integrate off-the-shelf foundation models for the OVCD task. The insight behind the M-C-I~framework is to discover all potential changes and then classify these changes, while the insight of I-M-C~framework is to identify all targets of interest and then determine whether their states have changed. Based on these two frameworks, we instantiate to obtain several methods, e.g., SAM-DINOv2-SegEarth-OV, Grounding-DINO-SAM2-DINO, etc. Extensive evaluations on 4 benchmark datasets demonstrate the superior generalization and robustness of our OVCD methods over existing supervised and unsupervised methods. To support continued exploration, we release DynamicEarth, a dedicated codebase designed to advance research and application of OVCD.
Kaiyu Li 0001, Xiangyong Cao, Yupeng Deng 0001, Chao Pang 0001, Zepeng Xin, Tieliang Gong, Deyu Meng, Zhi Wang 0002
AAAI2
2026 HSIGene: A Foundation Model for Hyperspectral Image Generation
abstract
Hyperspectral image (HSI) plays a vital role in various fields such as agriculture and environmental monitoring. However, due to the expensive acquisition cost, the number of hyperspectral images is limited, degenerating the performance of downstream tasks. Although some recent studies have attempted to employ diffusion models to synthesize HSIs, they still struggle with the scarcity of HSIs, affecting the reliability and diversity of the generated images. Some studies propose to incorporate multi-modal data to enhance spatial diversity, but spectral fidelity cannot be ensured. In addition, existing HSI synthesis models are typically uncontrollable or only support single-condition control, limiting their ability to generate accurate and reliable HSIs. To alleviate these issues, we propose HSIGene, a novel HSI generation foundation model which is based on latent diffusion and supports multi-condition control, allowing for more precise and reliable HSI generation. To enhance the spatial diversity of the training data while preserving spectral fidelity, we propose a new data augmentation method based on spatial super-resolution, in which HSIs are upscaled first, and thus abundant training patches could be obtained by cropping the high-resolution HSIs. In addition, to improve the perceptual quality of the augmented data, we introduce a novel two-stage HSI super-resolution framework, which first applies RGB bands super-resolution and then utilizes our proposed Rectangular Guided Attention Network (RGAN) for guided HSI super-resolution. Experiments demonstrate that the proposed model is capable of generating a vast quantity of realistic HSIs for downstream tasks such as denoising and super-resolution.
Li Pang, Xiangyong Cao, Datao Tang, Xueru Bai, Feng Zhou 0001, Deyu Meng
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 SegEarth-OV: Towards Training-Free Open-Vocabulary Segmentation for Remote Sensing Images
abstract
Current remote sensing semantic segmentation methods are mostly built on the close-set assumption, meaning that the model can only recognize pre-defined categories that exist in the training set. However, in practical Earth observation, there are countless new categories, and manual annotation is impractical. To address this challenge, we first attempt to introduce training-free1open-vocabulary semantic segmentation (OVSS) into the remote sensing context. However, due to the sensitivity of remote sensing images to low-resolution features, distorted target shapes and ill-fitting boundaries are exhibited in the prediction mask. To tackle these issues, we propose a simple and universal upsampler, i.e. SimFeatUp, to restore lost spatial information of deep features. Specifically, SimFeatUp only needs to learn from a few unlabeled images, and can upsample arbitrary remote sensing image features. Furthermore, based on the observation of the abnormal response
Kaiyu Li 0001, Ruixun Liu, Xiangyong Cao, Xueru Bai, Feng Zhou 0001, Deyu Meng, Zhi Wang 0002
CVPR3
2025 AeroGen: Enhancing Remote Sensing Object Detection with Diffusion-Driven Data Generation
abstract
Remote sensing image object detection (RSIOD) aims to identify and locate specific objects within satellite or aerial imagery. However, there is a scarcity of labeled data in current RSIOD datasets, which significantly limits the performance of current detection algorithms. Although existing techniques, e.g., data augmentation and semi-supervised learning, can mitigate this scarcity issue to some extent, they are heavily dependent on high-quality labeled data and perform worse in rare object classes. To address this issue, this paper proposes a layout-controllable diffusion generative model (i.e. AeroGen) tailored for RSIOD. To our knowledge, AeroGen is the first model to simultaneously support horizontal and rotated bounding box condition generation, thus enabling the generation of high-quality synthetic images that meet specific layout and object category requirements. Additionally, we propose an end-to-end data augmentation framework that integrates a diversity-conditioned generator and a filtering mechanism to enhance both the diversity and quality of generated data. Experimental results demonstrate that the synthetic data produced by our method are of high quality and diversity. Furthermore, the synthetic RSIOD data can significantly improve the detection performance of existing RSIOD models, i.e., the mAP metrics on DIOR, DIOR-R, and HRSC datasets are improved by 3.7%, 4.3%, and 2.43%, respectively. The code is available at here.
Datao Tang, Xiangyong Cao, Jing Yao 0002, Xueru Bai, Dongsheng Jiang, Deyu Meng
CVPR2
2025 Towards Satellite Image Road Graph Extraction: A Global-Scale Dataset and A Novel Method
abstract
Recently, road graph extraction has garnered increasing attention due to its crucial role in autonomous driving, navigation, etc. However, accurately and efficiently extracting road graphs remains a persistent challenge, primarily due to the severe scarcity of labeled data. To address this limitation, we collect a global-scale satellite road graph extraction dataset, i.e. Global-Scale dataset. Specifically, the Global-Scale dataset is ∼ 20× larger than the largest existing public road extraction dataset and spans over 13,800 km2globally. Additionally, we develop a novel road graph extraction model, i.e. SAM-Road++, which adopts a node-guided resampling method to alleviate the mismatch issue between training and inference in SAM-Road [17], a pioneering state-of-the-art road graph extraction model. Furthermore, we propose a simple yet effective "extended-line" strategy in SAM-Road++ to mitigate the occlusion issue on the road. Extensive experiments demonstrate the validity of the collected Global-Scale dataset and the proposed SAM-Road++ method, particularly highlighting its superior predictive power in unseen regions. The dataset and code are available at https://github.com/earth-insights/samroadplus.
Pan Yin, Kaiyu Li 0001, Xiangyong Cao, Jing Yao 0002, Lei Liu 0014, Xueru Bai, Feng Zhou 0001, Deyu Meng
CVPR3
2025 Graph Domain Adaptation With Dual-Branch Encoder and Two-Level Alignment for Whole Slide Image-Based Survival Prediction
Yuntao Shou, Xiangyong Cao, Peiqiang Yan, Qiaohui, Qian Zhao 0002, Deyu Meng
ICCV2
2025 Hipandas: Hyperspectral Image Joint Denoising and Super-Resolution by Image Fusion with the Panchromatic Image
abstract
Hyperspectral images (HSIs) are frequently noisy and of low resolution due to the constraints of imaging devices. Recently launched satellites can concurrently acquire HSIs and panchromatic (PAN) images, enabling the restoration of HSIs to generate clean and high-resolution imagery through fusing PAN images for denoising and super-resolution. However, previous studies treated these two tasks as independent processes, resulting in accumulated errors. This paper introduces \textbf{H}yperspectral \textbf{I}mage Joint \textbf{Pand}enoising \textbf{a}nd Pan\textbf{s}harpening (Hipandas), a novel learning paradigm that reconstructs HRHS images from noisy low-resolution HSIs (LRHS) and high-resolution PAN images. The proposed zero-shot Hipandas framework consists of a guided denoising network, a guided super-resolution network, and a PAN reconstruction network, utilizing an HSI low-rank prior and a newly introduced detail-oriented low-rank prior. The interconnection of these networks complicates the training process, necessitating a two-stage training strategy to ensure effective training. Experimental results on both simulated and real-world datasets indicate that the proposed method surpasses state-of-the-art algorithms, yielding more accurate and visually pleasing HRHS images.
Zixiang Zhao, Haowen Bai, Jiangjun Peng, Xiangyong Cao, Deyu Meng
ICCV6
2025 MoE-based Mamba for Multi-scene Universal Remote Sensing Semantic Segmentation
abstract
Remote sensing semantic segmentation (RSSS) aims to achieve pixel-level classification of remote sensing imagery for land cover identification. However, most existing RSSS methods are tailored for single-scene tasks and lack generalizability across diverse scenes. Extending these models to multi-scene tasks often results in decreased accuracy, increased training time and computational demands. In this paper, we propose MoE-SegMamba, a universal model for efficient multi-scene RSSS. Specifically, we propose an efficient encoder based on the novel TMoESSM Block, which includes a 2D Selective Scan module for capturing global information and a Task-aware Mixture-of-Experts (TMoE) Block to address multi-scene segmentation challenges. To mitigate task interference, we introduce the MoE Guidance Instructions (MGI) module, which generates task-related instructions to assist the TMoE Block in reducing spatial-dimension interference and support the Instruction Channel Gating Module (ICGM) in mitigating channel-dimension interference. Experimental results demonstrate that our proposed MoE-SegMamba achieves State-of-the-Art performance across four semantic segmentation scenes. The code is available at https://github.com/quanquans931225/MoE-SegMamba.
Jie Zhang 0133, Ming-Wen Shao, Xiaodong Tan 0002, Xiangyong Cao
ICME4
2025 Beyond Low-rankness: Guaranteed Matrix Recovery via Modified Nuclear Norm
abstract
The nuclear norm (NN) has been widely explored in matrix recovery problems, such as Robust PCA and matrix completion, leveraging the inherent global low-rank structure of the data. In this study, we introduce a new modified nuclear norm (MNN) framework, where the MNN family norms are defined by adopting suitable transformations and performing the NN on the transformed matrix. The MNN framework offers two main advantages: (1) it jointly captures both local information and global low-rankness without requiring trade-off parameter tuning; (2) under mild assumptions on the transformation, we provide theoretical recovery guarantees for both Robust PCA and MC tasks—an achievement not shared by existing methods that combine local and global information. Thanks to its general and flexible design, MNN can accommodate various proven transformations, enabling a unified and effective approach to structured low-rank recovery. Extensive experiments demonstrate the effectiveness of our method. Code and supplementary material are available at https://github.com/andrew-pengjj/modified_nuclear_norm.
Jiangjun Peng, Yi-Si Luo, Xiangyong Cao, Deyu Meng
IJCAI3
2025 Fast Guaranteed Tensor Recovery with Adaptive Tensor Nuclear Norm
abstract
Real-world datasets like multi-spectral images and videos are naturally represented as tensors. However, limitations in data acquisition often lead to corrupted or incomplete tensor data, making tensor recovery a critical challenge. Solving this problem requires exploiting inherent structural patterns, with the low-rank property being particularly vital. An important category of existing low-rank tensor recovery methods relies on the tensor nuclear norms. However, these methods struggle with either computational inefficiency or weak theoretical guarantees for large-scale data. To address these issues, we propose a fast guaranteed tensor recovery framework based on a new tensor nuclear norm. Our approach adaptively extracts a column-orthogonal matrix from the data, reducing a large-scale tensor into a smaller subspace for efficient processing. This dimensionality reduction enhances speed without compromising accuracy. The recovery theories of two typical models are established by introducing an adjusted incoherence condition. Extensive experiments demonstrate the effectiveness of the proposed method, showing improved accuracy and speed over existing approaches. Our code and supplementary material are available at https://github.com/andrew-pengjj/adaptive_tensor_nuclear_norm.
Jiangjun Peng, Hailin Wang 0001, Xiangyong Cao
IJCAI3
2025 Open-CD: A Comprehensive Toolbox for Change Detection
abstract
We present Open-CD, a change detection toolbox that contains a rich set of change detection methods as well as related components and modules. The toolbox started from a series of open source general vision task tools, including OpenMMLab Toolkits, PyTorch Image Models (Timm), etc. It gradually evolves into a unified platform that covers many popular change detection methods and contemporary modules. It not only includes training and inference codes, but also provides some useful scripts for data analysis. We believe this toolbox is by far the most comprehensive change detection toolbox. In this report, we introduce the features, supported methods and applications of Open-CD. In addition, we also conduct a benchmarking study on different methods and components. We wish that the toolbox and benchmark could serve the growing research community by providing a flexible toolkit to re-implement existing methods and develop their own new change detectors. Code and models are available at https://github.com/likyoo/open-cd.
Kaiyu Li 0001, Chengxi Han, Yupeng Deng 0001, Keyan Chen 0001, Zhuo Zheng, Hao Chen 0045, Ziyuan Liu 0006, Yuantao Gu, Zhengxia Zou, Zhenwei Shi 0001, Sheng Fang 0001, Deyu Meng, Zhi Wang 0002, Xiangyong Cao
ACM Multimedia15
2025 Parameterized Low-Rank Regularizer for High-dimensional Visual Data
Zixiang Zhao, Xiangyong Cao, Jiangjun Peng, Xi-Le Zhao, Deyu Meng, Yulun Zhang 0001, Radu Timofte, Luc Van Gool
Int. J. Comput. Vis.3
2025 Contrastive Graph Representation Learning with Adversarial Cross-View Reconstruction and Information Bottleneck
Yuntao Shou, Haozhi Lan, Xiangyong Cao
Neural Networks3
2025 Masked contrastive graph representation learning for age estimation
Yuntao Shou, Xiangyong Cao, Huan Liu 0012, Deyu Meng
Pattern Recognit.2
2025 A Low-Rank Matching Attention Based Cross-Modal Feature Fusion Method for Conversational Emotion Recognition
abstract
Conversational emotion recognition (CER) is an important research topic in human-computer interactions. Although recent advancements in transformer-based cross-modal fusion methods have shown promise in CER tasks, they tend to overlook the crucial intra-modal and inter-modal emotional interaction or suffer from high computational complexity. To address this, we introduce a novel and lightweight cross-modal feature fusion method called Low-Rank Matching Attention Method (LMAM). LMAM effectively captures contextual emotional semantic information in conversations while mitigating the quadratic complexity issue caused by the self-attention mechanism. Specifically, by setting a matching weight and calculating inter-modal features attention scores row by row, LMAM requires only one-third of the parameters of self-attention methods. We also employ the low-rank decomposition method on the weights to further reduce the number of parameters in LMAM. As a result, LMAM offers a lightweight model while avoiding overfitting problems caused by a large number of parameters. Moreover, LMAM is able to fully exploit the intra-modal emotional contextual information within each modality and integrates complementary emotional semantic information across modalities by computing and fusing similarities of intra-modal and inter-modal features simultaneously. Experimental results verify the superiority of LMAM compared with other popular cross-modal fusion methods on the premise of being more lightweight. Also, LMAM can be embedded into any existing state-of-the-art CER methods in a plug-and-play manner, and can be applied to other multi-modal recognition tasks, e.g., session recommendation and humour detection, demonstrating its remarkable generalization ability.
Yuntao Shou, Huan Liu 0012, Xiangyong Cao, Deyu Meng, Bo Dong 0001
IEEE Trans. Affect. Comput.3
2025 SemiCD-VL: Visual-Language Model Guidance Makes Better Semi-Supervised Change Detector
abstract
Change detection (CD) aims to identify pixels with semantic changes between images. However, annotating massive numbers of pixel-level images is labor-intensive and costly, especially for multitemporal images, which require pixel-wise comparisons by human experts. Considering the excellent performance of visual-language models (VLMs) for zero-shot, OV, etc., with prompt-based reasoning, it is promising to utilize VLMs to make better CD under limited labeled data. In this article, we propose a VLM guidance-based semi-supervised CD method, namely SemiCD-VL. The insight of SemiCD-VL is to synthesize free change labels using VLMs to provide additional supervision signals for unlabeled data. However, almost all current VLMs are designed for single-temporal images and cannot be directly applied to bi- or multitemporal images. Motivated by this, we first propose a VLM-based mixed change event generation (CEG) strategy to yield pseudo-labels for unlabeled CD data. Since the additional supervised signals provided by these VLM-driven pseudo-labels may conflict with the original pseudo-labels from the consistency regularization paradigm (e.g., FixMatch), we propose the dual projection head for de-entangling different signal sources. Further, we explicitly decouple the bitemporal images semantic representation through two auxiliary segmentation decoders, which are also guided by VLM. Finally, to make the model more adequately capture change representations, we introduce contrastive consistency regularization (CCR) by constructing feature-level contrastive loss in auxiliary branches. Extensive experiments show the advantage of SemiCD-VL. For instance, SemiCD-VL improves the FixMatch baseline by$+ 5.3~\text {IoU}^{c}$on WHU-CD and by$+ 2.4~\text {IoU}^{c}$on LEVIR-CD with 5% labels, and SemiCD-VL requires only 5%–10% of the labels to achieve performance similar to the supervised methods. In addition, our CEG strategy, in an unsupervised manner, can achieve performance far superior to state-of-the-art (SOTA) unsupervised CD methods (e.g., IoU improved from 18.8% to 46.3% on LEVIR-CD dataset). The code is available athttps://github.com/likyoo/SemiCD-VL.
Kaiyu Li 0001, Xiangyong Cao, Yupeng Deng 0001, Junmin Liu, Deyu Meng, Zhi Wang 0002
IEEE Trans. Geosci. Remote. Sens.2
2025 CTVNet: Gradient Prior-Guided Deep Unfolding Network for Infrared Small Target Detection
abstract
For infrared small target detection tasks, deep unfolding techniques have demonstrated effectiveness and practical value. However, existing methods generally emphasize the low-rankness of background and the sparsity of targets within the robust principal component analysis (RPCA) framework, which may overlook the intrinsic gradient prior information existed in background. To address the challenges of complex background estimation and accurate small target detection, we propose a gradient prior-guided deep unfolding network, termed the correlated total variation network (CTVNet). First, we introduce a correlated total variation regularization to simultaneously characterize the low-rankness and local smoothness of the background, and transform it into the estimation of gradient maps. Subsequently, we employ a multi-scale feature fusion network to thoroughly extract gradient priors, replacing the complex and limited analytical computation of gradient correlations. Finally, we unfold the designed iterative algorithm using alternating direction method of multipliers (ADMM) into a learnable network, where each module corresponds to a specific operator within the iterative process, and all parameters are learnable. By training the network end-to-end, the learnable modules can be automatically optimized to better separate the background and the target. Extensive experimental results demonstrate that our proposed method achieves competitive performance compared to several state-of-the-art algorithms while exhibiting superior performance and generalization capabilities on both in-distribution and out-of-distribution data. Our code is available at https://github.com/AuroraPei/CTVNet.
Li Pang, Jiangjun Peng, Yi-Si Luo, Junmin Liu, Xiangyong Cao
IEEE Trans. Geosci. Remote. Sens.6
2025 DGAT: Dynamic Gaussian Attenuate Transformer for Remote Sensing Image Change Captioning
abstract
TheRemote Sensing Image Change Captioning(RSICC) technique is designed to enhance geospatial analysis by generating semantic descriptions of differences observed in bi-temporal remote sensing imagery. Although Transformer-based methods have achieved significant advancements in this field, their standard global attention mechanism allows pixels to pay equal attention to all spatial positions, which lacks explicit prior knowledge of spatial proximity correlation, making the model unable to strengthen local associations and weaken long-distance connections. Furthermore, most existing methods rely on global average pooling when reinforcing channel-wise representations, which often dilutes the features of key changed regions by the largely unchanged background, resulting in the inability of feature compression to focus on the critical positions. To address the aforementioned challenges, this paper proposes aDynamic Gaussian Attenuate Transformer(DGAT), which innovatively introduces aDynamic Gaussian Attenuation(DGA) mechanism to model an attenuation law between visual tokens from two perspectives based on Euclidean distance. At the spatial level, DGA constrains visual attention through distance-related Gaussian attenuation, allowing it to prioritize attention on adjacent regions, thereby enhancing the detection of changes in local continuity; at the channel level, DGA identifies the core area based on the visual attention kernel, and uses the joint optimization of dynamic Gaussian weighted pooling and channel modulation to focus on features at key positions, effectively enhancing the expression of important channels. Experimental results on three benchmark RSICC datasets demonstrate that our proposed DGAT achieves significantly superior results, verifying the effectiveness of the DGA mechanism. Our code and model weights are available at https://github.com/Pengfei1005/DGAT.
Pengfei Qin, Junmin Liu, Lanyu Li, Xiangyong Cao
IEEE Trans. Geosci. Remote. Sens.5
2025 Hyperspectral Anomaly Detection Fused Unified Nonconvex Tensor Ring Factors Regularization
abstract
In recent years, tensor decomposition-based approaches forhyperspectral anomaly detection(HAD) have gained significant attention in the field of remote sensing. However, existing methods often fail to flexibly and effectively extract both the global correlations and local smoothness of the background components inhyperspectral images(HSIs). To mitigate this critical issue, we put forward a novel HAD method named HAD-EUNTRFR, which incorporates an enhanced unified nonconvex tensor ring (TR) factors regularization. In the HAD-EUNTRFR framework, the raw HSIs are first decomposed into background and anomaly components using the idea of tensor robust principal component analysis. The TR decomposition is then employed to capture the spatial-spectral correlations within the background component. Additionally, we introduce a unified and efficient nonconvex regularizer, induced bytensor singular value decomposition(T-SVD), to simultaneously encode the low-rankness and sparsity of the 3-D gradient TR factors into a unique concise form. The above characterization scheme enables the interpretable gradient TR factors to inherit the low-rankness and smoothness of the original background. To further enhance anomaly detection, we design a generalized nonconvex regularization term to exploit the group sparsity of the anomaly component. Based upon the above, we ultimately propose a scalable and reliable nonconvex HAD model. To solve the resulting doubly nonconvex model, we develop a highly efficient optimization algorithm based on thealternating direction method of multipliers(ADMM) framework. Theoretical results on convergence analysis for the proposed algorithm are derived. Experimental results on several benchmark datasets demonstrate that our proposed method outperforms existingstate-of-the-art(SOTA) approaches in terms of detection accuracy.
Wenjin Qin, Hailin Wang 0001, Feng Zhang 0023, Jianjun Wang 0003, Xiangyong Cao, Xi-Le Zhao, Gemine Vivone
IEEE Trans. Geosci. Remote. Sens.6
2025 Mask Approximation Net: A Novel Diffusion Model Approach for Remote Sensing Change Captioning
abstract
Remote sensing (RS) image change description represents an innovative multimodal task within the realm of RS processing. This task not only facilitates the detection of alterations in surface conditions but also provides comprehensive descriptions of these changes, thereby improving human interpretability and interactivity. Current deep learning methods typically adopt a three-stage framework consisting of feature extraction, feature fusion, and change localization, followed by text generation. Most approaches focus heavily on designing complex network modules but lack solid theoretical guidance, relying instead on extensive empirical experimentation and iterative tuning of network components. This experience-driven design paradigm may lead to overfitting and design bottlenecks, thereby limiting the model’s generalizability and adaptability. To address these limitations, this article proposes a paradigm that shifts toward data distribution learning using diffusion models, reinforced by frequency-domain noise filtering, to provide a theoretically motivated and practically effective solution to multimodal RS change description. The proposed method primarily includes a simple multiscale change detection (CD) module, whose output features are subsequently refined by a well-designed diffusion model. Furthermore, we introduce a frequency-guided complex filter module to boost the model’s performance by managing high-frequency noise throughout the diffusion process. We validate the effectiveness of our proposed method across several datasets for RS CD and description, showcasing its superior performance compared to existing techniques. The code will be available athttps://github.com/sundongweiMaskApproxNet
Dongwei Sun, Jing Yao 0002, Wu Xue, Changsheng Zhou, Pedram Ghamisi, Xiangyong Cao
IEEE Trans. Geosci. Remote. Sens.6
2025 PromptSeg: Prompt for Universal Remote Sensing Semantic Segmentation
Jie Zhang 0133, Ming-Wen Shao, Lingzhuang Meng, Xiangyong Cao, Shuigen Wang
IEEE Trans. Geosci. Remote. Sens.4
2025 Haar Nuclear Norms With Applications to Remote Sensing Imagery Restoration
abstract
Remote sensing image restoration, which aims to reconstruct corrupted or missing regions, heavily relies on low-rank models. A recent trend in this field is to jointly model low-rank and local smoothness priors using a single regularization term, in order to better recover fine textures. However, due to the entanglement of low- and high-frequency components in an image, existing methods often struggle to simultaneously capture both coarse-grained structures and fine-grained textures, while also suffering from high computational complexity. To address these issues, this paper proposes a novel regularization, the Haar Nuclear Norm (HNN), for efficient and effective remote sensing image restoration. HNN transforms images into wavelet coefficients that separate low-frequency (coarse-grained) and high-frequency (fine-grained) components, and enforces low-rankness via nuclear norms on the mode-3 unfolding matrices of these wavelet coefficients. Experimental evaluations conducted on hyperspectral image inpainting, multi-temporal image cloud removal, and hyperspectral image denoising have revealed the HNN's potential. Typically, HNN achieves a performance improvement of 1-4 dB and a speedup of 10-28x compared to some state-of-the-art methods (e.g., tensor correlated total variation, and fully-connected tensor network) for inpainting tasks. The code is available at https://github.com/isyuchang/HNN.
Jiangjun Peng, Shichao Chen, Xiangyong Cao, Deyu Meng
IEEE Trans. Image Process.5
2025 Polar R-CNN: End-to-End Lane Detection With Fewer Anchors
abstract
Lane detection is a critical and challenging task in autonomous driving, particularly in real-world scenarios where traffic lanes can be slender, lengthy, and often obscured by other vehicles, complicating detection efforts. Existing anchor-based methods typically rely on prior lane anchors to extract features and subsequently refine the location and shape of lanes. While these methods achieve high performance, manually setting prior anchors is cumbersome, and ensuring sufficient coverage across diverse datasets often requires a large amount of dense anchors. Furthermore, the use ofNon-Maximum Suppression(NMS) to eliminate redundant predictions complicates real-world deployment and may under-perform in complex scenarios. In this paper, we proposePolar R-CNN, an end-to-end anchor-based method for lane detection. By incorporating both local and global polar coordinate systems, Polar R-CNN facilitates flexible anchor proposals and significantly reduces the number of anchors required without compromising performance. Additionally, we introduce a triplet head with heuristic structure that supports NMS-free paradigm, enhancing deployment efficiency and performance in scenarios with dense lanes. Our method achieves competitive results on five popular lane detection benchmarks—Tusimple,CULane,LLAMAS,CurveLanes, andDL-Rail—while maintaining a lightweight design and straightforward structure. Our source code is available athttps://github.com/ShqWW/PolarRCNN
Shengqi Wang, Junmin Liu, Xiangyong Cao, Zengjie Song, Kai Sun 0007
IEEE Trans. Intell. Transp. Syst.3
2025 Multi-Scale Retinex Unfolding Network for Low-Light Image Enhancement
abstract
Retinex theory-based low-light image enhancement methods have received increasing attention and achieved tremendous advancements. However, there still exist two seldom-explored issues: 1) The above methods only formally simulate the Retinex decomposition, resulting in lacking explicit interpretability. 2) They usually are performed in single-scale space, leading to suboptimal enhancement results. In this paper, we propose an interpretable Multi-scale Retinex Unfolding Network (MRUNet) for low-light image enhancement, which can tackle both of the aforementioned issues simultaneously. Specifically, we formulate low-light image enhancement as a multi-scale Retinex optimization problem and design an iteration minimization solution to solve it. The optimization solution is further unfolded to fabricate MRUNet, which is empowered with clear physical significance and multi-scale prior knowledge in favor of image enhancement. However, it will aggravate model size and efficiency when exploiting multiple proximal mapping networks to extract multi-scale prior from multi-scale inputs. To surmount the issue, we propose a Scale-Aware Proximal mapping Module (SAPM), which efficiently collect multi-scale prior knowledge via the weight sharing strategy. In SAPM, we tailor a scale-aware transformer to model the specific scale-similarity among different scales. Extensive experiments manifest that MRUNet surpasses other Retinex-based low-light image enhancement methods on multiple benchmarks.
Huake Wang, Xingsong Hou, Jutao Li, Yadi Yan, Wenke Sun, Kaibing Zhang, Xiangyong Cao
IEEE Trans. Multim.8
2025 SpeGCL: Self-Supervised Graph Spectrum Contrastive Learning Without Positive Samples
abstract
Graph contrastive learning (GCL) has emerged as a powerful method for dealing with noise and fluctuations in graph-structured data, and can be applied to social networks and knowledge graphs. Although various graph augmentation strategies have emerged in the field of GCL, traditional graph convolutional network (GCN) mainly tends to preserve smooth features and has difficulty capturing fine-grained changes between different views. To address the above issue, we first construct Fourier graph neural network (FourierGNN) from the perspective of graph spectrum learning, which captures different frequency components by stacking multiple Fourier graph operations (FGO) layers in Fourier space. Then, we find that the difference between the high-frequency information of two augmented graphs should be larger than the difference between the low-frequency information. Next, we theoretically prove that focusing only on pushing negative pairs farther away can more effectively achieve performance advantages. By leveraging these discoveries, we propose a novel self-supervised graph spectrum contrastive learning framework, i.e., SpeGCL, and design an effective contrastive strategy to optimize this goal. We also provide a theoretical justification for the efficacy of using only negative samples in SpeGCL. Extensive experiments have been conducted on unsupervised, transfer, and semi-supervised learning tasks to show that SpeGCL outperforms existing state-of-the-art (SOTA) GCL methods.
Yuntao Shou, Xiangyong Cao, Deyu Meng
IEEE Trans. Neural Networks Learn. Syst.2
2024 HIR-Diff: Unsupervised Hyperspectral Image Restoration Via Improved Diffusion Models
abstract
Hyperspectral image (HSI) restoration aims at recov-ering clean images from degraded observations and plays a vital role in downstream tasks. Existing model-based methods have limitations in accurately modeling the com-plex image characteristics with handcraft priors, and deep learning-based methods suffer from poor generalization ability. To alleviate these issues, this paper proposes an unsupervised HSI restoration framework with pre-trained diffusion model (HIR-Diff), which restores the clean HSls from the product of two low-rank components, i.e., the re-duced image and the coefficient matrix. Specifically, the re-duced image, which has a low spectral dimension, lies in the image field and can be inferred from our improved diffusion model where a new guidance function with total variation (TV) prior is designed to ensure that the reduced image can be well sampled. The coefficient matrix can be effectively pre-estimated based on singular value decomposition (SVD) and rank-revealing QR (RRQR) factorization. Fur-thermore, a novel exponential noise schedule is proposed to accelerate the restoration process (about 5 x acceleration for denoising) with little performance decrease. Ex-tensive experimental results validate the superiority of our method in both performance and speed on a variety of HSI restoration tasks, including HSI denoising, noisy HSI super-resolution, and noisy HSI inpainting. The code is available at https://github.com/LiPang/HIRDiff.
Li Pang, Xiangyu Rui, Long Cui, Hongzhong Wang, Deyu Meng, Xiangyong Cao
CVPR6
2024 Stable Local-Smooth Principal Component Pursuit
abstract
Abstract. Recently, the CTV-RPCA model proposed the first recoverable theory for separating low-rank and local-smooth matrices and sparse matrices based on the correlated total variation (CTV) regularizer. However, the CTV-RPCA model ignores the influence of noise, which makes the model unable to effectively extract low-rank and local-smooth principal components under noisy circumstances. To alleviate this issue, this article extends the CTV-RPCA model by considering the influence of noise and proposes two robust models with parameter adaptive adjustment, i.e., Stable Principal Component Pursuit based on CTV (CTV-SPCP) and Square Root Principal Component Pursuit based on CTV (CTV-[Formula: see text]). Furthermore, we present a statistical recoverable error bound for the proposed models, which allows us to know the relationship between the solution of the proposed models and the ground-truth. It is worth mentioning that, in the absence of noise, our theory degenerates back to the exact recoverable theory of the CTV-RPCA model. Finally, we develop the effective algorithms with the strict convergence guarantees. Extensive experiments adequately validate the theoretical assertions and also demonstrate the superiority of the proposed models over many state-of-the-art methods on various typical applications, including video foreground extraction, multispectral image denoising, and hyperspectral image denoising. The source code is released at https://github.com/andrew-pengjj/CTV-SPCP .
Jiangjun Peng, Hailin Wang 0001, Xiangyong Cao, Xixi Jia, Hong-Ying Zhang 0001, Deyu Meng
SIAM J. Imaging Sci.3
2024 A New Learning Paradigm for Foundation Model-Based Remote-Sensing Change Detection
abstract
Change detection (CD) is a critical task to observe and analyze dynamic processes of land cover. Although numerous deep learning-based CD models have performed excellently, their further performance improvements are constrained by the limited knowledge extracted from the given labelled data. On the other hand, the foundation models that emerged recently contain a huge amount of knowledge by scaling up across data modalities and proxy tasks. In this paper, we propose a Bi-Temporal Adapter Network (BAN), which is a universal foundation model-based CD adaptation framework aiming to extract the knowledge of foundation models for CD. The proposed BAN contains three parts, i.e. frozen foundation model (e.g., CLIP), bi-temporal adapter branch (Bi-TAB), and bridging modules between them. Specifically, BAN extracts general features through a frozen foundation model, which are then selected, aligned, and injected into Bi-TAB via the bridging modules. Bi-TAB is designed as a model-agnostic concept to extract task/domain-specific features, which can be either an existing arbitrary CD model or some hand-crafted stacked blocks. Beyond current customized models, BAN is the first extensive attempt to adapt the foundation model to the CD task. Experimental results show the effectiveness of our BAN in improving the performance of existing CD methods (e.g., up to 4.08% IoU improvement) with only a few additional learnable parameters. More importantly, these successful practices show us the potential of foundation models for remote sensing CD. The code is available at https://github.com/likyoo/BAN and will be supported in our Open-CD.
Kaiyu Li 0001, Xiangyong Cao, Deyu Meng
IEEE Trans. Geosci. Remote. Sens.2
2024 Infrared Small Target Detection via Joint Low Rankness and Local Smoothness Prior
abstract
Infrared small target detection (ISTD) is a challenging task in the computer vision field due to factors such as target scale variations and strong clutter. The existing infrared patch tensor (IPT) models achieve good detection performance but still have several limitations, such as inaccurate background modeling results and poor robustness against noise. To alleviate these issues, in this article, we propose a new IPT model (dubbed as IPT-TCTV) by fully exploiting prior background knowledge. We construct an improved spatial-temporal (STT) model by sliding a 3-D window, which could better preserve the spatial correlation and temporal continuity of multiframe infrared images in the constructed tensor. Specifically, a joint low-rank and local smoothness regularization, i.e., tensor correlated total variation (TCTV), is utilized to characterize the background since the background exhibits not only the low-rank property but also the local smoothness property, without introducing additional trade-off parameters. Furthermore, considering the effect of edge structures, the${l} _{2,1}$norm is adopted as a noise constraint to eliminate strong residuals, which can help to extract real targets from the background with more precision. Finally, we design an efficient alternating direction method of multipliers (ADMMs) approach to solve the proposed model. Experimental results on some benchmark datasets illustrate that our IPT-TCTV model can achieve better detection performance than other state-of-the-art (SOTA) methods in various real scenes. The source code is released athttps://github.com/AuroraPei/IPT-TCTV.
Jiangjun Peng, Hailin Wang 0001, Danfeng Hong, Xiangyong Cao
IEEE Trans. Geosci. Remote. Sens.5
2024 Learnable Representative Coefficient Image Denoiser for Hyperspectral Image
abstract
Fully characterizing the spatial-spectral priors of hyperspectral images (HSI) is crucial for HSI denoising tasks. Recently, HSI denoising models based on representative coefficient images (RCIs) under the spectral low-rank decomposition framework have garnered significant attention due to their clever utilization of spatial-spectral information in HSI at a low cost. However, current methods either employ handcrafted classical denoisers or off-the-shelf deep denoisers to denoise RCIs, failing to fully capture the structural information of RCIs. In this paper, we propose a specific optimization framework for learning an RCI denoiser under the low-rank decomposition framework for the first time. Since low-rank decomposition can characterize the global low-rank property of HSI, our RCI denoiser only needs to learn the spatial prior of RCIs. Consequently, our optimization framework is inclined to learn a more powerful RCI denoiser. However, learning an RCI denoiser is not an easy task, primarily due to the lack of paired clean-noisy RCI data. To address this issue, we employ parametric techniques to represent the to-be-restored HSI as a function of RCI denoiser network parameters. In this way, the parameters of the RCI denoiser can thus be updated using noisy-clean HSI pairs. Furthermore, we adopt residual learning and Gaussian whitening techniques to enhance the RCI denoiser’s denoising ability for HSIs with various noise levels and different rank settings. Extensive experiments demonstrate that our method can achieve significant improvements in both denoising effectiveness and speed compared to state-of-the-art methods. The code of our algorithm is released at https://github.com/andrew-pengjj/RCILD.git.
Jiangjun Peng, Hailin Wang 0001, Xiangyong Cao, Qian Zhao 0002, Jing Yao 0002, Hong-Ying Zhang 0001, Deyu Meng
IEEE Trans. Geosci. Remote. Sens.3
2024 Tensor Ring Decomposition-Based Generalized and Efficient Nonconvex Approach for Hyperspectral Anomaly Detection
abstract
Anomaly detection in hyperspectral images (HSIs) aims to identify sparse, interesting anomalies against the background, which has become a significant topic in remote sensing. Although the existing tensor-based methods have achieved commendable performance to some extent, there is still room for further improvement. In combination with three key techniques, i.e., gradient map-based modeling, circular tensor ring (TR) unfolding, and nonconvex regularization, this article proposes a novel generalized nonconvex method for hyperspectral anomaly detection (HAD) tasks within the TR framework. For the implementation of our proposed approach, abbreviated as TR-GNHAD, we first develop an effective and reliable HAD model in virtue of two newly unified nonconvex regularizers. The first regularizer is devised under a new prior characterization paradigm, which has a strong ability to encode two insightful prior information underlying the HSI’s background simultaneously, i.e., global low rankness and local smoothness. The other regularizer can well capture the structured sparsity of the abnormal component. Then, we derive an efficient optimization algorithm to solve the proposed model based on the alternating direction method of multipliers (ADMMs) framework. Experiments conducted on 12 HSI datasets illustrate that the proposed approach achieves highly competitive performance in both qualitative and quantitative metrics compared with several state-of-the-art HAD methods.
Wenjin Qin, Hailin Wang 0001, Feng Zhang 0023, Jianjun Wang 0003, Xiangyong Cao, Xi-Le Zhao
IEEE Trans. Geosci. Remote. Sens.5
2024 Variational Zero-Shot Multispectral Pansharpening
abstract
Pansharpening aims to generate a high spatial-resolution multispectral image (HRMS) by fusing a low spatial-resolution multispectral image (LRMS) and a panchromatic image (PAN). The most challenging issue for this task is that only the to-be-fused LRMS and PAN are available, and the existing deep learning (DL)-based methods are unsuitable since they rely on many training pairs. Traditional variational optimization (VO) based methods are well-suited for addressing such a problem. They focus on carefully designing explicit fusion rules and regularizations for an optimization problem, which are based on the researcher’s discovery of the image relationships and image structures. Unlike previous VO-based methods, in this work, we explore such complex relationships by a parameterized term rather than a manually designed one. Specifically, we propose a zero-shot pansharpening method by introducing a neural network into the optimization objective. This network estimates a representation component of HRMS, which mainly describes the relationship between HRMS and PAN. In this way, the network achieves a similar goal to the so-called deep image prior (DIP) because it implicitly regulates the relationship between the HRMS and PAN images through its inherent structure. We directly minimize this optimization objective via network parameters and the expected HRMS image through alternating minimization. Extensive experiments on various benchmark datasets demonstrate that our proposed method can achieve better performance compared with other state-of-the-art (SOTA) methods. The codes are available athttps://github.com/xyrui/PSDip.
Xiangyu Rui, Xiangyong Cao, Deyu Meng
IEEE Trans. Geosci. Remote. Sens.2
2024 Low-Rank Prompt-Guided Transformer for Hyperspectral Image Denoising
abstract
Hyperspectral image (HSI) denoising is an essential preprocessing step for downstream applications. Although vision transformer (ViT)-based approaches show impressive denoising performance through self-similarity modeling, these methods still fail to exploit spatial and spectral correlations while ensuring flexibility and efficacy. To address this issue, we propose a hyperspectral denoising transformer using low-rank prompt (HyLoRa), simultaneously taking the spatial self-similarity and spectral low-rank property into account for HSI denoising. Specifically, to fully utilize intrinsic similarity in spatial domain, we perform cross-shaped window-based spatial self-attention for effectively modeling local and global similarity. Moreover, to exploit low-rank inductive bias, we integrate a low-rank prompt module into attention calculation for counting corrected low-dimensional vectors from a large collection of HSIs. This helps to better refine underlying noise-free structure representations. Compared to existing works, powerful capabilities for modeling spatial and spectral correlations can be built to correct low-rank representation in the feature space. Extensive experiments on both simulated and real remote sensing noise demonstrate that our HyLoRa consistently surpasses the state-of-the-art methods.
Xiaodong Tan 0002, Ming-Wen Shao, Yuanjian Qiao 0001, Tiyao Liu, Xiangyong Cao
IEEE Trans. Geosci. Remote. Sens.5
2024 CRS-Diff: Controllable Remote Sensing Image Generation With Diffusion Model
abstract
The emergence of generative models has revolutionized the field of remote sensing (RS) image generation. Despite generating high-quality images, existing methods are limited in relying mainly on text control conditions, and thus do not always generate images accurately and stably. In this article, we propose CRS-Diff, a new RS generative framework specifically tailored for RS image generation, leveraging the inherent advantages of diffusion models while integrating more advanced control mechanisms. Specifically, CRS-Diff can simultaneously support text-condition, metadata-condition, and image-condition control inputs, thus enabling more precise control to refine the generation process. To effectively integrate multiple condition control information, we introduce a new conditional control mechanism to achieve multiscale feature fusion (FF), thus enhancing the guiding effect of control conditions. To the best of our knowledge, CRS-Diff is the first multiple-condition controllable RS generative model. Experimental results in single-condition and multiple-condition cases have demonstrated the superior ability of our CRS-Diff to generate RS images both quantitatively and qualitatively compared with previous methods. Additionally, our CRS-Diff can serve as a data engine that generates high-quality training data for downstream tasks, e.g., road extraction. The code is available athttps://github.com/Sonettoo/CRS-Diff.
Datao Tang, Xiangyong Cao, Xingsong Hou, Zhongyuan Jiang, Junmin Liu, Deyu Meng
IEEE Trans. Geosci. Remote. Sens.2
2024 Pan-Denoising: Guided Hyperspectral Image Denoising via Weighted Represent Coefficient Total Variation
abstract
This article introduces a novel paradigm for hyperspectral image (HSI) denoising, which is termed pan-denoising. In a given scene, panchromatic (PAN) images capture similar structures and textures to HSIs but with less noise. This enables the utilization of PAN images to guide the HSI denoising process. Consequently, pan-denoising, which incorporates an additional prior, has the potential to uncover underlying structures and details beyond the internal information modeling of traditional HSI denoising methods. However, the proper modeling of this additional prior poses a significant challenge. To alleviate this issue, the article proposes a novel regularization term, panchromatic weighted representation coefficient total variation (PWRCTV). It employs the gradient maps of PAN images to automatically assign different weights of total variation (TV) regularization for each pixel, resulting in larger weights for smooth areas and smaller weights for edges. This regularization forms the basis of a pan-denoising model, which is solved using the alternating direction method of multipliers (ADMM). Extensive experiments on synthetic and real-world datasets demonstrate that PWRCTV outperforms several state-of-the-art methods in terms of metrics and visual quality. Furthermore, an HSI classification experiment confirms that PWRCTV, as a preprocessing method, can enhance the performance of downstream classification tasks. The code and data are available athttps://github.com/shuangxu96/PWRCTV.
Qiao Ke, Jiangjun Peng, Xiangyong Cao, Zixiang Zhao
IEEE Trans. Geosci. Remote. Sens.4
2024 Stacked Tucker Decomposition With Multi-Nonlinear Products for Remote Sensing Imagery Inpainting
abstract
In the field of remote sensing (RS) imaging, the occurrence of adverse meteorological conditions or sensor malfunctions can lead to missing data, posing a substantial impediment. Low-rank tensor decomposition has emerged as a promising strategy for resolving this issue, as it enables the integration of diverse data priors within a unified framework. Although various decomposition techniques, such as Tucker decomposition and tensor ring decomposition (TRD), have been developed based on multilinear products, they may not adequately capture the complex structure of RS imagery. Therefore, there is a need for tensor decompositions that incorporate nonlinear operations. To alleviate this challenge, a multi-nonlinear product is defined, which enables the construction of a nonlinear Tucker decomposition (NTD) model. To enhance the model’s capability, a stacked Tucker decomposition (STD) model is formulated, by representing a tensor as the product of a core tensor and a collection of factor matrices along each mode, utilizing the multi-nonlinear product, which potentially regulates the distribution of singular values, thereby achieving a more accurate characterization of textures. The proposed model, integrated with total variation regularization, is subsequently applied to the task of RS imagery inpainting. Extensive experimental results demonstrate the superiority of the proposed model over state-of-the-art (SOTA) methods across various tasks. This validates its effectiveness and adaptability in mitigating the challenges associated with RS imagery inpainting. The code is available athttps://github.com/shuangxu96/STDTV.
Jiangjun Peng, Teng-Yu Ji, Xiangyong Cao, Kai Sun 0007, Rongrong Fei, Deyu Meng
IEEE Trans. Geosci. Remote. Sens.4
2024 Boundary-Aware Spatial and Frequency Dual-Domain Transformer for Remote Sensing Urban Images Segmentation
abstract
Semantic segmentation of remote sensing (RS) images refers to labeling each pixel with a class to identify objects or land cover types. Existing mainstream spatial-domain semantic segmentation methods are mainly categorized into convolutional neural network (CNN)-based and vision transformer (ViT)-based approaches. The former excels at capturing local features, while the latter is adept at extracting global features. Several recent approaches consider combining CNN and ViT to efficiently capture local and global features. However, these approaches still struggle to capture complete features of the RS images, resulting in inaccurate segmentation. To address this issue, we introduce the fast Fourier transform (FFT), which transforms images into the frequency domain for feature extraction, acquiring the image-size receptive field that can complement spatial-domain methods. Based on this, we propose a boundary-aware spatial and frequency dual-domain transformer, termed dual-domain transformer. Specifically, our dual-domain transformer incorporates a dual-domain mixer (DualM), where the spatial-domain branch combines depthwise convolution and the attention mechanism to extract local and global features effectively, while the frequency-domain branch uses FFT to extract image-size features. The two branches complement each other, enabling a more comprehensive feature extraction of RS images. Meanwhile, a boundary-guided training strategy utilizing a boundary-aware module (BAM) is devised to constrain the model extract and predict boundary detail texture, which is an auxiliary task. In addition, the decoder incorporates a scale-feature fusion module (SFM) for adaptive information fusion between the encoder and decoder. Comprehensive experiments on the Zeebrugge and ISPRS datasets, including Vaihingen and Potsdam, showcase that the dual-domain transformer significantly outperforms state-of-the-art (SOTA) methods.
Jie Zhang 0133, Ming-Wen Shao, Yecong Wan, Lingzhuang Meng, Xiangyong Cao, Shuigen Wang
IEEE Trans. Geosci. Remote. Sens.5
2023 Probability-based Global Cross-modal Upsampling for Pansharpening
abstract
Pansharpening is an essential preprocessing step for remote sensing image processing. Although deep learning (DL) approaches performed well on this task, current upsampling methods used in these approaches only utilize the local information of each pixel in the low-resolution multispectral (LRMS) image while neglecting to exploit its global information as well as the cross-modal information of the guiding panchromatic (PAN) image, which limits their performance improvement. To address this issue, this paper develops a novel probability-based global cross-modal upsampling (PGCU) method for pan-sharpening. Precisely, we first formulate the PGCU method from a probabilistic perspective and then design an efficient network module to implement it by fully utilizing the information mentioned above while simultaneously considering the channel specificity. The PGCU module consists of three blocks, i.e., information extraction (IE), distribution and expectation estimation (DEE), and fine adjustment (FA). Extensive experiments verify the superiority of the PGCU method compared with other popular upsampling methods. Additionally, experiments also show that the PGCU module can help improve the performance of existing SOTA deep learning pansharpening methods. The codes are available at https://github.com/Zeyu-Zhu/PGCU.
Xiangyong Cao, Man Zhou 0003, Deyu Meng
CVPR2
2023 PanFlowNet: A Flow-Based Deep Network for Pan-sharpening
abstract
Pan-sharpening aims to generate a high-resolution multispectral (HRMS) image by integrating the spectral information of a low-resolution multispectral (LRMS) image with the texture details of a high-resolution panchromatic (PAN) image. It essentially inherits the ill-posed nature of the super-resolution (SR) task that diverse HRMS images can degrade into an LRMS image. However, existing deep learning-based methods recover only one HRMS image from the LRMS image and PAN image using a deterministic mapping, thus ignoring the diversity of the HRMS image. In this paper, to alleviate this ill-posed issue, we propose a flow-based pan-sharpening network (PanFlowNet) to directly learn the conditional distribution of HRMS image given LRMS image and PAN image instead of learning a deterministic mapping. Specifically, we first transform this unknown conditional distribution into a given Gaussian distribution by an invertible network, and the conditional distribution can thus be explicitly defined. Then, we design an invertible Conditional Affine Coupling Block (CACB) and further build the architecture of PanFlowNet by stacking a series of CACBs. Finally, the PanFlowNet is trained by maximizing the log-likelihood of the conditional distribution given a training set and can then be used to predict diverse HRMS images. The experimental results verify that the proposed PanFlowNet can generate various HRMS images given an LRMS image and a PAN image. Additionally, the experimental results on different kinds of satellite datasets also demonstrate the superiority of our PanFlowNet compared with other state-of-the-art methods both visually and quantitatively. Code is available at Github.
Xiangyong Cao, Wenzhe Xiao, Man Zhou 0003, Aiping Liu, Xun Chen 0001, Deyu Meng
ICCV2
2023 Memory-Augmented Deep Unfolding Network for Guided Image Super-resolution
Man Zhou 0003, Jinshan Pan, Wenqi Ren, Qi Xie 0002, Xiangyong Cao
Int. J. Comput. Vis.6
2023 Hyperspectral Image Denoising Via Texture-Preserved Total Variation Regularizer
abstract
The total variation (TV) regularizer is a widely used technique in image processing tasks to model an image’s local smoothness property. Intrinsically, the TV regularizer imposes sparsity constraints on the gradient maps of the image, which inevitably weakens the image texture structure and thus affects the quality of image restoration. To alleviate this issue, we propose a novel texture-preserved total variation (TPTV) regularizer for hyperspectral image (HSI) by introducing a weighting scheme. Specifically, the weights are assigned to the gradient maps of HSI, which help slack the sparsity constraint for the pixels with large variations, thus preserving the texture structure. Additionally, we elaborate an empirical method to learn the weights adaptively from observed HSI. Then, we propose an HSI denoising method based on the TPTV regularizer. Experimental results on synthetic and real HSI illustrate the superiority of our proposed method over other state-of-the-art methods. In addition, the proposed weighting scheme can be finely embedded into other TV regularizers and protect the image texture. The experiment results also demonstrate that the denoising performance of the original method is significantly improved after embedding the weighting scheme.
Yang Chen 0057, Wenfei Cao, Li Pang, Jiangjun Peng, Xiangyong Cao
IEEE Trans. Geosci. Remote. Sens.5
2022 Proximal PanNet: A Model-Based Deep Network for Pansharpening
abstract
Recently, deep learning techniques have been extensively studied for pansharpening, which aims to generate a high resolution multispectral (HRMS) image by fusing a low resolution multispectral (LRMS) image with a high resolution panchromatic (PAN) image. However, existing deep learning-based pansharpening methods directly learn the mapping from LRMS and PAN to HRMS. These network architectures always lack sufficient interpretability, which limits further performance improvements. To alleviate this issue, we propose a novel deep network for pansharpening by combining the model-based methodology with the deep learning method. Firstly, we build an observation model for pansharpening using the convolutional sparse coding (CSC) technique and design a proximal gradient algorithm to solve this model. Secondly, we unfold the iterative algorithm into a deep network, dubbed as Proximal PanNet, by learning the proximal operators using convolutional neural networks. Finally, all the learnable modules can be automatically learned in an end-to-end manner. Experimental results on some benchmark datasets show that our network performs better than other advanced methods both quantitatively and qualitatively.
Xiangyong Cao, Yang Chen 0057, Wenfei Cao
AAAI1
2022 PanCSC-Net: A Model-Driven Deep Unfolding Method for Pansharpening
abstract
Recently, deep learning (DL) approaches have been widely applied to the pansharpening problem, which is defined as fusing a low-resolution multispectral (LRMS) image with a high-resolution panchromatic (PAN) image to obtain a high-resolution multispectral (HRMS) image. However, most DL-based methods handle this task by designing black-box network architectures to model the mapping relationship from LRMS and PAN to HRMS. These network architectures always lack sufficient interpretability, which limits their further performance improvements. To address this issue, we adopt the model-driven method to design an interpretable deep network structure for pansharpening. First, we present a new pansharpening model using the convolutional sparse coding (CSC), which is quite different from the current pansharpening frameworks. Second, an alternative algorithm is developed to optimize this model. This algorithm is further unfolded to a network, where each network module corresponds to a specific operation of the iterative algorithm. Therefore, the proposed network has clear physical interpretations, and all the learnable modules can be automatically learned in an end-to-end way from the given dataset. Experimental results on some benchmark datasets show that our network performs better than other advanced methods both quantitatively and qualitatively.
Xiangyong Cao, Xueyang Fu, Danfeng Hong, Zongben Xu, Deyu Meng
IEEE Trans. Geosci. Remote. Sens.1
2022 Deep Spatial-Spectral Global Reasoning Network for Hyperspectral Image Denoising
abstract
Although deep neural networks (DNNs) have been widely applied to hyperspectral image (HSI) denoising, most DNN-based HSI denoising methods are designed by stacking convolution layer, which can only model and reason local relations, and thus ignore the global contextual information. To address this issue, we propose a deep spatial-spectral global reasoning network to consider both the local and global information for HSI noise removal. Specifically, two novel modules are proposed to model and reason global relational information. The first one aims to model global spatial relations between pixels in feature maps, and the second one models the global relations across the channels. Compared to traditional convolution operations, the two proposed modules enable the network to extract representations from new dimensions. For the HSI denoising task, the two modules, as well as the densely connected structures, are embedded into the U-Net architecture. Thus, the new-designed global reasoning network can help tackle complex noise by exploiting multiple representations, e.g., hierarchical local feature, global spatial coherence, cross-channel correlation, and multi-scale abstract representation. Experiments on both synthetic and real HSI data demonstrate that our proposed network can obtain comparable or even better denoising results than other state-of-the-art methods.
Xiangyong Cao, Xueyang Fu, Chen Xu 0007, Deyu Meng
IEEE Trans. Geosci. Remote. Sens.1
2022 Hyperspectral Image Denoising With Weighted Nonlocal Low-Rank Model and Adaptive Total Variation Regularization
abstract
Hyperspectral image (HSI) is always corrupted by various types of noise during image capturing, such as Gaussian noise, stripe noise, deadline noise, impulse noise, and more. Such complicated noise significantly degrades imaging quality and thus limits the performance of downstream vision tasks. Current HSI denoising methods tackle this problem by modeling either the spectral-spatial prior of HSI or the noise characteristic of HSI, and few work consider the two aspects simultaneously. In this paper, we propose a new HSI denoising method by simultaneously modeling the HSI prior and the HSI noise characteristic. Specifically, we firstly utilize the non independent and identically distributed (non i.i.d.) mixture of Gaussian (MoG) assumption to characterize the complex noise, which corresponds to optimize a weighted fidelity function. Secondly, we exploit HSI’s non-local similarity and spatial-spectral correlation priors by applying non-local low rank model. Thirdly, we design an adaptive edge preserving total variation regularization term to characterize the non-local smooth property of HSI. Finally, we propose a new denoising model and develop effective ADMM algorithm to solve it. Extensive experiments on simulated data and real data substantiate the superiority of the proposed method beyond state-of-the-arts.
Yang Chen 0057, Wenfei Cao, Li Pang, Xiangyong Cao
IEEE Trans. Geosci. Remote. Sens.4
2022 SRAF-Net: A Scene-Relevant Anchor-Free Object Detection Network in Remote Sensing Images
abstract
Object detection is a fundamental and important task in the analysis ofremote sensing images(RSIs), and existing deep learning-based object detection models in this literature strongly rely on predefined anchor boxes and encounter redesigned difficulties related to anchors. In addition, they often ignore the scene-contextual information that objects are usually closely related to their surrounding scene. To deal with these problems, we propose an anchor-free network, referred to asscene-relevant anchor-free network(SRAF-Net), for object detection in RSIs. The SRAF-Net first captures the scene-contextual features of objects by using a designedscene-enhanced feature pyramid network(SE-FPN) and then performs more accurate detection by implementing ascene auxiliary detection head(SADH), which can predict the existence of the objects with the help of the scene-contextual features extracted from the SE-FPN. To deal with insufficient scene diversity in the training stage, a simple yet effective data augmentation module, termedbalanced mixup data augment(BMDA), is introduced by linearly expanding the training dataset to improve the generalization of SRAF-Net. Comprehensive experiments on three publicly available challenging remote sensing datasets demonstrate the effectiveness of the proposed method. The codes will be made publicly available athttps://github.com/Complicateddd/SRAF-Net.
Junmin Liu, Changsheng Zhou, Xiangyong Cao
IEEE Trans. Geosci. Remote. Sens.4
2022 Fast Noise Removal in Hyperspectral Images via Representative Coefficient Total Variation
abstract
Mining structural priors in data is a widely recognized technique for hyperspectral image (HSI) denoising tasks, whose typical ways include model-based methods and data-based methods. The model-based methods have good generalization ability, while the runtime can hardly meet the fast processing requirements of the practical situations due to the large size of an HSI${\mathbf {X}}\in \mathbb {R}^{\textrm {MN}\times B}$. For the data-based methods, they perform relatively fast on new test data once they have been trained. However, their generalization ability is always insufficient. In this article, we propose a fast model-based approach via a novel regularizer named the representative coefficient total variation (RCTV) to simultaneously characterize the low-rank and local smooth properties. The RCTV regularizer is proposed based on the observation that the representative coefficient matrix${\mathbf {U}}\in \mathbb {R}^{\textrm {MN}\times R} (R\ll B)$obtained by orthogonally transforming the original HSI${\mathbf {X}}$can inherit the strong local-smooth prior of${\mathbf {X}}$. Since$R/B$is very small, the model based on the RCTV regularizer has lower time complexity. In addition, we find that the representative coefficient matrix${\mathbf {U}}$is robust to noise, and thus, the RCTV regularizer can somewhat promote the robustness of the HSI denoising model. Extensive experiments on mixed noise removal demonstrate that the proposed method realizes a perfect compromise between denoising performance and denoising speed compared with other state-of-the-art methods. Remarkably, the denoising speed of our proposed method outperforms all competing model-based techniques and is comparable with the deep learning-based approaches. The code of our algorithm is released athttps://github.com/andrew-pengjj/rctv.git.
Jiangjun Peng, Hailin Wang 0001, Xiangyong Cao, Xinling Liu, Xiangyu Rui, Deyu Meng
IEEE Trans. Geosci. Remote. Sens.3
2022 Hyperspectral Image Denoising by Asymmetric Noise Modeling
abstract
In general, hyperspectral images (HSIs) are degraded by a mixture of complicated noise (i.e., mixture of Gaussian and sparse noise), and how to precisely model HSI noise plays a vital role in the task of HSI denoising. The most popular choices for encoding the noise distribution are Gaussian, Laplacian, and the mixture of Gaussians, but they are always incompatible with real-world HSI noise. By investigating histograms of the error map, we first explore that asymmetry is a typical and general feature of HSI noise. Inspired by this discovery, we find that a bandwise asymmetric Laplacian (AL) distribution can be finely used to model this type of noise. Equipped with the low-rank matrix factorization (LRMF) framework, we formulate a novel model by the maximum likelihood estimation (MLE) principle, which can be efficiently solved using the iterative optimization algorithm. Extensive experimental results on synthetic and real datasets demonstrate that the proposed model outperforms other counterparts. It is also found that scale and asymmetry parameters in the AL distribution can well interpret the pattern of real-world HSI noise.
Xiangyong Cao, Jiangjun Peng, Qiao Ke, Cong Ma 0005, Deyu Meng
IEEE Trans. Geosci. Remote. Sens.2
2022 Sequential ISAR Target Classification Based on Hybrid Transformer
abstract
To make full use of the sequential information obtained by continuous inverse synthetic aperture radar (ISAR) imaging, this article proposes a sequential ISAR target classification network based on hybrid transformer (HT). First, a temporal–spatial encoder based on the attention mechanism is designed to extract long-term and global features from sequential images. Meanwhile, a local feature encoder based on the 3-D convolution neural network is designed to extract short-term and local features. Then, the above two features are fused and the classification labels are obtained by a channel encoder–decoder. In 4-satellite target classification experiments, the proposed HT shows high accuracy and robustness to the unknown image scaling, rotation, and combined deformations.
Ruihang Xue, Xueru Bai, Xiangyong Cao, Feng Zhou 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 MD³Net: Integrating Model-Driven and Data-Driven Approaches for Pansharpening
abstract
Pansharpening is a special image fusion task of reconstructing a high-resolution multispectral (HRMS) image by integrating a panchromatic (PAN) image of high spatial resolution and a low-resolution multispectral (LRMS) image. To handle such an ill-posed multi-modal fusion task, in this paper, we propose a novel pansharpening method, referred to as model-driven and data-driven network (MD3Net), which combines model-driven and data-driven approaches. The architecture design of MD3Net is inspired from the traditional model constructed based on domain knowledge and thus making its network topology explainable and its input/output predictable. In order to further explore the powerful learning ability of deep learning based approaches, we introduce the deep prior into the MD3Net as its implicit regularization, thus improving its data adaptability and representation capability. Comprehensive experiments conducted on both reduced and full resolution of several acknowledged datasets have qualitatively and quantitatively verified the superiority of our network compared to a benchmark consisting of several state-of-the-art approaches. The code can be downloaded from https://github.com/YinsongYan/M3DNet..
Yinsong Yan, Junmin Liu, Xiangyong Cao
IEEE Trans. Geosci. Remote. Sens.5
2022 Semi-Active Convolutional Neural Networks for Hyperspectral Image Classification
abstract
Owing to the powerful data representation ability of deep learning (DL) techniques, tremendous progress has been recently made in hyperspectral image (HSI) classification. Convolutional neural network (CNN), as a main part of the DL family, has been proven to be considerably effective to extract spatial-spectral features for HSIs. Nevertheless, its classification performance, to a great extent, depends on the quality and quantity of samples in the network training process. To select those samples, either labeled or unlabeled, that can be used to enhance the generalization ability of CNNs and further improve the classification accuracy, we propose an iterative semi-supervised CNNs framework by means of active learning and superpixel segmentation techniques, dubbed as semi-active CNNs (SA-CNNs), for HSI classification. More specifically, we start to pre-train a CNNs-based model on a small-scale unbiased labeled set and infer unlabeled data using the trained model, i.e., generating pseudo-labels. Then, the reliable samples, which consist of two parts: high label-homogeneity and most informativeness, are actively selected from superpixel segments. These selected labeled and unlabeled samples with their labels and pseudo-labels are re-fed into the next-round network training. Moreover, three different schedules, i.e.,log-,exp-, andlinear-schedules, are progressively adopted to fully explore their potentials in sample selection, until a labeling budget is finally reached. Extensive experiments are conducted on three benchmark HSI datasets, demonstrating substantial performance improvements of the proposed SA-CNNs over other similar competitors.
Jing Yao 0002, Xiangyong Cao, Danfeng Hong, Xin Wu 0001, Deyu Meng, Jocelyn Chanussot, Zongben Xu
IEEE Trans. Geosci. Remote. Sens.2
2022 A Model-Driven Deep Unfolding Method for JPEG Artifacts Removal
abstract
Deep learning-based methods have achieved notable progress in removing blocking artifacts caused by lossy JPEG compression on images. However, most deep learning-based methods handle this task by designing black-box network architectures to directly learn the relationships between the compressed images and their clean versions. These network architectures are always lack of sufficient interpretability, which limits their further improvements in deblocking performance. To address this issue, in this article, we propose a model-driven deep unfolding method for JPEG artifacts removal, with interpretable network structures. First, we build a maximum posterior (MAP) model for deblocking using convolutional dictionary learning and design an iterative optimization algorithm using proximal operators. Second, we unfold this iterative algorithm into a learnable deep network structure, where each module corresponds to a specific operation of the iterative algorithm. In this way, our network inherits the benefits of both the powerful model ability of data-driven deep learning method and the interpretability of traditional model-driven method. By training the proposed network in an end-to-end manner, all learnable modules can be automatically explored to well characterize the representations of both JPEG artifacts and image content. Experiments on synthetic and real-world datasets show that our method is able to generate competitive or even better deblocking results, compared with state-of-the-art methods both quantitatively and qualitatively.
Xueyang Fu, Menglu Wang 0003, Xiangyong Cao, Xinghao Ding, Zhengjun Zha
IEEE Trans. Neural Networks Learn. Syst.3
2021 Learning an Explicit Weighting Scheme for Adapting Complex HSI Noise
abstract
An efficient approach for handling hyperspectral image (HSI) denoising issue is to impose weights on different HSI pixels to suppress negative influence brought by noisy elements. Such weighting scheme, however, largely depends on the prior understanding or subjective distribution assumption on HSI noises, making them easily biased to complicated real noises, and hardly generalizable to diverse practical scenarios. Against this issue, this paper proposes a new scheme aiming to capture general weighting principle in a data-driven manner. Specifically, such weighting principle is delivered by an explicit function, called hyper-weight-net (HWnet), mapping from an input noisy image to its properly imposed weights. A Bayesian framework as well as a variational inference algorithm for inferring HWnet parameters is elaborately designed, expecting to extract the latent weighting rule for general diverse and complicated noisy HSIs. Comprehensive experiments substantiate that the learned HWnet can be not only finely generalized to different noise types from those used in training, but also effectively transferred to other weighted models. Besides, as a sounder guidance, HWnet can help to more faithfully and robustly achieve deep hyperspectral prior(DHP). The extracted weights by HWnet are verified to be able to effectively capture complex noise knowledge underlying input HSI, revealing its working insight in experiments.
Xiangyu Rui, Xiangyong Cao, Qi Xie 0002, Zongsheng Yue, Qian Zhao 0002, Deyu Meng
CVPR2
2021 An Enhanced 3-D Discrete Wavelet Transform for Hyperspectral Image Classification
abstract
In the classification of hyperspectral image (HSI), there exists a common issue that the collected HSI data set is always contaminated by various noise (e.g., Gaussian, stripe, and deadline), degrading the classification results. To tackle this issue, we modify the 3-dimensional discrete wavelet transform (3DDWT) method by considering the noise effect on feature quality and propose an enhanced 3DDWT (E-3DDWT) approach to extract the feature and meanwhile alleviate the noise. Specifically, the proposed E-3DDWT method first applies classical 3DDWT method to the HSI data cube and thus can generate eight subcubes in each level. Then, the stripe noise is concentrated into several subcubes due to its spatial vertical property. Finally, we abandon these subcubes and obtain the feature cube by stacking the remaining ones. After acquiring the feature, we then adopt the convolutional neural network (CNN) model with an active learning strategy for classification since CNN has been verified to be a state-of-the-art feature extraction method for HSI classification, and active learning strategy can alleviate the insufficient labeled sample issue to some extent. In addition, we apply the Markov random field to enhance the final categorized results. Experiments on two synthetically striped data sets show that our proposed approach achieves better categorized results than other advanced methods.
Xiangyong Cao, Jing Yao 0002, Xueyang Fu, Haixia Bi, Danfeng Hong
IEEE Geosci. Remote. Sens. Lett.1
2021 Online Rain/Snow Removal From Surveillance Videos
abstract
Video rain/snow removal from surveillance videos is an important task in the computer vision community since rain/snow existed in videos can severely degenerate the performance of many surveillance system. Various methods have been investigated extensively, but most only consider consistent rain/snow under stable background scenes. Rain/snow captured from practical surveillance camera, however, is always highly dynamic in time, and those videos also include occasionally transformed background scenes and background motions caused by waving leaves or water surfaces. To this issue, this paper proposes a novel rain/snow removal approach, which fully considers dynamic statistics of both rain/snow and background scenes taken from a video sequence. Specifically, the rain/snow is encoded as an online multi-scale convolutional sparse coding (OMS-CSC) model, which not only finely delivers the sparse scattering and multi-scale shapes of real rain/snow, but also well distinguish the components of background motion from rain/snow layer. The real-time ameliorated parameters in the model well encodes their temporally dynamic configurations. Furthermore, a transformation operator imposed on the background scenes is further embedded into the proposed model, which finely conveys the background transformations, such as rotations, scalings and distortions, inevitably existed in a real video sequence. The approach so constructed can naturally better adapt to the dynamic rain/snow as well as background changes, and also suitable to deal with the streaming video attributed its online learning mode. The proposed model is formulated in a concise maximum a posterior (MAP) framework and is readily solved by the alternating direction method of multipliers (ADMM). Compared with the state-of-the-art online and offline video rain/snow removal methods, the proposed method achieves best performance on synthetic and real videos datasets both visually and quantitatively. Specifically, our method can be implemented in relatively high efficiency, showing its potential to real-time video rain/snow removal. The code page is at: https://github.com/MinghanLi/OTMSCSC_matlab_2020.
Minghan Li 0001, Xiangyong Cao, Qian Zhao 0002, Lei Zhang 0006, Deyu Meng
IEEE Trans. Image Process.2
2020 New Interpretations of Normalization Methods in Deep Learning
abstract
In recent years, a variety of normalization methods have been proposed to help training neural networks, such as batch normalization (BN), layer normalization (LN), weight normalization (WN), group normalization (GN), etc. However, some necessary tools to analyze all these normalization methods are lacking. In this paper, we first propose a lemma to define some necessary tools. Then, we use these tools to make a deep analysis on popular normalization methods and obtain the following conclusions: 1) Most of the normalization methods can be interpreted in a unified framework, namely normalizing pre-activations or weights onto a sphere; 2) Since most of the existing normalization methods are scaling invariant, we can conduct optimization on a sphere with scaling symmetry removed, which can help to stabilize the training of network; 3) We prove that training with these normalization methods can make the norm of weights increase, which could cause adversarial vulnerability as it amplifies the attack. Finally, a series of experiments are conducted to verify these claims.
Xiangyong Cao, Hanwen Liang, Weiran Huang 0001, Zewei Chen, Zhenguo Li
AAAI2
2020 Discovering influential factors in variational autoencoders
Shiqi Liu 0001, Qian Zhao 0002, Xiangyong Cao, Huibin Li 0001, Deyu Meng, Hongying Meng, Sheng Liu 0033
Pattern Recognit.4
2020 Underwater image enhancement with global-local networks and compressed-histogram equalization
Xueyang Fu, Xiangyong Cao
Signal Process. Image Commun.2
2020 Hyperspectral Image Classification With Convolutional Neural Network and Active Learning
abstract
Deep neural network has been extensively applied to hyperspectral image (HSI) classification recently. However, its success is greatly attributed to numerous labeled samples, whose acquisition costs a large amount of time and money. In order to improve the classification performance while reducing the labeling cost, this article presents an active deep learning approach for HSI classification, which integrates both active learning and deep learning into a unified framework. First, we train a convolutional neural network (CNN) with a limited number of labeled pixels. Next, we actively select the most informative pixels from the candidate pool for labeling. Then, the CNN is fine-tuned with the new training set constructed by incorporating the newly labeled pixels. This step together with the previous step is iteratively conducted. Finally, Markov random field (MRF) is utilized to enforce class label smoothness to further boost the classification performance. Compared with the other state-of-the-art traditional and deep learning-based HSI classification methods, our proposed approach achieves better performance on three benchmark HSI data sets with significantly fewer labeled samples.
Xiangyong Cao, Jing Yao 0002, Zongben Xu, Deyu Meng
IEEE Trans. Geosci. Remote. Sens.1
2020 Polarimetric SAR Image Semantic Segmentation With 3D Discrete Wavelet Transform and Markov Random Field
abstract
Polarimetric synthetic aperture radar (PolSAR) image segmentation is currently of great importance in image processing for remote sensing applications. However, it is a challenging task due to two main reasons. Firstly, the label information is difficult to acquire due to high annotation costs. Secondly, the speckle effect embedded in the PolSAR imaging process remarkably degrades the segmentation performance. To address these two issues, we present a contextual PolSAR image semantic segmentation method in this paper. With a newly defined channel-wise consistent feature set as input, the three-dimensional discrete wavelet transform (3D-DWT) technique is employed to extract discriminative multi-scale features that are robust to speckle noise. Then Markov random field (MRF) is further applied to enforce label smoothness spatially during segmentation. By simultaneously utilizing 3D-DWT features and MRF priors for the first time, contextual information is fully integrated during the segmentation to ensure accurate and smooth segmentation. To demonstrate the effectiveness of the proposed method, we conduct extensive experiments on three real benchmark PolSAR image data sets. Experimental results indicate that the proposed method achieves promising segmentation accuracy and preferable spatial consistency using a minimal number of labeled pixels.
Haixia Bi, Lin Xu 0001, Xiangyong Cao, Yong Xue, Zongben Xu
IEEE Trans. Image Process.3
2018 Robust subspace clustering via penalized mixture of Gaussians
Jing Yao 0002, Xiangyong Cao, Qian Zhao 0002, Deyu Meng, Zongben Xu
Neurocomputing2
2018 Denoising Hyperspectral Image With Non-i.i.d. Noise Structure
abstract
Hyperspectral image (HSI) denoising has been attracting much research attention in remote sensing area due to its importance in improving the HSI qualities. The existing HSI denoising methods mainly focus on specific spectral and spatial prior knowledge in HSIs, and share a common underlying assumption that the embedded noise in HSI is independent and identically distributed (i.i.d.). In real scenarios, however, the noise existed in a natural HSI is always with much more complicated non-i.i.d. statistical structures and the under-estimation to this noise complexity often tends to evidently degenerate the robustness of current methods. To alleviate this issue, this paper attempts the first effort to model the HSI noise using a non-i.i.d. mixture of Gaussians (NMoGs) noise assumption, which finely accords with the noise characteristics possessed by a natural HSI and thus is capable of adapting various practical noise shapes. Then we integrate such noise modeling strategy into the low-rank matrix factorization (LRMF) model and propose an NMoG-LRMF model in the Bayesian framework. A variational Bayes algorithm is then designed to infer the posterior of the proposed model. As substantiated by our experiments implemented on synthetic and real noisy HSIs, the proposed method performs more robust beyond the state-of-the-arts.
Yang Chen 0057, Xiangyong Cao, Qian Zhao 0002, Deyu Meng, Zongben Xu
IEEE Trans. Cybern.2
2018 Hyperspectral Image Classification With Markov Random Fields and a Convolutional Neural Network
abstract
This paper presents a new supervised classification algorithm for remotely sensed hyperspectral image (HSI) which integrates spectral and spatial information in a unified Bayesian framework. First, we formulate the HSI classification problem from a Bayesian perspective. Then, we adopt a convolutional neural network (CNN) to learn the posterior class distributions using a patch-wise training strategy to better use the spatial information. Next, spatial information is further considered by placing a spatial smoothness prior on the labels. Finally, we iteratively update the CNN parameters using stochastic gradient decent and update the class labels of all pixel vectors using -expansion min-cut-based algorithm. Compared with the other state-of-the-art methods, the classification method achieves better performance on one synthetic data set and two benchmark HSI data sets in a number of experimental settings.
Xiangyong Cao, Feng Zhou 0001, Lin Xu 0001, Deyu Meng, Zongben Xu, John W. Paisley
IEEE Trans. Image Process.1
2017 Polsar image classification based on three-dimensional wavelet texture features and Markov random field
abstract
The speckle effect embedded in polarimetric synthetic aperture radar (PolSAR) data damages the performance of PolSAR image classification greatly. To alleviate this issue, a new supervised classification method, which introduces spatial consistency in both feature extraction and classification steps is proposed. Specifically, three-dimensional discrete wavelet transform (3D-DWT) is used to extract spectral-spatial texture features, which are proved to be more discriminative than original ones. Afterward, label smoothness prior is incorporated in the classification, which is implemented using a Markov random field (MRF). To demonstrate the validity of the proposed method, real PolSAR image is used in experiments. Compared with the other state-of-the-art methods, this method achieves higher classification accuracy and better visual spatial connectivity.
Haixia Bi, Lin Xu 0001, Xiangyong Cao, Zongben Xu
IGARSS3
2017 Integration of 3-dimensional discrete wavelet transform and Markov random field for hyperspectral image classification
Xiangyong Cao, Lin Xu 0001, Deyu Meng, Qian Zhao 0002, Zongben Xu
Neurocomputing1
2016 Robust Low-Rank Matrix Factorization Under General Mixture Noise Distributions
abstract
Many computer vision problems can be posed as learning a low-dimensional subspace from high-dimensional data. The low rank matrix factorization (LRMF) represents a commonly utilized subspace learning strategy. Most of the current LRMF techniques are constructed on the optimization problems using L1-norm and L2-norm losses, which mainly deal with the Laplace and Gaussian noises, respectively. To make LRMF capable of adapting more complex noise, this paper proposes a new LRMF model by assuming noise as mixture of exponential power (MoEP) distributions and then proposes a penalized MoEP (PMoEP) model by combining the penalized likelihood method with MoEP distributions. Such setting facilitates the learned LRMF model capable of automatically fitting the real noise through MoEP distributions. Each component in this mixture distribution is adapted from a series of preliminary superor sub-Gaussian candidates. Moreover, by facilitating the local continuity of noise components, we embed Markov random field into the PMoEP model and then propose the PMoEP-MRF model. A generalized expectation maximization (GEM) algorithm and a variational GEM algorithm are designed to infer all parameters involved in the proposed PMoEP and the PMoEPMRF model, respectively. The superiority of our methods is demonstrated by extensive experiments on synthetic data, face modeling, hyperspectral image denoising, and background subtraction.
Xiangyong Cao, Qian Zhao 0002, Deyu Meng, Yang Chen 0057, Zongben Xu
IEEE Trans. Image Process.1
2015 Low-Rank Matrix Factorization under General Mixture Noise Distributions
abstract
Many computer vision problems can be posed as learning a low-dimensional subspace from high dimensional data. The low rank matrix factorization (LRMF) represents a commonly utilized subspace learning strategy. Most of the current LRMF techniques are constructed on the optimization problem using L_1 norm and L_2 norm, which mainly deal with Laplacian and Gaussian noise, respectively. To make LRMF capable of adapting more complex noise, this paper proposes a new LRMF model by assuming noise as Mixture of Exponential Power (MoEP) distributions and proposes a penalized MoEP model by combining the penalized likelihood method with MoEP distributions. Such setting facilitates the learned LRMF model capable of automatically fitting the real noise through MoEP distributions. Each component in this mixture is adapted from a series of preliminary super-or sub-Gaussian candidates. An Expectation Maximization (EM) algorithm is also designed to infer the parameters involved in the proposed PMoEP model. The advantage of our method is demonstrated by extensive experiments on synthetic data, face modeling and hyperspectral image restoration.
Xiangyong Cao, Yang Chen 0057, Qian Zhao 0002, Deyu Meng, Yao Wang 0003, Zongben Xu
ICCV1