Liang-Jian Deng

dblp:136/7368 · also Liangjian Deng · DBLP profile ↗
← Back
116ranked-venue papers
8as first author
98since 2021 · last 2026
0000-0003-3178-9772ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 54 · 5 first-author · 43 since 2021Artificial intelligence and machine learning · 47 · 45 since 2021Applied, interdisciplinary, general and emerging computing · 29 · 1 first-author · 26 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 NODiff: Neural Operator Diffusion for Multispectral Image Fusion
abstract
Pansharpening is a powerful technique for generating high-resolution multispectral (HRMS) images by fusing currently available image pairs of low-resolution multispectral (LRMS) and texture-rich panchromatic (PAN) data, effectively addressing the physical constraints of satellite sensors. While recent generative diffusion models have demonstrated impressive performance gains in this domain, their prohibitive computational demands and training costs hinder practicality in resource-constrained remote sensing satellite systems. In this work, we propose NODiff, a novel diffusion framework that replaces the conventional attention-based denoising backbone with a neural operator, seamlessly integrating operator learning and generative modeling into an efficient yet effective solution for pansharpening. In practice, we implement our approach through a two-stage learning paradigm: First, we pretrain the proposed Neural Operator-based diffusion model to learn the high-resolution texture priors essential for pansharpening. Afterward, we freeze the pretrained parameters, and design a lightweight conditional detail guidance adapter to enable efficient fine-tuning for generating desired HRMS images. Meanwhile, a time-aware low-rank adaptation is introduced to dynamically refine high-frequency details potentially affected by spectral mode truncation. Extensive experiments on multiple benchmark datasets demonstrate that NODiff achieves competitive pansharpening performance while significantly reducing training and inference costs. Beyond pansharpening, our method provides new insights into building resource-efficient generative models.
Junming Hou, Ran Ran 0001, Sixing Chen, Xiaofeng Cong, Junling Li, Liang-Jian Deng
AAAI7
2026 SWIFT:A General Sensitive Weight Identification Framework for Fast Sensor-Transfer Pansharpening
abstract
Although deep learning-based methods have achieved promising performance in Pansharpening, they generally suffer from severe performance degradation when applied to data from unseen sensors. Existing cross-domain strategies, including retraining, fine-tuning, and zero-shot methods, fail to simultaneously preserve model architecture and maintain low adaptation costs. Therefore, we are the first to define and address a novel task in the pansharpening field: enhancing a model's cross-sensor generalization at an extremely low cost while keeping the model architecture invariant. To tackle this task, we propose SWIFT (Sensitive Weight Identification for Fast Transfer), a plug-and-play framework. SWIFT first employs an unsupervised manifold-based sampling strategy to efficiently select a high-fidelity subset the most informative target-domain samples. It then leverages this subset to probe a source-domain pre-trained model, identifying and updating only the weight subset most sensitive to the domain shift by analyzing the gradient behavior of its parameters. Extensive experiments demonstrate that SWIFT can be applied to various deep learning models, boosting adaptation efficiency by up to 30-fold. On a single NVIDIA RTX 4090 GPU, this reduces adaptation time from hours to as little as one minute. The adapted models not only substantially outperform direct-transfer baselines but also achieve performance competitive with, or even superior to full retraining while using only 3% of the target domain dataset and adapting nearly 10% to 30% of the model’s parameters. This establishs a new state-of-the-art on the WorldView-2 and QuickBird datasets.
Tianyu Xin, Yubo Zeng, Liang-Jian Deng
AAAI6
2026 Training and Inference Within 1 Second - Tackle Cross-Sensor Degradation of Real-World Pansharpening with Efficient Residual Feature Tailoring
abstract
Deep learning methods for pansharpening have advanced rapidly, yet models pretrained on data from a specific sensor often generalize poorly to data from other sensors. Existing methods to tackle such cross-sensor degradation include retraining model or zero-shot methods, but they are highly time-consuming or even need extra training data. To address these challenges, our method first performs modular decomposition on deep learning-based pansharpening models, revealing a general yet critical interface where high-dimensional fused features begin mapping to the channel space of the final image. % may need revisement A Feature Tailor is then integrated at this interface to address cross-sensor degradation at the feature level, and is trained efficiently with physics-aware unsupervised losses. Moreover, our method operates in a patch-wise manner, training on partial patches and performing parallel inference on all patches to boost efficiency. Our method offer two key advantages: (1) Improved Generalization Ability: it significantly enhance performance in cross-sensor cases. (2) Low Generalization Cost: it achieves sub-second training and inference, requiring only partial test inputs and no external data, whereas prior methods often take minutes or even hours. Experiments on the real-world data from multiple datasets demonstrate that our method achieves state-of-the-art quality and efficiency in tackling cross-sensor degradation. For example, training and inference of 512 times 512 times 8 image within 0.2 seconds and 4000 times 4000 times 8 image within 3 seconds at the fastest setting on a commonly used RTX 3090 GPU, which is over 100 times faster than zero-shot methods.
Tianyu Xin, Jin-Liang Xiao, Liang-Jian Deng
AAAI5
2026 Information-Epidemic Dynamics in Cyber-Physical Systems: A Hypergraph Framework With Interpersonal Relationships
abstract
Understanding how information propagation affects epidemic dynamics has become an emerging topic of interest. However, the influence of interpersonal relationship heterogeneity on information acquisition and disease transmission has been largely overlooked. In this work, we introduce a hypergraph structure for Cyber-Physical Systems (CPSs) with two distinct layers. The upper layer, referred to as the cyber layer, consists of a mixed hypergraph, capturing both pairwise propagation and higher-order diffusion of epidemic-related information. The lower layer, referred to as the physical layer, employs a Susceptible-Infected-Susceptible (SIS) process to capture epidemic spreading. This work introduces an adaptive perception-protection mechanism based on Jaccard similarity, which accounts for interpersonal heterogeneity. In this mechanism, individuals receive information based on their relationships with neighbors and take protective measures accordingly. We analyze the impact of interpersonal relationships and the adoption of neighborhood-based self-protection strategies on epidemic dynamics. Furthermore, we conduct a theoretical analysis based on the Microscopic Markov Chain Approach (MMCA), analytically derive the outbreak threshold, and confirm the results with extensive Monte Carlo (MC) simulations. The results show that stronger interpersonal relationships can promote information propagation, significantly increase the threshold for epidemic outbreaks, and effectively suppress the scale of the epidemic. The study provides theoretical support for designing epidemic control strategies considering interpersonal heterogeneity and improves the understanding of epidemic spreading on hypergraphs.
Shanchao Peng, Minyu Feng, Liang-Jian Deng, Matjaz Perc, Jürgen Kurths
IEEE Internet Things J.3
2026 Reweighted low-rank quaternion matrix factorization with deep denoising prior for color image inpainting
Liangtian He, Shaobing Gao, Jifei Miao, Liang-Jian Deng, Jun Liu 0012
Inf. Sci.5
2026 A General Image Fusion Approach Exploiting Gradient Transfer Learning and Fusion Rule Unfolding
abstract
The goal of a deep learning-based general image fusion method is to solve multiple image fusion tasks with a single model, thereby facilitating the deployment of models in practical applications. However, existing methods fail to provide an efficient and comprehensive solution from both model training and network design perspectives. Regarding model training, current approaches cannot effectively leverage complementary information across different tasks. In terms of network design, they rely on experience-based network designs. To address these issues, we propose a comprehensive framework for general image fusion using the newly proposed gradient transfer learning and fusion rule unfolding. To leverage complementary information across different tasks during training, we propose a sequential gradient-transfer framework based on the idea that different image fusion tasks often exhibit complementary structural details and that image gradients effectively capture these details. To move beyond heuristic-based network design, we evolved a fundamental image fusion rule and integrated it into a deep equilibrium model, resulting in a more efficient and versatile image fusion network capable of uniformly handling various fusion tasks. Considering three different image fusion tasks, i.e., multi-focus image fusion, multi-exposure image fusion, and infrared and visible image fusion, our method not only produces images with richer structural information but also achieves highly competitive objective metrics. Furthermore, the results of generalization experiments on previously unseen image fusion tasks, i.e., medical image fusion, demonstrate that our method significantly outperforms competing approaches.
Liang-Jian Deng, Gemine Vivone
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 Low-rank reduced biquaternion matrix completion with application to color image inpainting
Liangtian He, Jifei Miao, Liang-Jian Deng, Jun Liu 0012
Pattern Recognit.4
2026 Zero-shot unsupervised learning with unfolded equilibrium network for remote sensing pansharpening
Jieyi Zhu, Zihan Cao, Liang-Jian Deng
Pattern Recognit.4
2026 Evolutionary Dynamics of Variable Games in Structured Populations
abstract
The game interactions among individuals in nature are often uncertain and dynamically evolving, significantly influencing the persistence of cooperation. However, it remains a formidable challenge to effectively characterize these dynamic properties in structured populations, derive theoretical conditions for cooperation, and identify the optimal game distribution for promoting cooperation. To address these issues, we propose the variable game framework in a structured population, where the game interactions between different individuals change over time. By means of the Markov chain and the pair approximation method, we derive theoretical conditions under which cooperation is favored by natural selection and when it is favored over defection under weak selection. Furthermore, we, respectively, formulate and solve two optimization problems to determine the optimal game distribution that most effectively fosters the evolution of cooperation by maximizing the gradient of cooperation selection and minimizing the fitness difference between defectors and cooperators. The theoretical predictions regarding both the conditions for cooperation and optimal game distribution are further validated by numerical calculations and extensive Monte Carlo simulations. Our findings offer novel insights into the mechanisms driving cooperative behavior in complex systems and provide theoretical guidance for designing optimal game environments that facilitate the evolution of cooperation.
Bin Pi, Minyu Feng, Liang-Jian Deng, Xiaojie Chen 0003, Attila Szolnoki
IEEE Trans. Cybern.3
2026 Semantic-Decoupled and Knowledge-Shared Probabilistic Mapping Network for Multi-Grained Cross-Modal Retrieval
abstract
Cross-modal retrieval is essential for exploring semantic correlations between multimodal data. However, existing approaches face challenges in resolving semantic ambiguity and transferring knowledge with sparse sample generalization. To address these challenges, we propose a new Semantic-Decoupled and Knowledge-Shared Probabilistic Mapping Network (SKPMN). Specifically, the Semantic Decoupling and Distinction (SDD) module decomposes complex word-region relationships into relevance-driven representations. The Deep Probability Mapping (DPM) module introduces a paradigm shift by mapping multimodal features into probabilistic distributions, capturing the semantic similarities and the potential uncertainties that define sparse or ambiguous relationships. By combining the Attention Probabilistic Mapping (APM) module, the model can effectively transfer knowledge across similar samples while emphasizing critical distinctions, significantly enhancing generalization to sparse and ambiguous samples. Finally, the multi-grained alignment strategy establishes a novel integration of fine-grained patch-to-word alignment and coarse-grained global alignment. Experimental results show that SKPMN achieves superior retrieval accuracy across major benchmark datasets. Furthermore, we implement a channel resource allocation technique that allocates more transmission resources to semantically significant information. In resource-constrained environments, our approach leverages Joint Source-Channel Coding (JSCC) to enhance the efficiency of visual feature transmission.
Wenrui Li 0001, Yeyu Chai, Liang-Jian Deng, Ruiqin Xiong, Xiaopeng Fan 0001, Yonghong Tian 0001
IEEE Trans. Image Process.3
2026 Adaptive 3D Convolution for Remote Sensing Image Fusion
abstract
Remote sensing image fusion aims to create a high-resolution multi/hyper-spectral image from a high-resolution image with limited spectral information and a low-resolution image with abundant spectral data. Recently, deep learning (DL) techniques have shown significant effectiveness in this area. Most DL-based methods approach image fusion as a 2D problem by encoding spectral information into feature map channels. However, our research suggests that this strategy introduces notable spectral distortions. In contrast, some methods consider spectral data as an additional dimension, utilizing standard 3D convolutions to preserve spectral information. Nevertheless, in a standard 3D convolutional layer, the same set of kernels is applied across all input regions, which we have found to be sub-optimal for image fusion. Furthermore, standard 3D convolutions necessitate substantial computational resources. To address these challenges, we propose a novel convolutional paradigm called Adaptive 3D Convolution (Ada3D) for remote sensing image fusion. Ada3D applies a unique set of 3D kernels to each input voxel, enabling the capture of fine-grained details. These adaptive kernels are generated through a two-step process: 1) spatial and spectral kernels are derived from their respective image sources and 2) these two types of kernels are then combined to form content-aware 3D kernels that effectively integrate spatial and spectral information. Additionally, adaptive biases are introduced to enhance the convolutional outcome at the voxel level. Furthermore, we incorporate the group convolution technique to reduce computational complexity. As a result, Ada3D offers full adaptivity in an efficient manner. Evaluation results across five datasets demonstrate that our method achieves state-of-the-art (SOTA) performance, underscoring the superiority of Ada3D. The code is available at https://github.com/PSRben/Ada3D.
Siran Peng, Xiangyu Zhu 0001, Shangqi Deng, Liang-Jian Deng, Zhen Lei 0001
IEEE Trans. Image Process.4
2026 Tensor Wheel Decomposition: Theory and Application to Tensor Completion
abstract
Recently, tensor network (TN) decompositions have gained prominence in computer vision and contributed promising results to tensor recovery for their capability of compactly and efficiently representing high-order tensors. However, current TN topologies are rather being developed towards more intricate structures to pursue incremental improvements, resulting in a drastically increased number of TN ranks, which requires laborious hyper-parameter selection, especially for higher-order cases. In this paper, we propose a novel TN decomposition, dubbed tensor wheel (TW) decomposition, in which a high-order tensor is represented by a set of latent factors mapped into a specific wheel topology. Such a decomposition is constructed starting from analyzing the graph structure, aiming to more accurately characterize the complex interactions inside objectives while maintaining a lower hyper-parameter scale, theoretically alleviating the above deficiencies. The comprehensive analysis of the mathematical properties fully demonstrates that TW decomposition can be more potential in representation capabilities and more flexible in controlling both parameter storage and computational costs. To compute the TW-format decomposition, the sequential singular value decomposition (SVD)-based and the alternating least squares (ALS)-based learning algorithms are developed. Furthermore, to investigate the validity of TW decomposition, we provide its one numerical application, i.e., tensor completion (TC), yet develop an efficient proximal alternating minimization-based solving algorithm with guaranteed convergence. Experimental results on both synthetic and real-world data reveal that TW decomposition significantly outperforms other state-of-the-art tensor decompositions for incomplete-tensor inference, especially under solely few observations, thus substantiating the superiority and reliability of TW decomposition.
Zhong-Cheng Wu, Liang-Jian Deng, Ting-Zhu Huang, Hong-Xia Dou, Gemine Vivone, Yu Liu 0023
IEEE Trans. Image Process.2
2026 Digital Epidemiology With Awareness-Based Event-Triggered Migration in Networked Cyber-Physical Systems
abstract
Understanding how human mobility and information propagation influence the course of an epidemic remains a key challenge in digital epidemiology. In this work, we develop a new awareness-based, event-triggered epidemic model embedded within a networked Cyber-Physical System (CPS). In our framework, disease transmission and the dissemination of epidemic-related information evolve together on two interconnected layers. In detail, the physical layer models dis ease spread through human movement between two types of locations–residences and transfer stations-forming a bipartite metapopulation network. This structure captures the rendezvous effect, which reflects how gatherings in shared locations contribute to infection spread. The cyber layer represents the flow of information through digital communication networks. We introduce an event-triggered migration regulation mechanism, whereby individuals adapt their movement patterns based on local awareness thresholds, leading to a decentralized control process embedded within the network. Using a microscopic Markov chain approach (MMCA), we derive the epidemic threshold analytically and validate our results through extensive Monte Carlo simulations. Our findings show that event-triggered migration effectively suppresses the overall spread of the disease and lowers infection peaks-especially in heterogeneous populations and densely connected gathering points. These results demonstrate the potential of CPS-based epidemic models to enable real-time, awareness-driven interventions and to inform the design of decentralized control strategies that leverage digital communication dynamics.
Minyu Feng, Liang-Jian Deng, Matjaz Perc, Jürgen Kurths
IEEE Trans. Netw.3
2025 OTIAS: OcTree Implicit Adaptive Sampling for Multispectral and Hyperspectral Image Fusion
abstract
Implicit Neural Representation (INR) methods have demonstrated great potential in arbitrary-scale super-resolution tasks. This success is primarily due to their ability to continuously represent images using coordinates. In the task of remote sensing image fusion, INR methods have also shown promising applications. However, the previous INR methods neglect channel-wise modeling, while sharing a single kernel across all channels at each position, resulting in a lack of sensitivity to data specificity. To address these issues, we propose the OcTree Implicit Adaptive Sampling (OTIAS) method, which innovatively applies the octree structure to restore data from both horizontal and vertical directions, effectively incorporating spatial and spectral information from hyperspectral data. Additionally, we introduce a novel method to adaptively generate interpolation kernels based on coordinates. This approach efficiently produces customized interpolation kernel parameters for octree nodes, tailored to different spectral information. Overall, our method achieves state-of-the-art performance on the CAVE and Harvard datasets with 4× and 8× scaling factors, outperforming existing approaches.
Shangqi Deng, Liang-Jian Deng, Ping Wei 0001
AAAI3
2025 Wavelet-Assisted Multi-Frequency Attention Network for Pansharpening
abstract
Pansharpening aims to combine a high-resolution panchromatic (PAN) image with a low-resolution multispectral (LRMS) image to produce a high-resolution multispectral (HRMS) image. Although pansharpening in the frequency domain offers clear advantages, most existing methods either continue to operate solely in the spatial domain or fail to fully exploit the benefits of the frequency domain. To address this issue, we innovatively propose Multi-Frequency Fusion Attention (MFFA), which leverages wavelet transforms to cleanly separate frequencies and enable lossless reconstruction across different frequency domains. Then, we generate Frequency-Query, Spatial-Key, and Fusion-Value based on the physical meanings represented by different features, which enables a more effective capture of specific information in the frequency domain. Additionally, we focus on the preservation of frequency features across different operations. On a broader level, our network employs a wavelet pyramid to progressively fuse information across multiple scales. Compared to previous frequency domain approaches, our network better prevents confusion and loss of different frequency features during the fusion process. Quantitative and qualitative experiments on multiple datasets demonstrate that our method outperforms existing approaches and shows significant generalization capabilities for real-world scenarios.
Jie Huang 0005, Jinghao Xu, Siran Peng, Yule Duan 0001, Liang-Jian Deng
AAAI6
2025 PanAdapter: Two-Stage Fine-Tuning with Spatial-Spectral Priors Injecting for Pansharpening
abstract
Pansharpening is a challenging image fusion task that involves restoring images using two different modalities: low-resolution multispectral images (LRMS) and high-resolution panchromatic (PAN). Many end-to-end specialized models based on deep learning (DL) have been proposed, yet the scale and performance of these models are limited by the size of dataset. Given the superior parameter scales and feature representations of pre-trained models, they exhibit outstanding performance when transferred to downstream tasks with small datasets. Therefore, we propose an efficient fine-tuning method, namely PanAdapter, which utilizes additional advanced semantic information from pre-trained models to alleviate the issue of small-scale datasets in pansharpening tasks. Specifically, targeting the large domain discrepancy between image restoration and pansharpening tasks, the PanAdapter adopts a two-stage training strategy for progressively adapting to the downstream task. In the first stage, we fine-tune the pre-trained CNN model and extract task-specific priors at two scales by proposed Local Prior Extraction (LPE) module. In the second stage, we feed the extracted two-scale priors into two branches of cascaded adapters respectively. At each adapter, we design two parameter-efficient modules for allowing the two branches to interact and be injected into the frozen pre-trained VisionTransformer (ViT) blocks. We demonstrate that by only training the proposed LPE modules and adapters with a small number of parameters, our approach can benefit from pre-trained image restoration models and achieve state-of-the-art performance in several benchmark pansharpening datasets.
RuoCheng Wu, Zien Zhang, Shangqi Deng, Yule Duan 0001, Liang-Jian Deng
AAAI5
2025 Binarized Neural Network for Multi-spectral Image Fusion
abstract
Pan-sharpening technology refers to generating a high-resolution (HR) multi-spectral (MS) image with broad applications by fusing a low-resolution (LR) MS image and HR panchromatic (PAN) image. While deep learning approaches have shown impressive performance in pan-sharpening, they generally require extensive hardware with high memory and computational power, limiting their deployment on resource-constrained satellites. In this study, we investigate the use of binary neural networks (BNNs) for pan-sharpening and observe that binarization leads to distinct information degradation across different frequency components of an image. Building on this insight, we propose a novel binary pan-sharpening network, termed BNNPan, structured around the Prior-Integrated Binary Frequency (PIBF) module that features three key ingredients: Binary Wavelet Transform Convolution, Latent Diffusion Prior Compensation, and Channel-wise Distribution Calibration. Specifically, the first decomposes input features into distinct frequency components using Wavelet Transform, then applies a "divide-and-conquer" strategy to optimize binary feature learning for each component, informed by the corresponding full-precision residual statistics. The second integrates a latent diffusion prior to compensate for compromised information during binarization, while the third performs channel-wise calibration to further refine feature representation. Our BNNPan, developed with the proposed techniques, achieves promising pan-sharpening performance on multiple remote sensing datasets, surpassing state-of-the-art binarization algorithms.
Junming Hou, Ran Ran 0001, Xiaofeng Cong, Jian Wei You, Liang-Jian Deng
CVPR7
2025 A General Adaptive Dual-level Weighting Mechanism for Remote Sensing Pansharpening
abstract
Currently, deep learning-based methods for remote sensing pansharpening have advanced rapidly. However, many existing methods struggle to fully leverage feature heterogeneity and redundancy, thereby limiting their effectiveness. We use the covariance matrix to model the feature heterogeneity and redundancy and propose Correlation-Aware Covariance Weighting (CACW) to adjust them. CACW captures these correlations through the covariance matrix, which is then processed by a nonlinear function to generate weights for adjustment. Building upon CACW, we introduce a general adaptive dual-level weighting mechanism (ADWM) to address these challenges from two key perspectives, enhancing a wide range of existing deep-learning methods. First, Intra-Feature Weighting (IFW) evaluates correlations among channels within each feature to reduce redundancy and enhance unique information. Second, Cross-Feature Weighting (CFW) adjusts contributions across layers based on inter-layer correlations, refining the final output. Extensive experiments demonstrate the superior performance of ADWM compared to recent state-of-the-art (SOTA) methods. Furthermore, we validate the effectiveness of our approach through generality experiments, redundancy visualization, comparison experiments, key variables and complexity analysis, and ablation studies. Our code is available at https://github.com/Jie-1203/ADWM.
Jie Huang 0005, Haorui Chen, Jiaxuan Ren, Siran Peng, Liang-Jian Deng
CVPR5
2025 Adaptive Rectangular Convolution for Remote Sensing Pansharpening
abstract
Recent advancements in convolutional neural network (CNN)-based techniques for remote sensing pansharpening have markedly enhanced image quality. However, conventional convolutional modules in these methods have two critical drawbacks. First, the sampling positions in convolution operations are confined to a fixed square window. Second, the number of sampling points is preset and remains unchanged. Given the diverse object sizes in remote sensing images, these rigid parameters lead to suboptimal feature extraction. To overcome these limitations, we introduce an innovative convolutional module, Adaptive Rectangular Convolution (ARConv). ARConv adaptively learns both the height and width of the convolutional kernel and dynamically adjusts the number of sampling points based on the learned scale. This approach enables ARConv to effectively capture scale-specific features of various objects within an image, optimizing kernel sizes and sampling locations. Additionally, we propose ARNet, a network architecture in which ARConv is the primary convolutional module. Extensive evaluations across multiple datasets reveal the superiority of our method in enhancing pansharpening performance over previous techniques. Ablation studies and visualization further confirm the efficacy of ARConv. The source code can be available at https://github.com/WangXueyang-uestc/ARConv.
Zhixin Zheng, Jiandong Shao, Yule Duan 0001, Liang-Jian Deng
CVPR5
2025 Hyperspectral Pansharpening via Diffusion Models with Iteratively Zero-Shot Guidance
abstract
Hyperspectral pansharpening refers to fusing a panchromatic image (PAN) and a low-resolution hyperspectral image (LR-HSI) to obtain a high-resolution hyperspectral image (HR-HSI). Recently, guiding pre-trained diffusion models (DMs) has demonstrated significant potential in this area, leveraging their powerful representational abilities while avoiding complex training processes. However, these DMs are often trained on RGB images, not well-suited for pansharpening tasks, limited in adapting to the hyperspectral images. In this work, we propose a novel guided diffusion scheme with zero-shot guidance and neural spatialspectral decomposition (NSSD) to iteratively generate the RGB detail image and map the RGB detail image to target HR-HSI. Specifically, zero-shot guidance employs an auxiliary neural network that trained only with a PAN and LR-HSI to guide pre-trained DMs in generating the RGB detail image, informed by specific prior knowledge. Then, NSSD establishes a spectral mapping from the generated RGB detail image to the final HR-HSI. Extensive experiments are conducted on Pavia, Washington DC, Chukusei, and FR1 datasets to demonstrate that the proposed method significantly enhances the performance of DMs for hyperspectral pansharpening tasks, outperforming existing methods across multiple metrics and achieving improvements in visualization results. The code is available at https://github.com/Jin-liangXiao/DM-zs.
Jin-Liang Xiao, Ting-Zhu Huang, Liang-Jian Deng, Guang Lin 0002, Zihan Cao, Chao Li 0013, Qibin Zhao
CVPR3
2025 A Knowledge-driven Adaptive Collaboration of LLMs for Enhancing Medical Decision-making
abstract
Medical decision-making often involves integrating knowledge from multiple clinical specialties, typically achieved through multidisciplinary teams.Inspired by this collaborative process, recent work has leveraged large language models (LLMs) in multi-agent collaboration frameworks to emulate expert teamwork.While these approaches improve reasoning through agent interaction, they are limited by static, pre-assigned roles, which hinder adaptability and dynamic knowledge integration.To address these limitations, we propose KAMAC, a Knowledge-driven Adaptive Multi-Agent Collaboration framework that enables LLM agents to dynamically form and expand expert teams based on the evolving diagnostic context.KAMAC begins with one or more expert agents and then conducts a knowledge-driven discussion to identify and fill knowledge gaps by recruiting additional specialists as needed.This supports flexible, scalable collaboration in complex clinical scenarios, with decisions finalized through reviewing updated agent comments.Experiments on two real-world medical benchmarks demonstrate that KAMAC significantly outperforms both single-agent and advanced multi-agent methods, particularly in complex clinical scenarios (i.e., cancer prognosis) requiring dynamic, cross-specialty expertise.
Ting-Zhu Huang, Liang-Jian Deng, Yanyuan Qiao, Muhammad Imran Razzak, Yutong Xie 0001
EMNLP3
2025 Taming Flow Matching With Unbalanced Optimal Transport Into Fast Pansharpening
abstract
Pansharpening, a pivotal task in remote sensing for fusing high-resolution panchromatic and multispectral imagery, has garnered significant research interest. Recent advancements employing diffusion models based on stochastic differential equations (SDEs) have demonstrated state-of-the-art performance. However, the inherent multi-step sampling process of SDEs imposes substantial computational overhead, hindering practical deployment. While existing methods adopt efficient samplers, knowledge distillation, or retraining to reduce sampling steps (e.g., from 1,000 to fewer steps), such approaches often compromise fusion quality. In this work, we propose the Optimal Transport Flow Matching (OTFM) framework, which integrates the dual formulation of unbalanced optimal transport (UOT) to achieve one-step, high-quality pansharpening. Unlike conventional OT formulations that enforce rigid distribution alignment, UOT relaxes marginal constraints to enhance modeling flexibility, accommodating the intrinsic spectral and spatial disparities in remote sensing data. Furthermore, we incorporate task-specific regularization into the UOT objective, enhancing the robustness of the flow model. The OTFM framework enables simulation-free training and single-step inference while maintaining strict adherence to pansharpening constraints. Experimental evaluations across multiple datasets demonstrate that OTFM matches or exceeds the performance of previous regression-based models and leading diffusion-based methods while only needing one sampling step. Codes are available at https://github.com/294coder/PAN-OTFM.
Zihan Cao, Liang-Jian Deng
ICCV3
2025 MMAIF: Multi-Task and Multi-Degradation All-in-One for Image Fusion with Language Guidance
abstract
Image fusion, a fundamental low-level vision task, aims to integrate multiple image sequences into a single output while preserving as much information as possible from the input. However, existing methods face several significant limitations: 1) requiring task- or dataset-specific models; 2) neglecting real-world image degradations (\textit{e.g.}, noise), which causes failure when processing degraded inputs; 3) operating in pixel space, where attention mechanisms are computationally expensive; and 4) lacking user interaction capabilities. To address these challenges, we propose a unified framework for multi-task, multi-degradation, and language-guided image fusion. Our framework includes two key components: 1) a practical degradation pipeline that simulates real-world image degradations and generates interactive prompts to guide the model; 2) an all-in-one Diffusion Transformer (DiT) operating in latent space, which fuses a clean image conditioned on both the degraded inputs and the generated prompts. Furthermore, we introduce principled modifications to the original DiT architecture to better suit the fusion task. Based on this framework, we develop two versions of the model: Regression-based and Flow Matching-based variants. Extensive qualitative and quantitative experiments demonstrate that our approach effectively addresses the aforementioned limitations and outperforms previous restoration+fusion and all-in-one pipelines. Codes are available at https://github.com/294coder/MMAIF.
Zihan Cao, Liang-Jian Deng
ICCV4
2025 Physics-informed Neural Operator for Pansharpening
abstract
Over the past decades, pansharpening has contributed greatly to numerous remote sensing applications, with methods evolving from theoretically grounded models to deep learning approaches and their hybrids. Though promising, existing methods rarely address pansharpening through the lens of underlying physical imaging processes. In this work, we revisit the spectral imaging mechanism and propose a novel physics‐informed neural operator framework for pansharpening, termed PINO, which faithfully models the end‐to‐end electro‐optical sensor process. Specifically, PINO operates as: (1) First, a spatial-spectral encoder pair is introduced to aggregate multi-granularity high-resolution panchromatic (PAN) and low-resolution multispectral (LRMS) features. (2) Subsequently, an iterative neural integral process utilizes these fused spatial-spectral characteristics to learn a continuous radiance field $L_i(x, y, \lambda)$ over spatial coordinates and wavelength, effectively emulating band-wise spectral integration. (3) Finally, the learned radiance field is modulated by the sensor’s spectral responsivity $R_b(\lambda)$ to produce physically consistent spatial–spectral fusion products. This physics-grounded fusion paradigm offers a principled solution for reconstructing high-resolution multispectral and hyperspectral images in accordance with sensor imaging physics, effectively harnessing the unique advantages of spectral data to better uncover real-world characteristics. Experiments on multiple benchmark datasets show that our method surpasses state-of-the-art fusion algorithms, achieving reduced spectral aberrations and finer spatial textures. Furthermore, extension to hyperspectral (HS) data demonstrates its generalizability and universality. The code will be available upon potential acceptance.
Junming Hou, Chenxu Wu, Xiaofeng Cong, Shangqi Deng, Junling Li, Liang-Jian Deng
NeurIPS8
2025 Loss-driven dynamic weight and residual transformation in physics-informed neural network
Liang-Jian Deng
Neural Networks2
2025 An Efficient Image Fusion Network Exploiting Unifying Language and Mask Guidance
abstract
Image fusion aims to merge image pairs collected by different sensors over the same scene, preserving their distinct features. Recent works have often focused on designing various image fusion losses, developing different network architectures, and leveraging downstream tasks (e.g., object detection) for image fusion. However, a few studies have explored how language and semantic masks can serve as guidance to aid image fusion. In this paper, we investigate how the combination of language and masks can guide image fusion tasks, discarding the previously complex frameworks, which rely on downstream tasks, GAN-based cycle training, diffusion models, or deep image priors. Additionally, we exploit a recurrent neural network-like architecture to build a lightweight network that avoids the quadratic-cost of traditional attention mechanisms. To adapt the receptance weighted key value (RWKV) model to an image modality, we modify it into a bidirectional version using an efficient scanning strategy (ESS). To guide image fusion by language and mask features, we introduce a multi-modal fusion module (MFM) to facilitate information exchange. Comprehensive experiments show that the proposed framework achieved state-of-the-art results in various image fusion tasks (i.e., visible-infrared image fusion, multi-focus image fusion, multi-exposure image fusion, medical image fusion, hyperspectral and multispectral image fusion, and pansharpening).
Zihan Cao, Yu-Jie Liang, Liang-Jian Deng, Gemine Vivone
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Dynamic Evolution of Complex Networks: A Reinforcement Learning Approach Applying Evolutionary Games to Community Structure
abstract
Complex networks serve as abstract models for understanding real-world complex systems and provide frameworks for studying structured dynamical systems. This article addresses limitations in current studies on the exploration of individual birth-death and the development of community structures within dynamic systems. To bridge this gap, we propose a networked evolution model that includes the birth and death of individuals, incorporating reinforcement learning through games among individuals. Each individual has a lifespan following an arbitrary distribution, engages in games with network neighbors, selects actions using Q-learning in reinforcement learning, and moves within a two-dimensional space. The developed theories are validated through extensive experiments. Besides, we observe the evolution of cooperative behaviors and community structures in systems both with and without the birth-death process. The fitting of real-world populations and networks demonstrates the practicality of our model. Furthermore, comprehensive analyses of the model reveal that exploitation rates and payoff parameters determine the emergence of communities, learning rates affect the speed of community formation, discount factors influence stability, and two-dimensional space dimensions dictate community size. Our model offers a novel perspective on real-world community development and provides a valuable framework for studying population dynamics behaviors.
Bin Pi, Liang-Jian Deng, Minyu Feng, Matjaz Perc, Jürgen Kurths
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Fully-Connected Transformer for Multi-Source Image Fusion
abstract
Multi-source image fusion combines the information coming from multiple images into one data, thus improving imaging quality. This topic has aroused great interest in the community. How to integrate information from different sources is still a big challenge, although the existing self-attention based transformer methods can capture spatial and channel similarities. In this paper, we first discuss the mathematical concepts behind the proposed generalized self-attention mechanism, where the existing self-attentions are considered basic forms. The proposed mechanism employs multilinear algebra to drive the development of a novel fully-connected self-attention (FCSA) method to fully exploit local and non-local domain-specific correlations among multi-source images. Moreover, we propose a multi-source image representation embedding it into the FCSA framework as a non-local prior within an optimization problem. Some different fusion problems are unfolded into the proposed fully-connected transformer fusion network (FC-Former). More specifically, the concept of generalized self-attention can promote the potential development of self-attention. Hence, the FC-Former can be viewed as a network model unifying different fusion tasks. Compared with state-of-the-art methods, the proposed FC-Former method exhibits robust and superior performance, showing its capability of faithfully preserving information.
Zihan Cao, Ting-Zhu Huang, Liang-Jian Deng, Jocelyn Chanussot, Gemine Vivone
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 A fast Lanczos-based hierarchical algorithm for tensor ring decomposition
Cheng-Wei Sun, Ting-Zhu Huang, Hong-Xia Dou, Liang-Jian Deng
Signal Process.5
2025 Quaternion-based deep image prior with regularization by denoising for color image restoration
Liangtian He, Shaobing Gao, Liang-Jian Deng, Jun Liu 0012
Signal Process.4
2025 Nonlocal Tensor Wheel Decomposition for Hyperspectral Image Super-Resolution
Ting-Zhu Huang, Liang-Jian Deng
IEEE Signal Process. Lett.3
2025 MDFormer: Multi-Scale Downsampling-Based Transformer for Low-Light Image Enhancement
abstract
Vision Transformers have achieved impressive performance in the field of low-light image enhancement. Some Transformer-based methods acquire attention maps within channel dimension, whereas the spatial resolutions of queries and keys involved in matrix multiplication are much larger than the dimensions of channels. During the key-query dot-product interaction to generate attention maps, massive information redundancy and expensive computational costs are incurred. Simultaneously, most previous feed-forward networks in Transformers do not model the multi-range information that plays an important role for feature reconstruction. Based on the above observations, we propose an effective Multi-Scale Downsampling-Based Transformer (MDFormer) for low-light image enhancement, which consists of multi-scale downsampling-based self-attention (MDSA) and multi-range gated extraction block (MGEB). MDSA employs downsampling with two different factors for queries and keys to save the computational cost when implementing self-attention operations within channel dimension. Furthermore, we introduce learnable parameters for the two generated attention maps to adjust the weights for fusion, which allows MDSA to adaptively retain the most significant attention scores from attention maps. The proposed MGEB captures multi-range information by virtue of the multi-scale depth-wise convolutions and dilated convolutions, to enhance modeling capabilities. Extensive experiments on four challenging low-light image enhancement datasets demonstrate that our method outperforms the state-of-the-art.
Liangtian He, Liang-Jian Deng, Hongming Chen 0003, Chao Wang 0091
IEEE Signal Process. Lett.3
2025 Pansharpening Variational Model Based on Internal Adaptive Spatial Fidelity and External Deep-Driven Injection
abstract
Pansharpening is an image fusion technique that fuses the high spatial resolution of panchromatic images (PAN) and the rich spectral information of multispectral images (MS) to produce high resolution multispectral images (HRMS). The preservation of spatial details is crucial for enhancing the quality of the final results. However, existing detail extraction methods often fail to capture spatial information effectively. Most approaches rely only on internal details from the PAN image while overlooking external information, such as the deep-driven prior. Additionally, they struggle to establish an accurate relationship between the HRMS and PAN images, leading to spatial distortions. To address these issues, in this article, we propose a novel variational model based on double detail injection. Specifically, it integrates internal details from an adaptive spatial fidelity term and external details from a deep-driven injection term. Furthermore, an alternating direction method of multipliers (ADMM)-based algorithm is developed to efficiently solve the proposed model. The effectiveness of the proposed method is demonstrated through extensive experiments, showing superior performance compared to some existing pansharpening techniques.
Hong-Xia Dou, Jia-Lu Xu, Jin-Liang Xiao, Liang-Jian Deng
IEEE Trans. Geosci. Remote. Sens.4
2025 Spiking Variational Graph Representation Inference for Video Summarization
abstract
With the rise of short video content, efficient video summarization techniques for extracting key information have become crucial. However, existing methods struggle to capture the global temporal dependencies and maintain the semantic coherence of video content. Additionally, these methods are also influenced by noise during multi-channel feature fusion. We propose a Spiking Variational Graph (SpiVG) Network, which enhances information density and reduces computational complexity. First, we design a keyframe extractor based on Spiking Neural Networks (SNN), leveraging the event-driven computation mechanism of SNNs to learn keyframe features autonomously. To enable fine-grained and adaptable reasoning across video frames, we introduce a Dynamic Aggregation Graph Reasoner, which decouples contextual object consistency from semantic perspective coherence. We present a Variational Inference Reconstruction Module to address uncertainty and noise arising during multi-channel feature fusion. In this module, we employ Evidence Lower Bound Optimization (ELBO) to capture the latent structure of multi-channel feature distributions, using posterior distribution regularization to reduce overfitting. Experimental results show that SpiVG surpasses existing methods across multiple datasets such as SumMe, TVSum, VideoXum, and QFVS. Our codes and pre-trained models are available at https://github.com/liwrui/SpiVG.
Wenrui Li 0001, Wei Han 0002, Liang-Jian Deng, Ruiqin Xiong, Xiaopeng Fan 0001
IEEE Trans. Image Process.3
2025 TCJA-SNN: Temporal-Channel Joint Attention for Spiking Neural Networks
abstract
Spiking neural networks (SNNs) are attracting widespread interest due to their biological plausibility, energy efficiency, and powerful spatiotemporal information representation ability. Given the critical role of attention mechanisms in enhancing neural network performance, the integration of SNNs and attention mechanisms exhibits tremendous potential to deliver energy-efficient and high-performance computing paradigms. In this article, we present a novel temporal-channel joint attention mechanism for SNNs, referred to as TCJA-SNN. The proposed TCJA-SNN framework can effectively assess the significance of spike sequence from both spatial and temporal dimensions. More specifically, our essential technical contribution lies on: 1) we employ the squeeze operation to compress the spike stream into an average matrix. Then, we leverage two local attention mechanisms based on efficient 1-D convolutions to facilitate comprehensive feature extraction at the temporal and channel levels independently and 2) we introduce the cross-convolutional fusion (CCF) layer as a novel approach to model the interdependencies between the temporal and channel scopes. This layer effectively breaks the independence of these two dimensions and enables the interaction between features. Experimental results demonstrate that the proposed TCJA-SNN outperforms the state-of-the-art (SOTA) on all standard static and neuromorphic datasets, including Fashion-MNIST, CIFAR10, CIFAR100, CIFAR10-DVS, N-Caltech 101, and DVS128 Gesture. Furthermore, we effectively apply the TCJA-SNN framework to image generation tasks by leveraging a variation autoencoder. To the best of our knowledge, this study is the first instance where the SNN-attention mechanism has been employed for high-level classification and low-level generation tasks. Our implementation codes are available at https://github.com/ridgerchu/TCJA.
Rui-Jie Zhu 0003, Malu Zhang, Qihang Zhao, Yule Duan 0001, Liang-Jian Deng
IEEE Trans. Neural Networks Learn. Syst.6
2024 Gated Attention Coding for Training High-Performance and Efficient Spiking Neural Networks
abstract
Spiking neural networks (SNNs) are emerging as an energy-efficient alternative to traditional artificial neural networks (ANNs) due to their unique spike-based event-driven nature. Coding is crucial in SNNs as it converts external input stimuli into spatio-temporal feature sequences. However, most existing deep SNNs rely on direct coding that generates powerless spike representation and lacks the temporal dynamics inherent in human vision. Hence, we introduce Gated Attention Coding (GAC), a plug-and-play module that leverages the multi-dimensional gated attention unit to efficiently encode inputs into powerful representations before feeding them into the SNN architecture. GAC functions as a preprocessing layer that does not disrupt the spike-driven nature of the SNN, making it amenable to efficient neuromorphic hardware implementation with minimal modifications. Through an observer model theoretical analysis, we demonstrate GAC's attention mechanism improves temporal dynamics and coding efficiency. Experiments on CIFAR10/100 and ImageNet datasets demonstrate that GAC achieves state-of-the-art accuracy with remarkable efficiency. Notably, we improve top-1 accuracy by 3.10% on CIFAR100 with only 6-time steps and 1.07% on ImageNet while reducing energy usage to 66.9% of the previous works. To our best knowledge, it is the first time to explore the attention-based dynamic coding scheme in deep SNNs, with exceptional effectiveness and efficiency on large-scale datasets. Code is available at https://github.com/bollossom/GAC.
Xuerui Qiu, Rui-Jie Zhu 0003, Yuhong Chou, Zhaorui Wang 0005, Liang-Jian Deng, Guoqi Li 0002
AAAI5
2024 Content-Adaptive Non-Local Convolution for Remote Sensing Pansharpening
abstract
Currently, machine learning-based methods for remote sensing pansharpening have progressed rapidly. However, existing pansharpening methods often do not fully exploit differentiating regional information in non-local spaces, thereby limiting the effectiveness of the methods and resulting in redundant learning parameters. In this pa-per, we introduce a socalled content-adaptive non-local convolution (CANConv), a novel method tailored for re-mote sensing image pansharpening. Specifically, CANConv employs adaptive convolution, ensuring spatial adaptability, and incorporates non-local self-similarity through the similarity relationship partition (SRP) and the partition-wise adaptive convolution (PWAC) sub-modules. Furthermore, we also propose a corresponding network architecture, called CANNet, which mainly utilizes the multi-scale self-similarity. Extensive experiments demonstrate the superior performance of CANConv, compared with recent promising fusion methods. Besides, we substantiate the method's effectiveness through visualization, ablation experiments, and comparison with existing methods on multiple test sets. The source code is publicly available at https://github.com/duany11/CANConv.
Yule Duan 0001, Liang-Jian Deng
CVPR4
2024 Exploring the Low-Pass Filtering Behavior in Image Super-Resolution
abstract
Deep neural networks for image super-resolution (ISR) have shown significant advantages over traditional approaches like the interpolation. However, they are often criticized as ’black boxes’ compared to traditional approaches with solid mathematical foundations. In this paper, we attempt to interpret the behavior of deep neural networks in ISR using theories from the field of signal processing. First, we report an intriguing phenomenon, referred to as ‘the sinc phenomenon.’ It occurs when an impulse input is fed to a neural network. Then, building on this observation, we propose a method named Hybrid Response Analysis (HyRA) to analyze the behavior of neural networks in ISR tasks. Specifically, HyRA decomposes a neural network into a parallel connection of a linear system and a non-linear system and demonstrates that the linear system functions as a low-pass filter while the non-linear system injects high-frequency information. Finally, to quantify the injected high-frequency information, we introduce a metric for image-to-image tasks called Frequency Spectrum Distribution Similarity (FSDS). FSDS reflects the distribution similarity of different frequency components and can capture nuances that traditional metrics may overlook. Code, videos and raw experimental results for this paper can be found in: https://github.com/RisingEntropy/LPFInISR.
Zijing Xu, Yule Duan 0001, Liang-Jian Deng
ICML6
2024 A Novel Fidelity Based on the Adaptive Domain for Pansharpening
abstract
Pansharpening aims to obtain the high resolution multispectral image (HRMS) using the panchromatic image (PAN) and low spatial resolution multispectral image (LRMS). The similarity between PAN and HRMS has shown powerful performance for spatial feature extraction. The prevailing methods usually describe the similarity on a fixed transformed domain. However, such domain, e.g., gradient domain, usually limits the preservation of spatial details and neglects flexibility. To overcome these challenges, we propose an adaptive transformed domain-based spatial fidelity to depict the similarity accurately and flexibly. Based on the proposed spatial fidelity, we build a novel variational pansharpening model that consists of spectral and spatial fidelity terms. We design an algorithm based on the alternating direction method of multiplier (ADMM) framework to solve the model. Experimental results on reduced- and full-resolution data verify the effectiveness of the proposed method.
Jin-Liang Xiao, Ting-Zhu Huang, Liang-Jian Deng
IGARSS3
2024 A Novel State Space Model with Local Enhancement and State Sharing for Image Fusion
abstract
In image fusion tasks, images from different sources possess distinct characteristics.This has driven the development of numerous methods to explore better ways of fusing them while preserving their respective characteristics.Mamba, as a state space model, has emerged in the field of natural language processing.Recently, many studies have attempted to extend Mamba to vision tasks.However, due to the nature of images different from causal language sequences, the limited state capacity of Mamba weakens its ability to model image information.Additionally, the sequence modeling ability of Mamba is only capable of spatial information and cannot effectively capture the rich spectral information in images.Motivated by these challenges, we customize and improve the vision Mamba network designed for the image fusion task.Specifically, we propose the local-enhanced vision Mamba block, dubbed as LEVM.The LEVM block can improve local information perception of the network and simultaneously learn local and global spatial information.Furthermore, we propose the state sharing technique to enhance spatial details and integrate spatial and spectral information.Finally, the overall network is a multi-scale structure based on vision Mamba, called LE-Mamba.Extensive experiments show the proposed methods achieve state-of-the-art results on multispectral pansharpening and multispectral and hyperspectral image fusion datasets, and demonstrate the effectiveness of the proposed approach.Code can be accessed at https://github.com/294coder/Efficient-MIF.
Zihan Cao, Liang-Jian Deng
ACM Multimedia3
2024 Illumination Distribution Prior for Low-light Image Enhancement
abstract
In this paper, we propose a simple but effective illumination distribution prior (IDP) for images to illuminate the darkness. The illumination distribution prior is the product of a statistical approach to low-light images. It is based on a key factor - the mean value and standard deviation of images are positively correlated with the illumination. Using IDP in combination with the dual-domain feature fusion network (DFFN), we can obtain images that are more consistent with the ground truth distribution. DFFN inserts the discrete wavelet transform (DWT) into the transformer architecture, aiming to recover the detailed texture of the image through local high-frequency information and global spatial information. We have conducted extensive experiments on five widely used low-light image enhancement datasets and the experimental results show the superior performance of our proposed network (IDP-Net) compared to other state-of-the-art methods.
Chao Wang 0091, Liangtian He, Fenglai Lin, Hongming Chen 0003, Liang-Jian Deng
ACM Multimedia6
2024 Fourier-enhanced Implicit Neural Fusion Network for Multispectral and Hyperspectral Image Fusion
abstract
Recently, implicit neural representations (INR) have made significant strides in various vision-related domains, providing a novel solution for Multispectral and Hyperspectral Image Fusion (MHIF) tasks. However, INR is prone to losing high-frequency information and is confined to the lack of global perceptual capabilities. To address these issues, this paper introduces a Fourier-enhanced Implicit Neural Fusion Network (FeINFN) specifically designed for MHIF task, targeting the following phenomena: The Fourier amplitudes of the HR-HSI latent code and LR-HSI are remarkably similar; however, their phases exhibit different patterns. In FeINFN, we innovatively propose a spatial and frequency implicit fusion function (Spa-Fre IFF), helping INR capture high-frequency information and expanding the receptive field. Besides, a new decoder employing a complex Gabor wavelet activation function, called Spatial-Frequency Interactive Decoder (SFID), is invented to enhance the interaction of INR features. Especially, we further theoretically prove that the Gabor wavelet activation possesses a time-frequency tightness property that favors learning the optimal bandwidths in the decoder. Experiments on two benchmark MHIF datasets verify the state-of-the-art (SOTA) performance of the proposed method, both visually and quantitatively. Also, ablation studies demonstrate the mentioned contributions. The code can be available at https://github.com/294coder/Efficient-MIF.
Yu-Jie Liang, Zihan Cao, Shangqi Deng, Hong-Xia Dou, Liang-Jian Deng
NeurIPS5
2024 SSDiff: Spatial-spectral Integrated Diffusion Model for Remote Sensing Pansharpening
abstract
Pansharpening is a significant image fusion technique that merges the spatial content and spectral characteristics of remote sensing images to generate high-resolution multispectral images. Recently, denoising diffusion probabilistic models have been gradually applied to visual tasks, enhancing controllable image generation through low-rank adaptation (LoRA). In this paper, we introduce a spatial-spectral integrated diffusion model for the remote sensing pansharpening task, called SSDiff, which considers the pansharpening process as the fusion process of spatial and spectral components from the perspective of subspace decomposition. Specifically, SSDiff utilizes spatial and spectral branches to learn spatial details and spectral features separately, then employs a designed alternating projection fusion module (APFM) to accomplish the fusion. Furthermore, we propose a frequency modulation inter-branch module (FMIM) to modulate the frequency distribution between branches. The two components of SSDiff can perform favorably against the APFM when utilizing a LoRA-like branch-wise alternative fine-tuning method. It refines SSDiff to capture component-discriminating features more sufficiently. Finally, extensive experiments on four commonly used datasets, i.e., WorldView-3, WorldView-2, GaoFen-2, and QuickBird, demonstrate the superiority of SSDiff both visually and quantitatively. The code is available at https://github.com/Z-ypnos/SSdiff_main.
Liang-Jian Deng, Zihan Cao, Hong-Xia Dou
NeurIPS3
2024 A General Paradigm with Detail-Preserving Conditional Invertible Network for Image Fusion
Liang-Jian Deng, Ran Ran 0001, Gemine Vivone
Int. J. Comput. Vis.2
2024 Tensor decomposition based attention module for spiking neural networks
Rui-Jie Zhu 0003, Xuerui Qiu, Yule Duan 0001, Malu Zhang, Liang-Jian Deng
Knowl. Based Syst.6
2024 Remote Sensing Image Destriping by an ℓ₀-Based Nonconvex Model With Overlapping Group Sparse Hyper-Laplacian Prior
abstract
In this paper, we propose an ℓ0-based nonconvex optimization model with overlapping group sparse hyper-Laplacian prior (ℓ0-OGSHL) to remove stripes from remote sensing images (RSIs) effectively. Specifically, we utilize the hyper-Laplacian prior with overlapping group sparsity (OGSHL) to characterize the properties of the underlying image. Additionally, the related ℓ0-quasi equivalent is transformed into an easily solvable form by employing a mathematical program with equilibrium constraints (MPEC). Furthermore, the alternating direction method of multipliers (ADMM) algorithm is employed for resolving the equivalent nonconvex optimization model, and the complex OGSHL subproblem is addressed through the majorization-minimization (MM) method. Finally, the experimental results on the simulated datasets conclusively demonstrate the superior performance of the proposed method over the compared methods (with 1 3dB higher MPSNR), both quantitatively and visually. The code will be available after possible acceptance.
Hong-Xia Dou, Yong Chen 0013, Jun Liu 0012, Liang-Jian Deng
IEEE Geosci. Remote. Sens. Lett.6
2024 SSCAConv: Self-Guided Spatial-Channel Adaptive Convolution for Image Fusion
abstract
Pansharpening, which attempts to obtain a high-resolution multispectral (HR-MS) image by fusing a panchromatic (PAN) image with a low-resolution multispectral (LR-MS) image, is a critical yet difficult remote sensing image processing task. In this study, we present a novel convolution operation, self-guided spatial-channel adaptive convolution (SSCAConv), for pansharpening. Unlike the reported adaptive convolutions that only focus on spatial details, our SSCAConv also considers channel specificity by generating an individual convolution kernel for each channel patch according to its own content and supplements the interchannel information by introducing a global bias. We further apply the designed SSCAConv to a simple residual network architecture to construct the image fusion network (SSCANet). Experimental results show that SSCANet outperforms state-of-the-art (SOTA) pansharpening algorithms and achieves better generalization ability with fewer parameters. In addition, our network also yields the best results when extended to the hyperspectral image super-resolution (HISR) problem. The code is available athttps://github.com/Pluto-wei/SSCAConv.
Xiaoya Lu, Yu-Wei Zhuo, Hongming Chen 0003, Liang-Jian Deng, Junming Hou
IEEE Geosci. Remote. Sens. Lett.4
2024 CMT: Cross Modulation Transformer With Hybrid Loss for Pansharpening
abstract
Pansharpening aims to enhance remote sensing image (RSI) quality by merging high-resolution panchromatic (PAN) with multispectral (MS) images. However, prior techniques struggled to optimally fuse PAN and MS images for enhanced spatial and spectral information, due to a lack of a systematic framework capable of effectively coordinating their individual strengths. In response, we present the cross modulation Transformer (CMT), a pioneering method that modifies the attention mechanism. This approach utilizes a robust modulation technique from signal processing, integrating it into the attention mechanism’s calculations. It dynamically tunes the weights of the carrier’s value (V) matrix according to the modulator’s features, thus resolving historical challenges and achieving a seamless integration of spatial and spectral attributes. Furthermore, considering that RSI exhibit large-scale features and edge details along with local textures, we crafted a hybrid loss function that combines Fourier and wavelet transforms to effectively capture these characteristics, thereby enhancing both spatial and spectral accuracy in pansharpening. Extensive experiments demonstrate our framework’s superior performance over existing state-of-the-art methods. The source code is publicly available athttps://github.com/WenjieShu/CMT.
Hong-Xia Dou, Liang-Jian Deng
IEEE Geosci. Remote. Sens. Lett.5
2024 Denoiser-guided image deconvolution with arbitrary boundaries and incomplete observations
Liangtian He, Shaobing Gao, Liang-Jian Deng, Yilun Wang 0004, Chao Wang 0091
Signal Process.3
2024 Quaternion weighted Schatten p-norm minimization for color image restoration with convergence guarantee
Liangtian He, Yilun Wang 0004, Liang-Jian Deng, Jun Liu 0012
Signal Process.4
2024 Rethinking Pan-Sharpening via Spectral-Band Modulation
abstract
Pan-sharpening aims to super-resolve the low-resolution (LR) multispectral (MS) image under the guidance of a high-resolution (HR) panchromatic (PAN) image. Existing deep learning (DL)-based pan-sharpening methods usually adhere to a common philosophy of learning complementary information between MS and PAN images. Despite remarkable advances, few studies consider the band-private characteristics which differ greatly from band to band. An ideal MS image, however, is jointly determined by its diverse spectral bands, thus the accurate restoration of every band will benefit the pan-sharpening performance. In this work, we propose a novel yet effective solution to reconstruct the HRMS image by explicitly modulating every spectral band under the conditions of the PAN image. As a result, we design a spatially-adaptive spectral modulation network, dubbed SSMNet, which consists of three core designs: source-aware spectral modulator (SSM), cross-band information aggregation (CBIA) module, and cross-stage feature integration (CSFI) module. The first predicts a series of spatially-adaptive kernels to capture the local information of every spectral band. Followed by, the second is responsible for facilitating the information communication among various bands to guarantee continuous spectral representations. Furthermore, the third attends to integrate the cross-stage output features to produce the pan-sharpened result. In addition, we also introduce the histogram loss to constrain the band-wise distribution of the final fused products. Extensive experiments demonstrate that our SSMNet achieves favorable performance against other state-of-the-art (SOTA) methods on multiple satellite datasets. The code is available athttps://github.com/ez4lionky/SSMNet/.
Junming Hou, Xiaofeng Cong, Hao Shen 0006, Zhuochen Lou, Liang-Jian Deng, Jian Wei You
IEEE Trans. Geosci. Remote. Sens.6
2024 FusionMamba: Efficient Remote Sensing Image Fusion With State Space Model
abstract
Remote sensing image fusion aims to generate a high-resolution multi/hyperspectral image by combining a high-resolution image with limited spectral data and a low-resolution image rich in spectral information. Current deep learning (DL) methods typically employ convolutional neural networks (CNNs) or Transformers for feature extraction and information integration. While CNNs are efficient, their limited receptive fields restrict their ability to capture global context. Transformers excel at learning global information but are computationally expensive. Recent advancements in the state space model (SSM), particularly Mamba, present a promising alternative by enabling global perception with low complexity. However, the potential of SSM for information integration remains largely unexplored. Therefore, we propose FusionMamba, an innovative method for efficient remote sensing image fusion. Our contributions are twofold. First, to effectively merge spatial and spectral features, we expand the single-input Mamba block to accommodate dual inputs, creating the FusionMamba block, which serves as a plug-and-play solution for information integration. Second, we incorporate Mamba and FusionMamba blocks into an interpretable network architecture tailored for remote sensing image fusion. Our designs utilize two U-shaped network branches, each primarily composed of four-directional (FD) Mamba blocks, to extract spatial and spectral features separately and hierarchically. The resulting feature maps are sufficiently merged in an auxiliary network branch constructed with FusionMamba blocks. Furthermore, we improve the representation of spectral information through an enhanced channel attention module. Quantitative and qualitative valuation results across six datasets demonstrate that our method achieves the state-of-the-art (SOTA) performance, underscoring the effectiveness of FusionMamba. The code is available athttps://github.com/PSRben/FusionMamba.
Siran Peng, Xiangyu Zhu 0001, Liang-Jian Deng, Zhen Lei 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 CroDoSR: Tensor Cross-Domain Rank for Hyperspectral Image Super-Resolution
abstract
Hyperspectral image super-resolution (HSI SR) aims to combine the detailed spectral information of hyperspectral images with the spatial resolution of multispectral images, thus enhancing the ability to extract valuable insights across various applications. Recently, the tensor singular value decomposition (t-SVD) has emerged as a powerful tool and has been introduced into the HSI SR field for exploring low-rank prior information. For t-SVD, the domain transform is crucial to acquiring more low-rank data characteristics. Nevertheless, previous efforts on domain transform have only involved the single transformed domain (i.e., single domain), while ignoring the potential pursuing the lower rankness in multiple successional transformed domains, termed cross-domain (CD). In this article, we propose a novel CD-based t-SVD and define the corresponding tensor CD rank based on a pivotal observation, i.e., the low-rank behavior of HSI in CD is more significant than that in single domain. More specifically, we first define a successional linear transform (SLT) to establish the CD concept, then develop a novel CD-based t-SVD and tensor CD rank, and theoretically deduce a new tensor CD-nuclear norm as the convex approximation of CD rank. Equipped with such a CD rank, we thus formulate a CD-rank-constrained minimization model for the HSI SR task, called CroDoSR, which is effectively solved by the alternating direction method of multipliers (ADMMs). Comprehensive experiments on several widely used datasets evidently demonstrate the superiority of the proposed CroDoSR method.
Zhong-Cheng Wu, Ting-Zhu Huang, Liang-Jian Deng, Gemine Vivone
IEEE Trans. Geosci. Remote. Sens.4
2024 A Coupled Tensor Double-Factor Method for Hyperspectral and Multispectral Image Fusion
abstract
Hyperspectral and multispectral image fusion, denoted as HSI-MSI fusion, involves merging a pair of hyperspectral (HSI) and multispectral (MSI) images to generate a high spatial resolution hyperspectral image (HR-HSI). The primary challenge in HSI-MSI fusion is to find the best way to extract one-dimensional spectral features and two-dimensional (2-D) spatial features from HSI and MSI and harmoniously combine them. In recent times, coupled tensor decomposition (CTD)-based methods have shown promising performance in the fusion task. However, the tensor decompositions (TDs) used by these CTD-based methods face difficulties in extracting complex features and capturing 2-D spatial features, resulting in suboptimal fusion results. To address these issues, we introduce a novel method called Coupled Tensor Double-Factor Decomposition (CTDF). Specifically, we propose a Tensor Double-Factor (TDF) decomposition, representing a 3rd-order HR-HSI as a 4th-order spatial factor and a 3rd-order spectral factor, connected through tensor contraction. Compared to other TDs, the TDF has better feature extraction capability since it has a higher order factor than that of HR-HSI, whereas the other TDs only have the same order factor as the HR-HSI. Moreover, the TDF can extract 2-D spatial features using the 4th-order spatial factor. We apply the TDF to the HSI-MSI fusion problem and formulate the CTDF model. Furthermore, we design an algorithm based on proximal alternating minimization to solve this model and provide insights into its computational complexity and convergence analysis. The simulated and real experiments validate the effectiveness and efficiency of the proposed CTDF method. The code is available at https://github.com/tingxu113/CTDF.
Ting-Zhu Huang, Liang-Jian Deng, Jin-Liang Xiao, Clifford Broni-Bediako, Junshi Xia, Naoto Yokoya
IEEE Trans. Geosci. Remote. Sens.3
2024 KNLConv: Kernel-Space Non-Local Convolution for Hyperspectral Image Super-Resolution
abstract
Pixel-level adaptive convolution, which overcomes the deficiency of the spatial-invariance of standard convolution, is always limited to performing feature extraction from local patches and ignores the latent long-range dependencies imperceptible in the feature space, which are more significant in pixel-level tasks such as hyperspectral image super-resolution (HSISR). To handle such limitations, we propose kernel-space non-local convolution (KNLConv), which explores non-local dependencies in the generated kernel space, to leverage these global information to guide the network to extract image features more flexibly. Technically, the proposed KNLConv first decomposes the convolutional kernel space into spatial and channel dimensions, and designs a depth-wise non-local expansion convolution (NLEC) in the spatial dimension of the kernel-space to explore underlying global correlations. Then introduce an adaptive point-wise convolution (APC), generalizing the NLEC to the pixel-level while integrating features in the channel dimension. In addition, applying KNLConv, we design an effective network architecture for hyperspectral image super-resolution. Extensive experiments demonstrate that our approach performs favorably against current state-of-the-art HSISR methods, both on quantitative indicators and visual quality.
Ran Ran 0001, Liang-Jian Deng, Tianjing Zhang, Jianlong Chang, Qi Tian 0001
IEEE Trans. Multim.2
2024 An Evolutionary Game With the Game Transitions Based on the Markov Process
abstract
The psychology of the individual is continuously changing in nature, which has a significant influence on the evolutionary dynamics of populations. To study the influence of the continuously changing psychology of individuals on the behavior of populations, in this article, we consider the game transitions of individuals in evolutionary processes to capture the changing psychology of individuals in reality, where the game that individuals will play shifts as time progresses and is related to the transition rates between different games. Besides, the individual’s reputation is taken into account and utilized to choose a suitable neighbor for the strategy updating of the individual. Within this model, we investigate the statistical number of individuals staying in different game states and the expected number fits well with our theoretical results. Furthermore, we explore the impact of transition rates between different game states, payoff parameters, the reputation mechanism, and different time scales of strategy updates on cooperative behavior, and our findings demonstrate that both the transition rates and reputation mechanism have a remarkable influence on the evolution of cooperation. Additionally, we examine the relationship between network size and cooperation frequency, providing valuable insights into the robustness of the model.
Minyu Feng, Bin Pi, Liang-Jian Deng, Jürgen Kurths
IEEE Trans. Syst. Man Cybern. Syst.3
2023 Modality-Fusion Spiking Transformer Network for Audio-Visual Zero-Shot Learning
abstract
Audio-visual zero-shot learning (ZSL), which learns to classify video data from the classes not being observed during training, is challenging. In audio-visual ZSL, both semantic and temporal information from different modalities is relevant to each other. However, effectively extracting and fusing information from audio and visual remains an open challenge. In this work, we propose an Audio-Visual Modality-fusion Spiking Transformer network (AVMST) for audio-visual ZSL. To be more specific, AVMST provides a spiking neural network (SNN) module for extracting conspicuous temporal information of each modality, a cross-attention block to effectively fuse the temporal and semantic information, and a transformer reasoning module to further explore the interrelationships of fusion features. To provide robust temporal features, the spiking threshold of the SNN module is adjusted dynamically based on the semantic cues of different modalities. The generated feature map is in accordance with the zero-shot learning property thanks to our proposed spiking transformer’s ability to combine the robustness of SNN feature extraction and the precision of transformer feature inference. Extensive experiments on three benchmark audio-visual datasets (i.e., VGGSound, UCF and ActivityNet) validate that the proposed AVMST outperforms existing state-of-the-art methods by a significant margin. The code and pre-trained models are available at https://github.com/liwr-hit/ICME23_AVMST.
Wenrui Li 0001, Zhengyu Ma, Liang-Jian Deng, Hengyu Man, Xiaopeng Fan 0001
ICME3
2023 Bidirectional Dilation Transformer for Multispectral and Hyperspectral Image Fusion
abstract
Transformer-based methods have proven to be effective in achieving long-distance modeling, capturing the spatial and spectral information, and exhibiting strong inductive bias in various computer vision tasks. Generally, the Transformer model includes two common modes of multi-head self-attention (MSA): spatial MSA (Spa-MSA) and spectral MSA (Spe-MSA). However, Spa-MSA is computationally efficient but limits the global spatial response within a local window. On the other hand, Spe-MSA can calculate channel self-attention to accommodate high-resolution images, but it disregards the crucial local information that is essential for low-level vision tasks. In this study, we propose a bidirectional dilation Transformer (BDT) for multispectral and hyperspectral image fusion (MHIF), which aims to leverage the advantages of both MSA and the latent multiscale information specific to MHIF tasks. The BDT consists of two designed modules: the dilation Spa-MSA (D-Spa), which dynamically expands the spatial receptive field through a given hollow strategy, and the grouped Spe-MSA (G-Spe), which extracts latent features within the feature map and learns local data behavior. Additionally, to fully exploit the multiscale information from both inputs with different spatial resolutions, we employ a bidirectional hierarchy strategy in the BDT, resulting in improved performance. Finally, extensive experiments on two commonly used datasets, CAVE and Harvard, demonstrate the superiority of BDT both visually and quantitatively. Furthermore, the related code will be available at the GitHub page of the authors.
Shangqi Deng, Liang-Jian Deng, Ran Ran 0001
IJCAI2
2023 LGPConv: Learnable Gaussian Perturbation Convolution for Lightweight Pansharpening
abstract
Pansharpening is a crucial and challenging task that aims to obtain a high spatial resolution image by merging a multispectral (MS) image and a panchromatic (PAN) image. Current methods use CNNs with standard convolution, but we've observed strong correlation among channel dimensions in the kernel, leading to computational burden and redundancy. To address this, we propose Learnable Gaussian Perturbation Convolution (LGPConv), surpassing standard convolution. LGPConv leverages two properties of standard convolution kernels: 1) correlations within channels, learning a premier kernel as a base to reduce parameters and training difficulties caused by redundancy; 2) introducing Gaussian noise perturbations to simulate randomness and enhance nonlinear representation within channels. We incorporate LGPConv into a well-designed pansharpening network and demonstrate its superiority through extensive experiments, achieving state-of-the-art performance with minimal parameters (27K). Code is available on the GitHub page of the authors.
Chen-Yu Zhao, Tianjing Zhang, Ran Ran 0001, Zhi-Xuan Chen, Liang-Jian Deng
IJCAI5
2023 Bidomain Modeling Paradigm for Pansharpening
abstract
Pansharpening is a challenging low-level vision task whose aim is to learn the complementary representation between spectral information and spatial detail. Despite the remarkable progress, existing deep neural network (DNN) based pansharpening algorithms are still confronted with common limitations. 1) These methods rarely consider the local specificity of different spectral bands; 2) They often extract the global detail in the spatial domain, which ignore the task-related degradation, e.g., the down-sampling process of MS image, and also suffer from limited receptive field. In this work, we propose a novel bidomain modeling paradigm for pansharpening problem (dubbed as BiMPan), which takes into both local spectral specificity and global spatial detail. More specifically, we first customize the specialized source-discriminative adaptive convolution (SDAConv) for every spectral band instead of sharing the identical kernels across all bands like prior works. Then, we devise a novel Fourier global modeling module (FGMM), which is capable of embracing global information while benefiting the disentanglement of image degradation. By integrating the band-aware local feature and Fourier global detail from these two functional designs, we can fuse a texture-rich while visually pleasing high-resolution MS image. Extensive experiments demonstrate that the proposed framework achieves favorable performance against current state-of-the-art pansharpening methods. The code is available at https://github.com/coder-qicao/BiMPan.
Junming Hou, Ran Ran 0001, Che Liu 0004, Junling Li, Liang-Jian Deng
ACM Multimedia6
2023 Reservoir Computing Transformer for Image-Text Retrieval
abstract
Although the attention mechanism in transformers has proven successful in image-text retrieval tasks, most transformer models suffer from a large number of parameters. Inspired by brain circuits that process information with recurrent connected neurons, we propose a novel Reservoir Computing Transformer Reasoning Network (RCTRN) for image-text retrieval. The proposed RCTRN employs a two-step strategy to focus on feature representation and data distribution of different modalities respectively. Specifically, we send visual and textual features through a unified meshed reasoning module, which encodes multi-level feature relationships with prior knowledge and aggregates the complementary outputs in a more effective way. The reservoir reasoning network is proposed to optimize memory connections between features at different stages and address the data distribution mismatch problem introduced by the unified scheme. To investigate the significance of the low power dissipation and low bandwidth characteristics of RRN in practical scenarios, we deployed the model in the wireless transmission system, demonstrating that RRN's optimization of data structures also has a certain robustness against channel noise. Extensive experiments on two benchmark datasets, Flickr30K and MS-COCO, demonstrate the superiority of RCTRN in terms of performance and low-power dissipation compared to state-of-the-art baselines.
Wenrui Li 0001, Zhengyu Ma, Liang-Jian Deng, Penghong Wang, Jinqiao Shi, Xiaopeng Fan 0001
ACM Multimedia3
2023 U2Net: A General Framework with Spatial-Spectral-Integrated Double U-Net for Image Fusion
abstract
In image fusion tasks, images obtained from different sources exhibit distinct properties. Consequently, treating them uniformly with a single-branch network can lead to inadequate feature extraction. Additionally, numerous works have demonstrated that multi-scaled networks capture information more sufficiently than single-scaled models in pixel-level computer vision problems. Considering these factors, we propose U2Net, a spatial-spectral-integrated double U-shape network for image fusion. The U2Net utilizes a spatial U-Net and a spectral U-Net to extract spatial details and spectral characteristics, which allows for the discriminative and hierarchical learning of features from diverse images. In contrast to most previous works that merely employ concatenation to merge spatial and spectral information, this paper introduces a novel spatial-spectral integration structure called S2Block, which combines feature maps from different sources in a logical and effective way. We conduct a series of experiments on two image fusion tasks, including remote sensing pansharpening and hyperspectral image super-resolution (HISR). The U2Net outperforms representative state-of-the-art (SOTA) approaches in both quantitative and qualitative evaluations, demonstrating the superiority of our method. The code is available at https://github.com/PSRben/U2Net.
Siran Peng, Chenhao Guo, Liang-Jian Deng
ACM Multimedia4
2023 Dynamical Fusion Model With Joint Variational and Deep Priors for Hyperspectral Image Super-Resolution
abstract
In this paper, we propose a novel dynamic fusion model (DFM) with joint variational and deep priors for the task of hyperspectral image super-resolution (HISR). The given model can benefit from both the advantages of traditional modeling and deep learning methods, thus achieving significant improvements based on existing deep pre-trained models. Specifically, the given model mainly contains two new designed terms, i.e., the weighted spatial fidelity (WSF) term and the deep fusion (DF) term. The WSF term focuses on the spatial recovery of the low-resolution hyperspectral image through the high-resolution multispectral image without the knowledge of the spectral response matrix, thus the proposed DFM can be viewed as a semi-blind model for HISR. Moreover, the DF term relied upon deep fusion with a designed adaptive weight matrix, which can effectively inject the deep priors into the traditional minimization model. Besides, the proposed DFM can be quickly and effectively solved using the alternating direction method of multipliers. Experimental results on widely used datasets demonstrate the superiority of our approach compared with state-of-the-art HISR methods.
Hong-Xia Dou, Zhong-Cheng Wu, Yu-Wei Zhuo, Liang-Jian Deng, Gemine Vivone
IEEE Geosci. Remote. Sens. Lett.4
2023 FDDN: frequency-guided network for single image dehazing
Haozhen Shen, Chao Wang 0102, Liang-Jian Deng, Liangtian He, Ming-Wen Shao, Deyu Meng
Neural Comput. Appl.3
2023 Neuron-Based Spiking Transmission and Reasoning Network for Robust Image-Text Retrieval
abstract
Most of the image-text retrieval methods carry out accurate results using fine-grained features for feature alignment. However, extracting the robustness features while maintaining the retrieval accuracy in wireless communication is still a challenge, especially with channel noises and limited transmission bandwidth. Inspired by spike signals of neurons in the human brain, we propose the neuron-based spiking transmission and reasoning network (NSTRN). In this way, the features are compressed into compacted efficient representations. In NSTRN, we construct the feature sender based on spiking activation function to selectively encode only important information in images and sentences into binary codes, and reduce the transmission cost. Moreover, the feature receiver is designed as a recurrent architecture and applies both temporal attention and global attention blocks to memorize long-term information. Finally, to compensate for the loss of visual concepts in transmission, we use the global textual features as coefficients to guide the formation of visual features in the training stage. The traditional CNN-based joint source-channel coding model outputs float-point encoded features, which requires additional quantization steps to convert features into binary bitstreams in the practical wireless communication system. Instead, the spiking neural networks (SNNs) directly use binary spike trains to reduce the computation complexity caused by the quantization steps. More importantly, SNNs can naturally encode the asynchronous event streams and inhibit the discrete noisy events to extract robust information. Even with binary bitstreams, NSTRN shows effectiveness compared with the state-of-the-art image-text retrieval methods. In the wireless communication scenario, NSTRN not only reduces the transmission bandwidth but also alleviates the “cliff effect” to a certain extent in the traditional separate encoding methods. To the best of our knowledge, this is the first work using SNNs on robust image-text retrieval.
Wenrui Li 0001, Zhengyu Ma, Liang-Jian Deng, Xiaopeng Fan 0001, Yonghong Tian 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 GuidedNet: A General CNN Fusion Framework via High-Resolution Guidance for Hyperspectral Image Super-Resolution
abstract
Hyperspectral image super-resolution (HISR) is about fusing a low-resolution hyperspectral image (LR-HSI) and a high-resolution multispectral image (HR-MSI) to generate a high-resolution hyperspectral image (HR-HSI). Recently, convolutional neural network (CNN)-based techniques have been extensively investigated for HISR yielding competitive outcomes. However, existing CNN-based methods often require a huge amount of network parameters leading to a heavy computational burden, thus, limiting the generalization ability. In this article, we fully consider the characteristic of the HISR, proposing a general CNN fusion framework with high-resolution guidance, called GuidedNet. This framework consists of two branches, including 1) the high-resolution guidance branch (HGB) that can decompose the high-resolution guidance image into several scales and 2) the feature reconstruction branch (FRB) that takes the low-resolution image and the multiscaled high-resolution guidance images from the HGB to reconstruct the high-resolution fused image. GuidedNet can effectively predict the high-resolution residual details that are added to the upsampled HSI to simultaneously improve spatial quality and preserve spectral information. The proposed framework is implemented using recursive and progressive strategies, which can promote high performance with a significant network parameter reduction, even ensuring network stability by supervising several intermediate outputs. Additionally, the proposed approach is also suitable for other resolution enhancement tasks, such as remote sensing pansharpening and single-image super-resolution (SISR). Extensive experiments on simulated and real datasets demonstrate that the proposed framework generates state-of-the-art outcomes for several applications (i.e., HISR, pansharpening, and SISR). Finally, an ablation study and more discussions assessing, for example, the network generalization, the low computational cost, and the fewer network parameters, are provided to the readers. The code link is: https://github.com/Evangelion09/GuidedNet.
Ran Ran 0001, Liang-Jian Deng, Tai-Xiang Jiang, Jin-Fan Hu, Jocelyn Chanussot, Gemine Vivone
IEEE Trans. Cybern.2
2023 PSRT: Pyramid Shuffle-and-Reshuffle Transformer for Multispectral and Hyperspectral Image Fusion
abstract
A Transformer has received a lot of attention in computer vision. Because of global self-attention, the computational complexity of Transformer is quadratic with the number of tokens, leading to limitations for practical applications. Hence, the computational complexity issue can be efficiently resolved by computing the self-attention in groups of smaller fixed-size windows. In this article, we propose a novel pyramid Shuffle-and-Reshuffle Transformer (PSRT) for the task of multispectral and hyperspectral image fusion (MHIF). Considering the strong correlation among different patches in remote sensing images and complementary information among patches with high similarity, we design Shuffle-and-Reshuffle (SaR) modules to consider the information interaction among global patches in an efficient manner. Besides, using pyramid structures based on window self-attention, the detail extraction is supported. Extensive experiments on four widely used benchmark datasets demonstrate the superiority of the proposed PSRT with a few parameters compared with several state-of-the-art approaches. The related code is available athttps://github.com/Deng-shangqi/PSRT.
Shangqi Deng, Liang-Jian Deng, Ran Ran 0001, Danfeng Hong, Gemine Vivone
IEEE Trans. Geosci. Remote. Sens.2
2023 Cascadic Multireceptive Learning for Multispectral Pansharpening
abstract
Pansharpening refers to the fusion of a panchromatic (PAN) image with high spatial resolution and a multispectral (LRMS) image with low spatial resolution to obtain a high spatial resolution multispectral (HRMS) image, which is beneficial to visual display and geographic research. Recently, many deep learning (DL) methods have been proposed to address the pansharpening problem, but still a few examples of DL-based techniques are designed from the perspective of a better receptive field while the scale of features greatly varies among different ground objects. In this paper, we mainly focus on designing a cascadic multi-receptive learning module (CML-resblock) relying on the ResNet block, which can efficiently extract multi-scale features from both the PAN and LRMS images. Moreover, we propose a novel multiplication network preserving a physical significance, which uses deep neural networks (DNNs) to learn the coefficients of the pixel-wise restoration mapping and multiplies the up-sampled LRMS image with the learned coefficients to get the HRMS image. The two parts mentioned above constitute our cascadic multi-receptive learning network (CMLNet). Extensive experiments on both reduced-resolution and full-resolution images acquired by the WorldView-3, GaoFen-2, and QuickBird satellites show that the proposed approach outperforms state-of-the-art methods. Furthermore, additional experiments have been conducted to prove the generality of the CML-resblock and multiplication network. The code is available at: https://github.com/wajuda/CML.
Jun-Da Wang, Liang-Jian Deng, Chen-Yu Zhao, Hongming Chen 0003, Gemine Vivone
IEEE Trans. Geosci. Remote. Sens.2
2023 A Novel Spatial Fidelity With Learnable Nonlinear Mapping for Panchromatic Sharpening
abstract
The purpose of panchromatic sharpening, i.e., pansharpening, is to fuse a low spatial resolution multispectral (LRMS) image with a high spatial resolution panchromatic (PAN) image, aiming to obtain a high spatial resolution multispectral (HRMS) image. Pansharpening models based on variational optimization consist of a spectral fidelity term, a spatial fidelity term, and a regularization term. Most of the methods assume that the existing PAN image and the homologous HRMS image satisfy the global or local linear relationship, which could be far from the real case, thus causing sub-optimal performance. Inspired by the nonlinear mapping ability of machine learning (ML) techniques, we propose a novel spatial fidelity term with learnable nonlinear mapping (LNM-SF), which trains an implicit functional operator via a specifically designed convolutional neural network (CNN) and efficiently constructs the nonlinear relationship between the known PAN and the latent HRMS images. Relying upon the above description of the spatial fidelity term, a new variational model with a learnable nonlinear mapping in the spatial fidelity term for pansharpening, named LNM-PS, is simply integrated by the conventional spectral fidelity term into the proposed LNM-SF. To effectively solve the resulting optimization problem, we develop an alternating direction method of multipliers (ADMM)-based algorithm with the fast iterative shrinkage-thresholding algorithm (FISTA) as inner solver. Extensive numerical experiments on different datasets, assessing the performance both at reduced-resolution and full-resolution, show the superiority of the proposed LNM-PS method. The code is available at https://github.com/liangjiandeng/-LNM-PS.
Liang-Jian Deng, Zhong-Cheng Wu, Gemine Vivone
IEEE Trans. Geosci. Remote. Sens.2
2023 Variational Pansharpening Based on Coefficient Estimation With Nonlocal Regression
abstract
Pansharpening (which stands for panchromatic sharpening) involves the fusion between a multispectral (MS) image with a higher spectral content than a fine spatial resolution panchromatic (PAN) image to generate a high spatial resolution multispectral (HRMS) image. A widely-used concept is the construction of the relationship between PAN and HRMS images by designing pixel-based coefficients. Previous pixel-based methods compute the coefficients pixel-by-pixel while suffering from inaccuracies in some areas leading to spatial distortion. However, we found that the coefficients inherit the spatial properties of the HRMS image, e.g., the local smoothness and nonlocal self-similarity, and the spatial correlation between the coefficients and the HRMS image can increase the accuracy of the estimation process. In this article, we propose a novel spatial fidelity with nonlocal regression (SFNLR) to describe the relationship between PAN and HRMS images. Unlike from the pixel-based perspective, the SFNLR can jointly utilize the local smoothness and nonlocal self-similarity of the coefficients for preserving spatial information. Besides, the SFNLR is integrated with a widely-used spectral fidelity to formulate a new variational model for the pansharpening problem. An effective algorithm based on the alternating direction method of multiplier (ADMM) framework is designed to solve the proposed model. Qualitative and quantitative assessments on reduced and full resolution datasets from different satellites demonstrate that the proposed approach outperforms several state-of-the-art methods. The code is available at: https://github.com/Jin-liangXiao/SFNLR.
Jin-Liang Xiao, Ting-Zhu Huang, Liang-Jian Deng, Zhong-Cheng Wu, Gemine Vivone
IEEE Trans. Geosci. Remote. Sens.3
2023 QIS-GAN: A Lightweight Adversarial Network With Quadtree Implicit Sampling for Multispectral and Hyperspectral Image Fusion
abstract
Multispectral and Hyperspectral Image Fusion (MHIF) involves the fusion of high spatial resolution multispectral images (HR-MSI) and low spatial resolution hyperspectral images (LR-HSI) to generate high spatial resolution hyperspectral images (HR-HSI), has gained significant attention in the field of remote sensing imaging. While CNN and Transformer models have shown effectiveness in MHIF, existing CNN or Transformer-based algorithms are overburdened with model size, making it difficult to achieve an effective trade-off between fusion accuracy and degree of lightweight. Recently, Implicit Neural Representation (INR) has been proven good interpretability and the ability to exploit coordinate information in 2D tasks. Nonetheless, INR-based fusion networks have certain limitations, such as the need for deeper super-resolution networks as shallow encoders, and insufficient representation capability on high upsampling ratios. To address these challenges, we present the Quadtree Implicit Sampling (QIS), which employs a hierarchical sampling from the perspective of the quadtree, to enhance the capacity of the overall network. Furthermore, the remarkable design of QIS allows us to adopt a lightweight structure as the shallow encoder, greatly alleviating the network burden and achieving lightweight. Inspired by generative adversarial models, we incorporate QIS as a lightweight generator into the GAN framework named QIS-GAN and leverage a discriminator to increase the fidelity of fused images. The results showcase the superior performance of QIS-GAN on the MHIF tasks with upsampling ratios of ×4, ×8, and ×16, surpassing the state-of-the-art in several datasets. The code for our approach will be available at https://github.com/chunyuzhu/QIS-GAN.
Chunyu Zhu, Shangqi Deng, Yingjie Zhou 0001, Liang-Jian Deng
IEEE Trans. Geosci. Remote. Sens.4
2023 LRTCFPan: Low-Rank Tensor Completion Based Framework for Pansharpening
abstract
Pansharpening refers to the fusion of a low spatial-resolution multispectral image with a high spatial-resolution panchromatic image. In this paper, we propose a novel low-rank tensor completion (LRTC)-based framework with some regularizers for multispectral image pansharpening, called LRTCFPan. The tensor completion technique is commonly used for image recovery, but it cannot directly perform the pansharpening or, more generally, the super-resolution problem because of the formulation gap. Different from previous variational methods, we first formulate a pioneering image super-resolution (ISR) degradation model, which equivalently removes the downsampling operator and transforms the tensor completion framework. Under such a framework, the original pansharpening problem is realized by the LRTC-based technique with some deblurring regularizers. From the perspective of regularizer, we further explore a local-similarity-based dynamic detail mapping (DDM) term to more accurately capture the spatial content of the panchromatic image. Moreover, the low-tubal-rank property of multispectral images is investigated, and the low-tubal-rank prior is introduced for better completion and global characterization. To solve the proposed LRTCFPan model, we develop an alternating direction method of multipliers (ADMM)-based algorithm. Comprehensive experiments at reduced-resolution (i.e., simulated) and full-resolution (i.e., real) data exhibit that the LRTCFPan method significantly outperforms other state-of-the-art pansharpening methods. The code is publicly available at: https://github.com/zhongchengwu/code_LRTCFPan.
Zhong-Cheng Wu, Ting-Zhu Huang, Liang-Jian Deng, Jie Huang 0005, Jocelyn Chanussot, Gemine Vivone
IEEE Trans. Image Process.3
2023 A Triple-Double Convolutional Neural Network for Panchromatic Sharpening
abstract
Pansharpening refers to the fusion of a panchromatic (PAN) image with a high spatial resolution and a multispectral (MS) image with a low spatial resolution, aiming to obtain a high spatial resolution MS (HRMS) image. In this article, we propose a novel deep neural network architecture with level-domain-based loss function for pansharpening by taking into account the following double-type structures, i.e., double-level, double-branch, and double-direction, called as triple-double network (TDNet). By using the structure of TDNet, the spatial details of the PAN image can be fully exploited and utilized to progressively inject into the low spatial resolution MS (LRMS) image, thus yielding the high spatial resolution output. The specific network design is motivated by the physical formula of the traditional multi-resolution analysis (MRA) methods. Hence, an effective MRA fusion module is also integrated into the TDNet. Besides, we adopt a few ResNet blocks and some multi-scale convolution kernels to deepen and widen the network to effectively enhance the feature extraction and the robustness of the proposed TDNet. Extensive experiments on reduced- and full-resolution datasets acquired by WorldView-3, QuickBird, and GaoFen-2 sensors demonstrate the superiority of the proposed TDNet compared with some recent state-of-the-art pansharpening approaches. An ablation study has also corroborated the effectiveness of the proposed approach. The code is available at https://github.com/liangjiandeng/TDNet.
Tianjing Zhang, Liang-Jian Deng, Ting-Zhu Huang, Jocelyn Chanussot, Gemine Vivone
IEEE Trans. Neural Networks Learn. Syst.2
2022 LAGConv: Local-Context Adaptive Convolution Kernels with Global Harmonic Bias for Pansharpening
abstract
Pansharpening is a critical yet challenging low-level vision task that aims to obtain a higher-resolution image by fusing a multispectral (MS) image and a panchromatic (PAN) image. While most pansharpening methods are based on convolutional neural network (CNN) architectures with standard convolution operations, few attempts have been made with context-adaptive/dynamic convolution, which delivers impressive results on high-level vision tasks. In this paper, we propose a novel strategy to generate local-context adaptive (LCA) convolution kernels and introduce a new global harmonic (GH) bias mechanism, exploiting image local specificity as well as integrating global information, dubbed LAGConv. The proposed LAGConv can replace the standard convolution that is context-agnostic to fully perceive the particularity of each pixel for the task of remote sensing pansharpening. Furthermore, by applying the LAGConv, we provide an image fusion network architecture, which is more effective than conventional CNN-based pansharpening approaches. The superiority of the proposed method is demonstrated by extensive experiments implemented on a wide range of datasets compared with state-of-the-art pansharpening methods. Besides, more discussions testify that the proposed LAGConv outperforms recent adaptive convolution techniques for pansharpening.
Zi-Rong Jin, Tianjing Zhang, Tai-Xiang Jiang, Gemine Vivone, Liang-Jian Deng
AAAI5
2022 Construction Method of Overlay Network for Cyber-Physical System
abstract
The real-time requirements of network communication in cyber-physical systems are high, and the network topology is complex, which makes the communication efficiency in the system low in real-time. However, the traditional routing algorithm has been unable to adapt to the increasingly complex intelligent communication network in today's society. In order to improve the network communication efficiency of the cyber-physical system, a cyber-physical fusion system overlay network construction method based on an improved variable neighborhood search algorithm is proposed. Reduce the communication time overhead of cyber-physical systems. In the network node selection method of the cyber physical system, the neighborhood structure and the dithering method of the variable neighborhood search algorithm are optimized. The objective function is used to further optimize the selection strategy of the covering node set. Finally, the network time cost, communication delay and algorithm stability of the proposed method are analyzed on a large number of random discrete node sets. The experimental results show that the optimized variable domain search algorithm can effectively reduce the network delay and time overhead, and improve the communication efficiency of the cyber-physical system.
Liang-Jian Deng, Yubing Chen
APNOMS1
2022 Context-Aware Element Filter for Hyperspectral Image Super-Resolution
abstract
Hyperspectral image super-resolution (HISR) aims to fuse a low-resolution image (LR-HSI) and a high-resolution multispectral image (HR-MSI), generating a high-resolution hyperspectral image (HR-HSI). Previous attempts to apply convolutional neural networks (CNNs) with spatial-variant adaptive filters for HISR tasks. Such filters overcome the spatial invariance and content-agnostic property of standard convolution. However, the current adaptive filters only consider pixellevel specificity, ignoring that each element of the features has unique close relationships with their neighbourhoods. To address the issue, we propose a context-aware element filter (CEF) operation, which generates adaptive filters for each element with sufficient perception of the specificity of each element to improve the representation capability. CEF can generate a single-channel filter to trade off the computational resource consumption for each element and is appropriate for HISR tasks with element-level dependencies. Specifically, we design a new network structure for HISR, which utilizes CEF to replace the standard convolution in the residual block. Extensive experiments demonstrate the superiority of the proposed CEF both visually and quantitatively compared with state-of-the-art (SOTA) methods.
Ran Ran 0001, Liang-Jian Deng, Chen-Yu Zhao
IGARSS2
2022 Cross-Frequency Detail Compensation Network for Pansharpening
abstract
Pansharpening is a fusion technique aiming at improving the spatial resolution of multispectral images while preserving spectral information. Previous attempts to adopt CNNs have led to significant progress in pansharpening, but always with a cumbersome network structure, and there exists redundancy in both spatial and channel of feature maps learned by CNNs. Considering the distinct properties of components with different frequencies in the feature map, we propose a cross-frequency detail compensation network (CFDCNet) by processing low, medium, and high frequency separately. Specifically, a cross-frequency convolution block is designed to produce a representation that captures the different frequency classes while achieving the more efficient detail extraction. Overall pipeline is progressive, and the learned features are fused in an interactive compensation manner to obtain the final output. Experimental results demonstrate the superiority of CFDCNet over state-of-the-art pansharpening methods in terms of visual quality and quantitative metrics.
Xiao-Nan Zhao, Chen-Yu Zhao, Tianjing Zhang, Liang-Jian Deng
IGARSS4
2022 SpanConv: A New Convolution via Spanning Kernel Space for Lightweight Pansharpening
abstract
Standard convolution operations can effectively perform feature extraction and representation but result in high computational cost, largely due to the generation of the original convolution kernel corresponding to the channel dimension of the feature map, which will cause unnecessary redundancy. In this paper, we focus on kernel generation and present an interpretable span strategy, named SpanConv, for the effective construction of kernel space. Specifically, we first learn two navigated kernels with single channel as bases, then extend the two kernels by learnable coefficients, and finally span the two sets of kernels by their linear combination to construct the so-called SpanKernel. The proposed SpanConv is realized by replacing plain convolution kernel by SpanKernel. To verify the effectiveness of SpanConv, we design a simple network with SpanConv. Experiments demonstrate the proposed network significantly reduces parameters comparing with benchmark networks for remote sensing pansharpening, while achieving competitive performance and excellent generalization. Code is available at https://github.com/zhi-xuan-chen/IJCAI-2022 SpanConv.
Zhi-Xuan Chen, Cheng Jin 0003, Tianjing Zhang, Liang-Jian Deng
IJCAI5
2022 Source-Adaptive Discriminative Kernels based Network for Remote Sensing Pansharpening
abstract
For the pansharpening problem, previous convolutional neural networks (CNNs) mainly concatenate high-resolution panchromatic (PAN) images and low-resolution multispectral (LR-MS) images in their architectures, which ignores the distinctive attributes of different sources. In this paper, we propose a convolution network with source-adaptive discriminative kernels, called ADKNet, for the pansharpening task. Those kernels consist of spatial kernels generated from PAN images containing rich spatial details and spectral kernels generated from LR-MS images containing abundant spectral information. The kernel generating process is specially designed to extract information discriminately and effectively. Furthermore, the kernels are learned in a pixel-by-pixel manner to characterize different information in distinct areas. Extensive experimental results indicate that ADKNet outperforms current state-of-the-art (SOTA) pansharpening methods in both quantitative and qualitative assessments, in the meanwhile only with about 60,000 network parameters. Also, the proposed network is extended to the hyperspectral image super-resolution (HSISR) problem, still yields SOTA performance, proving the universality of our model. The code is available at http://github.com/liangjiandeng/ADKNet.
Siran Peng, Liang-Jian Deng, Jin-Fan Hu, Yu-Wei Zhuo
IJCAI2
2022 A Decoder-free Transformer-like Architecture for High-efficiency Single Image Deraining
abstract
Despite the success of vision Transformers for the image deraining task, they are limited by computation-heavy and slow runtime. In this work, we investigate Transformer decoder is not necessary and has huge computational costs. Therefore, we revisit the standard vision Transformer as well as its successful variants and propose a novel Decoder-Free Transformer-Like (DFTL) architecture for fast and accurate single image deraining. Specifically, we adopt a cheap linear projection to represent visual information with lower computational costs than previous linear projections. Then we replace standard Transformer decoder block with designed Progressive Patch Merging (PPM), which attains comparable performance and efficiency. DFTL could significantly alleviate the computation and GPU memory requirements through proposed modules. Extensive experiments demonstrate the superiority of DFTL compared with competitive Transformer architectures, e.g., ViT, DETR, IPT, Uformer, and Restormer. The code is available at https://github.com/XiaoXiao-Woo/derain.
Ting-Zhu Huang, Liang-Jian Deng, Tianjing Zhang
IJCAI3
2022 Tensor Wheel Decomposition and Its Tensor Completion Application
abstract
Recently, tensor network (TN) decompositions have gained prominence in computer vision and contributed promising results to high-order data recovery tasks. However, current TN models are rather being developed towards more intricate structures to pursue incremental improvements, which instead leads to a dramatic increase in rank numbers, thus encountering laborious hyper-parameter selection, especially for higher-order cases. In this paper, we propose a novel TN decomposition, dubbed tensor wheel (TW) decomposition, in which a high-order tensor is represented by a set of latent factors mapped into a specific wheel topology. Such decomposition is constructed starting from analyzing the graph structure, aiming to more accurately characterize the complex interactions inside objectives while maintaining a lower hyper-parameter scale, theoretically alleviating the above deficiencies. Furthermore, to investigate the potentiality of TW decomposition, we provide its one numerical application, i.e., tensor completion (TC), yet develop an efficient proximal alternating minimization-based solving algorithm with guaranteed convergence. Experimental results elaborate that the proposed method is significantly superior to other tensor decomposition-based state-of-the-art methods on synthetic and real-world data, implying the merits of TW decomposition. The code is available at: https://github.com/zhongchengwu/code_TWDec.
Zhong-Cheng Wu, Ting-Zhu Huang, Liang-Jian Deng, Hong-Xia Dou, Deyu Meng
NeurIPS3
2022 Exemplar-based image inpainting using adaptive two-stage structure-tensor based priority function and nonlocal filtering
Ting-Zhu Huang, Liang-Jian Deng, Xi-Le Zhao, Jin-Fan Hu
J. Vis. Commun. Image Represent.3
2022 Fusformer: A Transformer-Based Fusion Network for Hyperspectral Image Super-Resolution
abstract
Hyperspectral image super-resolution (HISR) is to fuse a low-resolution hyperspectral image (LR-HSI) and a high-resolution multispectral image (HR-MSI), aiming to obtain a high-resolution hyperspectral image (HR-HSI). Recently, various convolution neural network (CNN) based techniques have been successfully applied to address the HISR problem. However, these methods generally only consider the relation of a local neighborhood by convolution kernels with a limited receptive field, thus ignoring the global relationship in a feature map. In this paper, we design a transformer-based architecture (called Fusformer) for the HISR problem, which is the first attempt to apply the transformer architecture to this task to the best of our knowledge. Thanks to the excellent ability of feature representations, especially by the self-attention in the transformer, our approach can globally explore the intrinsic relationship within features. Considering the specific HISR problem, since the LR-HSI holds the primary spectral information, our method estimates the spatial residual between the upsampled LR-MSI and the desired HR-HSI, reducing the burden of training the whole data in a smaller mapping space. Various experiments show that our approach outperforms current state-of-the-art HISR methods. The code is available at https://github.com/J-FHu/Fusformer.
Jin-Fan Hu, Ting-Zhu Huang, Liang-Jian Deng, Hong-Xia Dou, Danfeng Hong, Gemine Vivone
IEEE Geosci. Remote. Sens. Lett.3
2022 VO+Net: An Adaptive Approach Using Variational Optimization and Deep Learning for Panchromatic Sharpening
abstract
Pansharpening refers to a spatio-spectral fusion of a lower spatial resolution multispectral (MS) image with a high spatial resolution panchromatic image, aiming at obtaining an image with a corresponding high resolution both in the domains. In this article, we propose a generic fusion framework that is able to weightedly combine variational optimization (VO) with deep learning (DL) for the task of pansharpening, where these crucial weights directly determining the relative contribution of DL to each pixel are estimated adaptively. This framework can benefit from both VO and DL approaches, e.g., the good modeling explanation and data generalization of a VO approach with the high accuracy of a DL technique thanks to massive data training. The proposed method can be divided into three parts: 1) for the VO modeling, a general details injection term inspired by the classical multiresolution analysis is proposed as a spatial fidelity term and a spectral fidelity employing the MS sensor’s modulation transfer functions is also incorporated; 2) for the DL injection, a weighted regularization term is designed to introduce deep learning into the variational model; and 3) the final convex optimization problem is efficiently solved by the designed alternating direction method of multipliers. Extensive experiments both at reduced and full-resolution demonstrate that the proposed method outperforms recent state-of-the-art pansharpening methods, especially showing a higher accuracy and a significant generalization ability.
Zhong-Cheng Wu, Ting-Zhu Huang, Liang-Jian Deng, Jin-Fan Hu, Gemine Vivone
IEEE Trans. Geosci. Remote. Sens.3
2022 A New Context-Aware Details Injection Fidelity With Adaptive Coefficients Estimation for Variational Pansharpening
abstract
Pansharpening is related to the fusion of a low spatial resolution multispectral (MS) image retaining an abundant spectral content and a high spatial resolution panchromatic (PAN) image to obtain a product with both the abundant spectral content of the former and the high spatial resolution of the latter. Many previous studies are only focused on the global or local relationship between the PAN image and the corresponding high-resolution multispectral (HRMS) image. However, we found that the relationship between PAN and HRMS images in the gradient domain can be better explored through the image context. In this article, we propose context-aware details injection fidelity (CDIF) with adaptive coefficients estimation, which can fully explore the complicated relationship between the PAN image and the HRMS image in the gradient domain. More specifically, we apply a clustering method to divide the pixels of an image into different context-based regions. Afterward, the adaptive coefficients are estimated by using a regression-based method for each region. The CDIF is effective in extracting the main features from the two inputs to be fused. In addition, we integrate the CDIF with a conventional fidelity term and a total variation regularization to formulate a novel variational pansharpening model that is solved by designing an algorithm based on the alternating direction method of multiplier (ADMM) framework. Qualitative and quantitative assessments on different datasets support the effectiveness and robustness of the proposed method. The code is available athttps://github.com/liangjiandeng/CDIF.
Jin-Liang Xiao, Ting-Zhu Huang, Liang-Jian Deng, Zhong-Cheng Wu, Gemine Vivone
IEEE Trans. Geosci. Remote. Sens.3
2022 An Iterative Regularization Method Based on Tensor Subspace Representation for Hyperspectral Image Super-Resolution
abstract
Hyperspectral image super-resolution (HSI-SR) can be achieved by fusing a paired multispectral image (MSI) and hyperspectral image (HSI), which is a prevalent strategy. But, how to precisely reconstruct the high spatial resolution hyperspectral image (HR-HSI) by fusion technology is a challenging issue. In this paper, we propose an iterative regularization method based on tensor subspace representation (IR-TenSR) for MSI-HSI fusion, thus HSI-SR. First, we propose a tensor subspace representation (TenSR)-based regularization model that integrates the global spectral-spatial low-rank and the nonlocal self-similarity priors of HR-HSI. These two priors have been proven effective, but previous HSI-SR works cannot simultaneously exploit them. Subsequently, we design an iterative regularization procedure to utilize the residual information of acquired low-resolution images, which are ignored in other works that produce suboptimal results. Finally, we develop an effective algorithm based on the proximal alternating minimization method to solve the TenSR-regularization model. With that, we obtain the iterative regularization algorithm. Experiments implemented on the simulated and real datasets illustrate the advantages of the proposed IR-TenSR compared with state-of-the-art fusion approaches. The code is available at https://github.com/liangjiandeng/IR-TenSR.
Ting-Zhu Huang, Liang-Jian Deng, Naoto Yokoya
IEEE Trans. Geosci. Remote. Sens.3
2022 Hyperspectral Image Super-Resolution via Deep Spatiospectral Attention Convolutional Neural Networks
abstract
Hyperspectral images (HSIs) are of crucial importance in order to better understand features from a large number of spectral channels. Restricted by its inner imaging mechanism, the spatial resolution is often limited for HSIs. To alleviate this issue, in this work, we propose a simple and efficient architecture of deep convolutional neural networks to fuse a low-resolution HSI (LR-HSI) and a high-resolution multispectral image (HR-MSI), yielding a high-resolution HSI (HR-HSI). The network is designed to preserve both spatial and spectral information thanks to a new architecture based on: 1) the use of the LR-HSI at the HR-MSI's scale to get an output with satisfied spectral preservation and 2) the application of the attention and pixelShuffle modules to extract information, aiming to output high-quality spatial details. Finally, a plain mean squared error loss function is used to measure the performance during the training. Extensive experiments demonstrate that the proposed network architecture achieves the best performance (both qualitatively and quantitatively) compared with recent state-of-the-art HSI super-resolution approaches. Moreover, other significant advantages can be pointed out by the use of the proposed approach, such as a better network generalization ability, a limited computational burden, and the robustness with respect to the number of training samples. Please find the source code and pretrained models from https://liangjiandeng.github.io/Projects_Res/HSRnet_2021tnnls.html.
Jin-Fan Hu, Ting-Zhu Huang, Liang-Jian Deng, Tai-Xiang Jiang, Gemine Vivone, Jocelyn Chanussot
IEEE Trans. Neural Networks Learn. Syst.3
2021 Dynamic Cross Feature Fusion for Remote Sensing Pansharpening
abstract
Deep Convolution Neural Networks have been adopted for pansharpening and achieved state-of-the-art performance. However, most of the existing works mainly focus on single-scale feature fusion, which leads to failure in fully considering relationships of information between high-level semantics and low-level features, despite the network is deep enough. In this paper, we propose a dynamic cross feature fusion network (DCFNet) for pansharpening. Specifically, DCFNet contains multiple parallel branches, including a high-resolution branch served as the backbone, and the low-resolution branches progressively supplemented into the backbone. Thus our DCFNet can represent the overall information well. In order to enhance the relationships of inter-branches, dynamic cross feature transfers are embedded into multiple branches to obtain high-resolution representations. Then contextualized features will be learned to improve the fusion of information. Experimental results indicate that DCFNet significantly outperforms the prior arts in both quantitative indicators and visual qualities.
Ting-Zhu Huang, Liang-Jian Deng, Tianjing Zhang
ICCV3
2021 Weighted Shallow-Deep Feature Fusion Network for Pansharpening
abstract
In this paper, we propose a novel weighted shallow-deep feature fusion convolutional neural network (WSDFNet) for the task of multispectral image pansharpening. This network could effectively overcome the drawback of the common identity skip connection (ISC), and propagate shallow features scaled by a novel adaptive skip weighter (ASW) to deeper layers. By the technique, it could favor the feature fusion in different network depths adequately, as well as yield a promising outcome. Experimental results on reduced- and full-resolution WorldView-3 dataset demonstrate the superiority of the WSDFNet compared with recent state-of-the-art (SOTA) pansharpening approaches. Moreover, WSDFNet is also verified as a lightweight network.
Zi-Rong Jin, Tianjing Zhang, Cheng Jin 0003, Liang-Jian Deng
IGARSS4
2021 Progressive Band-Separated Convolutional Neural Network for Multispectral Pansharpening
abstract
Recently, convolutional neural networks (CNNs) have been introduced to pansharpening for enhancing fusion accuracy and overcoming the drawbacks of the conventional methods. However, most of methods based on CNN fail to distinguish the difference of multispectral bands, and only use a uniform set of convolutional kernels to extract features. In this paper, we design a progressive, band-separated convolutional network architecture for discriminatively learning the features and relation among spectral bands, aiming to address the problem mentioned before. More specifically, the proposed architecture mainly consists of three aspects. First, to accurately preserve the spectral peculiarities, we divide the multispectral input image in terms of its bands into several groups. Second, our original panchromatic and multispectral inputs are filtered by a high-pass operation to further yield more spatial details. Third, we use a spectral fusion module (SFM) for each group and associate them to progressively assemble the whole architecture. It is worth mentioning that the architecture could be integrated into any other competitive CNNs to improve the performance. Both visual and quantitative experiments have demonstrated that our proposed method outperforms recent state-of-the-art pansharpening techniques.
Shishi Xiao, Cheng Jin 0003, Tianjing Zhang, Ran Ran 0001, Liang-Jian Deng
IGARSS5
2021 A Variational Approach with Nonlocal Self-Similarity and Joint-Sparsity for Hyperspectral Image Super-Resolution
abstract
The aim of hyperspectral image super-resolution (HSI-SR) is to produce high spatial resolution hyperspectral image (HR - HSI) by exploiting the available high spatial resolution multispectral image (HR-MSI) and low spatial resolution hyperspectral image (LR-HSI). In this work, we develop a novel matrix factorization (MF)-based HSI -SR way, which formulates the HSI -SR problem as estimating the spectral dictionary from the observed LR - HSI and the coefficient matrix from both the observed HR-MSI and LR-HSI. Specifically, we first estimate the spectral dictionary from the observed LR - HSI by the dictionary learning algorithm with redundancy assumption. Moreover, based on the superpixel segmentation technology used in the observed HR-MSI, the coefficient vectors are grouped. By concatenating the joint-sparse, nonlocallow-rank, and nonnegative priors of the grouped coefficient vectors, we develop a novel coefficient matrix estimation variational model, which fully explores the nonlocal self-similarity of the desired HR-HSI. The proposed coefficient matrix estimation model is solved under the alternating direction method of multipliers (ADMM) framework. Experimental results prove the superiority of the proposed way from the quantitative and qualitative analysis.
Ting-Zhu Huang, Yong Chen 0013, Jie Huang 0005, Liang-Jian Deng
IGARSS5
2021 BAM: Bilateral Activation Mechanism for Image Fusion
abstract
As the conventional activation functions such as ReLU, LeakyReLU, and PReLU, the negative parts in feature maps are simply truncated or linearized, which may result in unflexible structure and undesired information distortion. In this paper, we propose a simple but effective Bilateral Activation Mechanism (BAM) which could be applied to the activation function to offer an efficient feature extraction model. Based on BAM, the Bilateral ReLU Residual Block (BRRB) that still sufficiently keeps the nonlinear characteristic of ReLU is constructed to separate the feature maps into two parts, i.e., the positive and negative components, then adaptively represent and extract the features by two independent convolution layers. Besides, our mechanism will not increase any extra parameters or computational burden in the network. We finally embed the BRRB into a basic ResNet architecture, called BRResNet, it is easy to obtain state-of-the-art performance in two image fusion tasks, i.e., pansharpening and hyperspectral image super-resolution (HISR). Additionally, deeper analysis and ablation study demonstrate the effectiveness of BAM, the lightweight property of the network, etc. Please find the code from the project page1 https://liangjiandeng.github.io/Projects_Res/bam_mm2021.html
Zi-Rong Jin, Liang-Jian Deng, Tianjing Zhang, Xiao-Xu Jin
ACM Multimedia2
2021 SSconv: Explicit Spectral-to-Spatial Convolution for Pansharpening
abstract
Pansharpening aims to fuse a high spatial resolution panchromatic (PAN) image and a low resolution multispectral (LR-MS) image to obtain a multispectral image with the same spatial resolution as the PAN image. Thanks to the flexible structure of convolution neural networks (CNNs), they have been successfully applied to the problem of pansharpening. However, most of the existing methods only simply feed the up-sampled LR-MS into the CNNs and ignore the spatial distortion caused by direct up-sampling. In this paper, we propose an explicit spectral-to-spatial convolution (SSconv) that aggregates spectral features into the spatial domain to perform the up-sampling operation, which can get better performance than the direct up-sampling. Furthermore, SSconv is embedded into a multiscale U-shaped convolution neural network (MUCNN) for fully utilizing the multispectral information of involved images. In particular, multiscale injection branch and mixed loss on cross-scale levels are employed to fuse pixel-wise image information. Benefiting from the distortion-free property of SSconv, the proposed MUCNN can generate state-of-the-art performance with a simple structure, both on reduced-resolution and full-resolution datasets acquired from WorldView-3 and GaoFen-2. Please find the code from the project page.
Liang-Jian Deng, Tianjing Zhang
ACM Multimedia2
2021 EAA-Net: A novel edge assisted attention network for single image dehazing
Chao Wang 0008, Haozhen Shen, Ming-Wen Shao, Chuan-Sheng Yang, Jiancheng Luo, Liang-Jian Deng
Knowl. Based Syst.7
2021 Endmember independence constrained hyperspectral unmixing via nonnegative tensor factorization
Jin-Ju Wang, Ding-Cheng Wang, Ting-Zhu Huang, Jie Huang 0005, Xi-Le Zhao, Liang-Jian Deng
Knowl. Based Syst.6
2021 Detail Injection-Based Deep Convolutional Neural Networks for Pansharpening
abstract
The fusion of high spatial resolution panchromatic (PAN) data with simultaneously acquired multispectral (MS) data with the lower spatial resolution is a hot topic, which is often called pansharpening. In this article, we exploit the combination of machine learning techniques and fusion schemes introduced to address the pansharpening problem. In particular, deep convolutional neural networks (DCNNs) are proposed to solve this issue. The latter is combined first with the traditional component substitution and multiresolution analysis fusion schemes in order to estimate the nonlinear injection models that rule the combination of the upsampled low-resolution MS image with the extracted details exploiting the two philosophies. Furthermore, inspired by these two approaches, we also developed another DCNN for pansharpening. This is fed by the direct difference between the PAN image and the upsampled low-resolution MS image. Extensive experiments conducted both at reduced and full resolutions demonstrate that this latter convolutional neural network outperforms both the other detail injection-based proposals and several state-of-the-art pansharpening methods.
Liang-Jian Deng, Gemine Vivone, Cheng Jin 0003, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.1
2021 Nonlocal Tensor-Based Sparse Hyperspectral Unmixing
abstract
Sparse unmixing is an important technique for analyzing and processing hyperspectral images (HSIs). Simultaneously exploiting spatial correlation and sparsity improves substantially abundance estimation accuracy. In this article, we propose to exploit nonlocal spatial information in the HSI for the sparse unmixing problem. Specifically, we first group similar patches in the HSI, and then unmix each group by imposing simultaneous a low-rank constraint and joint sparsity in the corresponding third-order abundance tensor. To this end, we build an unmixing model with a mixed regularization term consisting of the sum of the weighted tensor trace norm and the weighted tensor$\ell _{2,1}$-norm of the abundance tensor. The proposed model is solved under the alternating direction method of multipliers framework. We term the developed algorithm as the nonlocal tensor-based sparse unmixing algorithm. The effectiveness of the proposed algorithm is illustrated in experiments with both simulated and real hyperspectral data sets.
Jie Huang 0005, Ting-Zhu Huang, Xi-Le Zhao, Liang-Jian Deng
IEEE Trans. Geosci. Remote. Sens.4
2021 Rain Streaks Removal for Single Image via Kernel-Guided Convolutional Neural Network
abstract
Recently emerged deep learning methods have achieved great success in single image rain streaks removal. However, existing methods ignore an essential factor in the rain streaks generation mechanism, i.e., the motion blur leading to the line pattern appearances. Thus, they generally produce overderaining or underderaining results. In this article, inspired by the generation mechanism, we propose a novel rain streaks removal framework using a kernel-guided convolutional neural network (KGCNN), achieving state-of-the-art performance with a simple network architecture. More precisely, our framework consists of three steps. First, we learn the motion blur kernel by a plain neural network, termed parameter network, from the detail layer of a rainy patch. Then, we stretch the learned motion blur kernel into a degradation map with the same spatial size as the rainy patch. Finally, we use the stretched degradation map together with the detail patches to train a deraining network with a typical ResNet architecture, which produces the rain streaks with the guidance of the learned motion blur kernel. Experiments conducted on extensive synthetic and real data demonstrate the effectiveness of the proposed KGCNN, in terms of rain streaks removal and image detail preservation.
Ye-Tao Wang, Xi-Le Zhao, Tai-Xiang Jiang, Liang-Jian Deng, Yi Chang 0002, Ting-Zhu Huang
IEEE Trans. Neural Networks Learn. Syst.4
2019 Rain Streaks Removal for Single Image Via Directional Total Variation Regularization
abstract
Images captured in rainy conditions are often corrupted by unexpected rain streaks, which severely degrade the performance of subsequent processes in outdoor computer vision systems. In this paper, we exploit the directional smoothness of rain streaks for the single-image rain streaks removal and propose a convex model that uses the directional total variation (DTV) to characterize the smoothness of rain streaks in arbitrary orientations. The proposed model consists of four terms: the fidelity term, the ℓ1norm for the sparsity of rain streaks, and two DTV regularization terms for the directional smoothness and the piecewise smoothness of rain streaks and rain-free backgrounds, respectively. To solve the proposed model, we develop an efficient algorithm based on the alternating direction method of multipliers (ADMM) framework. Extensive experimental results on both synthetic and real rainy images show that our method outperforms the recent state-of-the-art methods visually and quantitatively.
Yugang Wang, Ting-Zhu Huang, Xi-Le Zhao, Liang-Jian Deng, Tai-Xiang Jiang
ICIP4
2019 Unidirectional Sparse Tensor Based Model for the Noise Removal of Remote Sensing Image
abstract
In this paper, we mainly focus on a quite challenging denoising problem in remote sensing images, which is to simultaneously remove Gaussian noise and sparse noise that mainly include stripes and salt-pepper noise. We propose a convex unidirectional sparse model based on mode-3 tensor modeling to remove the mixture noise. A proximal alternating direction method of multipliers (ADMM) based algorithm is designed to effectively solve the given minimization model. Comparing with some recent state-of-the-art denoising methods, the proposed method shows the best performance from visual and quantitative aspects.
Hong-Xia Dou, Ting-Zhu Huang, Liang-Jian Deng
IGARSS3
2019 Pan-Sharpening Via RoG-Based Filtering
abstract
In this paper, a pan-sharpening approach based on RoG filtering is proposed. This approach follows the framework of classic methods of pan-sharpening, i.e., component substitution and multi-resolution analysis. The filtering technique based on Relativity-of-Gaussian (RoG) regularization is first used in the process of upsampling the original multi-spectral image, and then in the detail extraction phase to obtain spatial details from the panchromatic image. Experiments on datasets acquired by Quickbird and IKONOS demonstrate that the proposed approach obtains competitive performance comparing with several popular pan-sharpening methods.
Ting-Zhu Huang, Liang-Jian Deng, Jie Huang 0005, Hong-Xia Dou
IGARSS3
2019 Bilateral filter based total variation regularization for sparse hyperspectral image unmixing
Jie Huang 0005, Liang-Jian Deng, Ting-Zhu Huang
Inf. Sci.3
2019 A New Operator Splitting Method for the Euler Elastica Model for Image Smoothing
abstract
Euler's elastica model has a wide range of applications in image processing and computer vision. However, the nonconvexity, the nonsmoothness, and the nonlinearity of the associated energy functional make its minimization a challenging task, further complicated by the presence of high order derivatives in the model. In this article we propose a new operator-splitting algorithm to minimize the Euler elastica functional. This algorithm is obtained by applying an operator-splitting based time discretization scheme to an initial value problem (dynamical flow) associated with the optimality system (a system of multivalued equations). The subproblems associated with the three fractional steps of the splitting scheme have either closed form solutions or can be handled by fast dedicated solvers. Compared with earlier approaches relying on ADMM (Alternating Direction Method of Multipliers), the new method has, essentially, only the time discretization step as free parameter to choose, resulting in a very robust and stable algorithm. The simplicity of the subproblems and its modularity make this algorithm quite efficient. Applications to the numerical solution of smoothing test problems demonstrate the efficiency and robustness of the proposed methodology.
Liang-Jian Deng, Roland Glowinski, Xue-Cheng Tai
SIAM J. Imaging Sci.1
2019 A total variation and group sparsity based tensor optimization model for video rain streak removal
Ye-Tao Wang, Xi-Le Zhao, Tai-Xiang Jiang, Liang-Jian Deng, Tian-Hui Ma, Yue-Tian Zhang, Ting-Zhu Huang
Signal Process. Image Commun.4
2019 Joint-Sparse-Blocks and Low-Rank Representation for Hyperspectral Unmixing
abstract
Hyperspectral unmixing has attracted much attention in recent years. Single sparse unmixing assumes that a pixel in a hyperspectral image consists of a relatively small number of spectral signatures from large, ever-growing, and available spectral libraries. Joint-sparsity (or row-sparsity) model typically enforces all pixels in a neighborhood to share the same set of spectral signatures. The two sparse models are widely used in the literature. In this paper, we propose a joint-sparsity-blocks model for abundance estimation problem. Namely, the abundance matrix of size m × n is partitioned to have one row block and s column blocks and each column block itself is joint-sparse. It generalizes both the single (i.e., s = n) and the joint (i.e., s = 1) sparsities. Moreover, concatenating the proposed joint-sparsity-blocks structure and low rankness assumption on the abundance coefficients, we develop a new algorithm called joint-sparseblocks and low-rank unmixing. In particular, for the joint-sparseblocks regression problem, we develop a two-level reweighting strategy to enhance the sparsity along the rows within each block. Simulated and real-data experiments demonstrate the effectiveness of the proposed algorithm.
Jie Huang 0005, Ting-Zhu Huang, Liang-Jian Deng, Xi-Le Zhao
IEEE Trans. Geosci. Remote. Sens.3
2019 FastDeRain: A Novel Video Rain Streak Removal Method Using Directional Gradient Priors
abstract
Rain streaks removal is an important issue in outdoor vision systems and has recently been investigated extensively. In this paper, we propose a novel video rain streak removal approach FastDeRain, which fully considers the discriminative characteristics of rain streaks and the clean video in the gradient domain. Specifically, on the one hand, rain streaks are sparse and smooth along the direction of the raindrops, whereas on the other hand, clean videos exhibit piecewise smoothness along the rain-perpendicular direction and continuity along the temporal direction. Theses smoothness and continuity results in the sparse distribution in the different directional gradient domain, respectively. Thus, we minimize 1) the ℓ1 norm to enhance the sparsity of the underlying rain streaks, 2) two ℓ1 norm of unidirectional Total Variation (TV) regularizers to guarantee the anisotropic spatial smoothness, and 3) an ℓ1 norm of the time-directional difference operator to characterize the temporal continuity. A split augmented Lagrangian shrinkage algorithm (SALSA) based algorithm is designed to solve the proposed minimization model. Experiments conducted on synthetic and real data demonstrate the effectiveness and efficiency of the proposed method. According to comprehensive quantitative performance measures, our approach outperforms other state-of-the-art methods, especially on account of the running time. The code of FastDeRain can be downloaded at https://github.com/TaiXiangJiang/FastDeRain.
Tai-Xiang Jiang, Ting-Zhu Huang, Xi-Le Zhao, Liang-Jian Deng, Yao Wang 0003
IEEE Trans. Image Process.4
2018 Matrix factorization for low-rank tensor completion using framelet prior
Tai-Xiang Jiang, Ting-Zhu Huang, Xi-Le Zhao, Teng-Yu Ji, Liang-Jian Deng
Inf. Sci.5
2018 A Variational Pansharpening Approach Based on Reproducible Kernel Hilbert Space and Heaviside Function
abstract
Pansharpening is an important application in remote sensing image processing. It can increase the spatial-resolution of a multispectral image by fusing it with a high spatial-resolution panchromatic image in the same scene, which brings great favor for subsequent processing such as recognition, detection, etc. In this paper, we propose a continuous modeling and sparse optimization based method for the fusion of a panchromatic image and a multispectral image. The proposed model is mainly based on reproducing kernel Hilbert space (RKHS) and approximated Heaviside function (AHF). In addition, we also propose a Toeplitz sparse term for representing the correlation of adjacent bands. The model is convex and solved by the alternating direction method of multipliers which guarantees the convergence of the proposed method. Extensive experiments on many real datasets collected by different sensors demonstrate the effectiveness of the proposed technique as compared with several state-of-the-art pansharpening approaches.
Liang-Jian Deng, Gemine Vivone, Weihong Guo 0002, Mauro Dalla Mura, Jocelyn Chanussot
IEEE Trans. Image Process.1
2017 A Novel Tensor-Based Video Rain Streaks Removal Approach via Utilizing Discriminatively Intrinsic Priors
abstract
Rain streaks removal is an important issue of the outdoor vision system and has been recently investigated extensively. In this paper, we propose a novel tensor based video rain streaks removal approach by fully considering the discriminatively intrinsic characteristics of rain streaks and clean videos, which needs neither rain detection nor time-consuming dictionary learning stage. In specific, on the one hand, rain streaks are sparse and smooth along the raindrops direction, and on the other hand, the clean videos possess smoothness along the rain-perpendicular direction and global and local correlation along time direction. We use the l1 norm to enhance the sparsity of the underlying rain, two unidirectional Total Variation (TV) regularizers to guarantee the different discriminative smoothness, and a tensor nuclear norm and a time directional difference operator to characterize the exclusive correlation of the clean video along time. Alternation direction method of multipliers (ADMM) is employed to solve the proposed concise tensor based convex model. Experiments implemented on synthetic and real data substantiate the effectiveness and efficiency of the proposed method. Under comprehensive quantitative performance measures, our approach outperforms other state-of-the-art methods.
Tai-Xiang Jiang, Ting-Zhu Huang, Xi-Le Zhao, Liang-Jian Deng, Yao Wang 0003
CVPR4
2017 A variational pansharpening approach based on reproducible kernel Hilbert space and heaviside function
abstract
In this paper, we propose a continuous modeling and sparse optimization based method for the fusion of a panchromatic (PAN) image and a multispectral (MS) image. The proposed model is mainly based on reproducing kernel Hilbert space (RKHS) and approximated Heaviside function (AHF). In addition, we also design an iterative strategy to recover more image details. The final model is a convex one and solved by the designed alternating direction method of multipliers (ADMM) which guarantees the convergence of the proposed method. Experimental results on two real datasets corresponding to different sensors and different resolutions demonstrate the effectiveness of the proposed approach as compared with several state-of-the-art pansharpening approaches.
Liang-Jian Deng, Gemine Vivone, Weihong Guo 0002, Mauro Dalla Mura, Jocelyn Chanussot
ICIP1
2017 Stripe noise removal of remote sensing image with a directional l0 sparse model
abstract
This paper commits to remove the stripe noise to enhance the visual quality of remote sensing images, in the meanwhile preserves image details of stripe-free regions. Instead of solving the underlying image as most of researches, we propose a non-convex l0model for remote sensing image destriping by taking full consideration of the intrinsically directional and structural priors of stripe noise. Moreover, the proposed non-convex model can be solved by the proximal alternating direction method of multipliers (PADMM) method which theoretically guarantees converging to a KKT point. Extensively experimental results on simulated and real data demonstrate that the proposed method outperforms recent state-of-the-art destriping methods, both visually and quantitatively.
Hong-Xia Dou, Ting-Zhu Huang, Liang-Jian Deng, Yong Chen 0013
ICIP3
2017 Image fusion via dynamic gradient sparsity and anisotropic spectral-spatial total variation
abstract
In this paper, we develop a sparsity based model for the fusion of a high spatial-resolution image and a multispectral image. The given model is based on the combination of a dynamic gradient sparsity (DGS) and an anisotropic spectral-spatial total variation (ASSTV). We design an alternating direction method of multipliers (ADMM) based algorithm to solve the proposed model. In contrast to existing approaches, the proposed method can generate more spatial details as well as preserve favorable spectral information. Experimental results demonstrate that the proposed approach outperforms several state-of-the-art image fusion methods both quantitatively and visually, in terms of both pansharpening application of remote sensing images and fusion application of natural color images.
Chao-Chao Zheng, Ting-Zhu Huang, Liang-Jian Deng, Xi-Le Zhao, Hong-Xia Dou
ICIP3
2017 Group sparsity based regularization model for remote sensing image stripe noise removal
Yong Chen 0013, Ting-Zhu Huang, Liang-Jian Deng, Xi-Le Zhao, Min Wang 0022
Neurocomputing3
2016 Single image super-resolution by approximated Heaviside functions
Liang-Jian Deng, Weihong Guo 0002, Ting-Zhu Huang
Inf. Sci.1
2016 Single-Image Super-Resolution via an Iterative Reproducing Kernel Hilbert Space Method
abstract
Image super-resolution (SR), a process to enhance image resolution, has important applications in satellite imaging, high-definition television, medical imaging, and so on. Many existing approaches use multiple low-resolution (LR) images to recover one high-resolution (HR) image. In this paper, we present an iterative scheme to solve single-image SR problems. It recovers a high-quality HR image from solely one LR image without using a training data set. We solve the problem from image intensity function estimation perspective and assume that the image contains smooth and edge components. We model the smooth components of an image using a thin-plate reproducing kernel Hilbert space and the edges using approximated Heaviside functions. The proposed method is applied to image patches, aiming to reduce computation and storage. Visual and quantitative comparisons with some competitive approaches show the effectiveness of the proposed method.
Liang-Jian Deng, Weihong Guo 0002, Ting-Zhu Huang
IEEE Trans. Circuits Syst. Video Technol.1
2015 Heaviside image edge sharpening
abstract
In this paper, we propose an automatic and efficient method to enhance edge sharpness of images. Starting from an image with blur edges, we improve the edges using transformed Heaviside functions for better visualization. In addition, we provide an efficient method to directly compute the scaling and shifting factors of the transformed Heaviside functions, so that blur edges can be improved accurately. Experimental results show that the proposed method is fast and can get sharper image edges than some recent state-of-the-art edge enhancement methods. We also apply the edge sharpening method to image super-resolution and obtained promising results.
Liang-Jian Deng, Weihong Guo 0002, Ting-Zhu Huang, Xi-Le Zhao
MMSP1