Haijin Zeng

dblp:261/8056 · DBLP profile ↗
← Back
29ranked-venue papers
13as first author
27since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 first-author · 16 since 2021Artificial intelligence and machine learning · 15 · 8 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 7 since 2021
YearPublicationVenuePosition
2026 Self-Supervised One-Step Diffusion Refinement for Snapshot Compressive Imaging
abstract
Snapshot compressive imaging (SCI) captures multispectral images (MSIs) using a single coded two-dimensional (2-D) measurement, but reconstructing high-fidelity MSIs from these compressed inputs remains a fundamentally ill-posed challenge. Recent diffusion-based methods improve quality but are limited by scarce MSI training data, domain shifts from RGB-pretrained models, and slow multi-step sampling. These drawbacks restrict their practicality in real-world applications. Unlike prior approaches that rely on expensive iterative refinement or subspace-based diffusion embeddings (e.g., DiffSCI, PSR-SCI)—we introduce a fundamentally different paradigm: a self-supervised One-Step Diffusion (OSD) framework designed specifically for SCI. The key novelty lies in using a single-step diffusion refiner to correct an initial reconstruction, eliminating iterative denoising entirely while preserving generative quality. Moreover, we adopt a self-supervised equivariant learning strategy to train both the predictor and refiner directly from raw 2-D measurements, enabling generalization to unseen domains without ground-truth MSI. To further address limited MSI data, we design a band-selection–driven distillation strategy that transfers core generative priors from large-scale RGB datasets, effectively bridging the domain gap. Extensive experiments confirm that our approach sets a new standard—yielding PSNR gains of 3.44dB, 1.61dB, and 0.28dB on the Harvard, NTIRE, and ICVL datasets respectively, while cutting reconstruction time from 8.9s to just 0.22s per image. These gains in efficiency and adaptability advance SCI reconstruction, enabling accurate and practical real-world deployment.
Shaoguang Huang, Yunzhen Wang, Haijin Zeng, Hongyu Chen 0003, Hongyan Zhang 0001
AAAI3
2026 Deep LoRA-Unfolding Networks for Image Restoration
abstract
Deep unfolding networks (DUNs), combining conventional iterative optimization algorithms and deep neural networks into a multi-stage framework, have achieved remarkable accomplishments in Image Restoration (IR), such as spectral imaging reconstruction, compressive sensing and super-resolution. It unfolds the iterative optimization steps into a stack of sequentially linked blocks. Each block consists of a Gradient Descent Module (GDM) and a Proximal Mapping Module (PMM) which is equivalent to a denoiser from a Bayesian perspective, operating on Gaussian noise with a known level. However, existing DUNs suffer from two critical limitations: 1) their PMMs share identical architectures and denoising objectives across stages, ignoring the need for stage-specific adaptation to varying noise levels; and 2) their chain of structurally repetitive blocks results in severe parameter redundancy and high memory consumption, hindering deployment in large-scale or resource-constrained scenarios. To address these challenges, we introduce generalized Deep Low-rank Adaptation (LoRA) Unfolding Networks for image restoration, named LoRun, harmonizing denoising objectives and adapting different denoising levels between stages with compressed memory usage for more efficient DUN. LoRun introduces a novel paradigm where a single pretrained base denoiser is shared across all stages, while lightweight, stage-specific LoRA adapters are injected into the PMMs to dynamically modulate denoising behavior according to the noise level at each unfolding step. This design decouples the core restoration capability from task-specific adaptation, enabling precise control over denoising intensity without duplicating full network parameters and achieving up to $N$ times parameter reduction for an $N$ -stage DUN with on-par or better performance. Extensive experiments conducted on three IR tasks validate the efficiency of our method.
Xiangming Wang, Haijin Zeng, Benteng Sun, Jiezhang Cao, Kai Zhang 0008, Qiangqiang Shen, Yongyong Chen
IEEE Trans. Image Process.2
2025 OTLRM: Orthogonal Learning-based Low-Rank Metric for Multi-Dimensional Inverse Problems
abstract
In real-world scenarios, complex data such as multispectral images and multi-frame videos inherently exhibit robust low-rank property. This property is vital for multi-dimensional inverse problems, such as tensor completion, spectral imaging reconstruction, and multispectral image denoising. Existing tensor singular value decomposition (t-SVD) definitions rely on hand-designed or pre-given transforms, which lack flexibility for defining tensor nuclear norm (TNN). The TNN-regularized optimization problem is solved by the singular value thresholding (SVT) operator, which leverages the t-SVD framework to obtain the low-rank tensor. However, it's quite complicated to introduce SVT into deep neural network due to the numerical instability problem in solving the derivatives of the eigenvectors. In this paper, we introduce a novel data-driven generative low-rank t-SVD model based on the learnable orthogonal transform, which can be naturally solved under its representation. Prompted by the linear algebra theorem of the Householder transformation, our learnable orthogonal transform is achieved by constructing an endogenously orthogonal matrix adaptable to neural networks, optimizing it as arbitrary orthogonal matrices. Additionally, we propose a low-rank solver as a generalization of SVT, which utilizes an efficient representation of generative networks to obtain low-rank structures. Extensive experiments highlight its significant restoration enhancements.
Xiangming Wang, Haijin Zeng, Jiaoyang Chen, Sheng Liu 0033, Yongyong Chen, Guoqing Chao
AAAI2
2025 Gradient of White Matter Functional Variability via fALFF Differential Identifiability
abstract
Functional variability in both gray matter (GM) and white matter (WM) is closely associated with human brain cognitive and developmental processes, and is commonly assessed using functional connectivity (FC). However, as a correlationbased approach, FC captures the co-fluctuation between brain regions rather than the intensity of neural activity in each region. Consequently, FC provides only a partial view of functional variability, and this limitation is particularly pronounced in WM, where functional signals are weaker and more susceptible to noise. To tackle this limitation, we introduce fractional amplitude of low-frequency fluctuation (fALFF) to measure the intensity of spontaneous neural activity and analyze functional variability in WM. Specifically, we propose a novel method to quantify WM functional variability by estimating the differential identifiability of fALFF. Higher differential identifiability is observed in WM fALFF compared to FC, which indicates that fALFF is more sensitive to WM functional variability. Through fALFF differential identifiability, we evaluate the functional variabilities of both WM and GM, and find the overall functional variability pattern is similar although WM shows slightly lower variability than GM. The regional functional variabilities of WM are associated with structural connectivity, where commissural fiber regions generally exhibit higher variability than projection fiber regions. Furthermore, we discover that WM functional variability demonstrates a spatial gradient ascending from the brainstem to the cortex by hypothesis testing, which aligns well with the evolutionary expansion. The gradient of functional variability in WM provides novel insights for understanding WM function. To the best of our knowledge, this is the first attempt to investigate WM functional variability via fALFF. Our code is available at https://github.com/Xinle-Chang/WM-fALFF-Idiff-Gradient.
Xinle Chang, Yang Yang 0002, Yueran Li, Zhengcen Li, Haijin Zeng, Jingyong Su
BIBM5
2025 BrainCognizer: Brain Decoding with Human Visual Cognition Simulation for fMRI-to-Image Reconstruction
abstract
Brain decoding is a key neuroscience field that reconstructs the visual stimuli from brain activity with fMRI, which helps illuminate how the brain represents the world. fMRI-to-image reconstruction has achieved impressive progress by leveraging diffusion models. However, brain signals infused with prior knowledge and associations exhibit a significant information asymmetry when compared to raw visual features, still posing challenges for decoding fMRI representations under the supervision of images. Consequently, the reconstructed images often lack fine-grained visual fidelity, such as missing attributes and distorted spatial relationships. To tackle this challenge, we propose BrainCognizer, a novel brain decoding model inspired by human visual cognition, which explores multilevel semantics and correlations without fine-tuning of generative models. Specifically, BrainCognizer introduces two modules: the Cognitive Integration Module which incorporates prior human knowledge to extract hierarchical region semantics; and the Cognitive Correlation Module which captures contextual semantic relationships across regions. Incorporating these two modules enhances intra-region semantic consistency and maintains interregion contextual associations, thereby facilitating fine-grained brain decoding. Moreover, we quantitatively interpret our components from a neuroscience perspective and analyze the associations between different visual patterns and brain functions. Extensive quantitative and qualitative experiments demonstrate that BrainCognizer outperforms state-of-the-art approaches on multiple evaluation metrics. Our code is released publicly at https://github.com/Grace160/BrainCognizer.
Guoying Sun, Weiyu Guo, Tong Shao, Yang Yang 0002, Haijin Zeng, Jingyong Su
BIBM5
2025 Vision-Language Gradient Descent-driven All-in-One Deep Unfolding Networks
abstract
Dynamic image degradations, including noise, blur and lighting inconsistencies, pose significant challenges in image restoration, often due to sensor limitations or adverse environmental conditions. Existing Deep Unfolding Networks (DUNs) offer stable restoration performance but require manual selection of degradation matrices for each degradation type, limiting their adaptability across diverse scenarios. To address this issue, we propose the Vision-Language-guided Unfolding Network (VLU-Net), a unified DUN framework for handling multiple degradation types simultaneously. VLU-Net leverages a VisionLanguage Model (VLM) refined on degraded image-text pairs to align image features with degradation descriptions, selecting the appropriate transform for target degradation. By integrating an automatic VLM-based gradient estimation strategy into the Proximal Gradient Descent (PGD) algorithm, VLU-Net effectively tackles complex multi-degradation restoration tasks while maintaining interpretability. Furthermore, we design a hierarchical feature unfolding structure to enhance VLU-Net framework, efficiently synthesizing degradation patterns across various levels. VLU-Net is the first all-in-one DUN framework and outperforms current leading one-by-one and all-in-one end- to-end methods by 3.74 dB on the SOTS dehazing dataset and 1.70 dB on the Rain100L deraining dataset.
Haijin Zeng, Xiangming Wang, Yongyong Chen, Jingyong Su
CVPR1
2025 Binarized Mamba-Transformer for Lightweight Quad Bayer HybridEVS Demosaicing
abstract
Quad Bayer demosaicing is the central challenge for enabling the widespread application of Hybrid Event-based Vision Sensors (HybridEVS). Although existing learning-based methods that leverage long-range dependency modeling have achieved promising results, their complexity severely limits deployment on mobile devices for real-world applications. To address these limitations, we propose a lightweight Mamba-based binary neural network designed for efficient and high-performing demosaicing of HybridEVS RAW images. First, to effectively capture both global and local dependencies, we introduce a hybrid Binarized Mamba-Transformer architecture that combines the strengths of the Mamba and Swin Transformer architectures. Next, to significantly reduce computational complexity, we propose a binarized Mamba (Bi-Mamba), which binarizes all projections while retaining the core Selective Scan in full precision. Bi-Mamba also incorporates additional global visual information to enhance global context and mitigate precision loss. We conduct quantitative and qualitative experiments to demonstrate the effectiveness of BMTNet in both performance and computational efficiency, providing a lightweight demosaicing solution suited for real-world edge devices. Our codes and models are available at https://github.com/Clausy9/BMTNet.
Haijin Zeng, Yunfan Lu, Tong Shao, Yongyong Chen, Jingyong Su
CVPR2
2025 Spectral Compressive Imaging via Unmixing-driven Subspace Diffusion Refinement
abstract
Spectral Compressive Imaging (SCI) reconstruction is inherently ill-posed because a single observation admits multiple plausible reconstructions. Traditional deterministic methods struggle to effectively recover high-frequency details. Although diffusion models offer promising solutions to this challenge, their application is constrained by the limited training data and high computational demands associated with multispectral images (MSIs), making direct diffusion training impractical. To address these issues, we propose a novel Predict-and-unmixing-driven-Subspace-Refine framework (PSR-SCI). This framework begins with a light-weight predictor that produces an initial, rough estimate of the MSI. Subsequently, we introduce a unmixing-driven reversible spectral embedding module that decomposes the MSI into subspace images and spectral coefficients. This compact representation facilitates the adaptation of pre-trained RGB diffusion models and focuses refinement processes on high-frequency details, thereby enabling efficient diffusion generation with minimal MSI data. Additionally, we design a high-dimensional guidance mechanism enforcing SCI consistency during sampling. The refined subspace image is then reconstructed back into an MSI using the reversible embedding, yielding the final MSI with full spectral resolution. Experimental results on the standard KAIST and zero-shot datasets NTIRE, ICVL, and Harvard show that PSR-SCI enhances overall visual quality and delivers PSNR and SSIM results competitive with state-of-the-art diffusion, transformer, and deep-unfolding baselines. This framework provides a robust alternative to traditional deterministic SCI reconstruction methods. Code and models are available at [https://github.com/SMARK2022/PSR-SCI](https://github.com/SMARK2022/PSR-SCI).
Haijin Zeng, Benteng Sun, Yongyong Chen, Jingyong Su, Yong Xu 0001
ICLR1
2025 Degradation-Noise-Aware Deep Unfolding Transformer for Hyperspectral Image Denoising
abstract
Hyperspectral images (HSIs) play a pivotal role in fields, such as medical diagnosis and agriculture. However, it often contends with significant noise stemming from narrowband spectral filtering. Existing denoising techniques have their limitations: model-driven methods rely on manual priors and hyperparameters, while learning-based methods struggle to discern intrinsic noise patterns, as they require paired images with specific example noise for training, fail to capture critical noise distribution information, leading to unrobust denoising results. This work addresses the issue by presenting a degradation-noise-aware unfolding network (DNA-Net). Unlike training directly with the simulated noise, DNA-Net initially models general sparse and Gaussian noise through statistic distributions. It then explicitly represents image priors with a customized spectral transformer. The model is subsequently unfolded into an end-to-end (E2E) network, with hyperparameters adaptively estimated from noisy HSI and degradation models, effectively regulating each iteration. Furthermore, a novel U-shaped local-nonlocal–spectral transformer (U-LNSA) is introduced, simultaneously capturing spectral correlations, local features, and nonlocal dependencies. The integration of U-LNSA into DNA-Net establishes the first Transformer-based deep unfolding method for HSI denoising. Experimental results on synthetic and real noise validate DNA-Net’s superior performance over state-of-the-art (SOTA) methods. Moreover, the DNA-Net, trained exclusively on mixed Gaussian noise and impulse noise, demonstrates the ability to generalize to unseen noise present in real images. Code and models will be released at:https://github.com/NavyZeng/DNA-Net.
Haijin Zeng, Xudong Zhao 0003, Jiezhang Cao, Shaoguang Huang, Hiêp Quang Luong, Wilfried Philips
IEEE Trans. Geosci. Remote. Sens.1
2025 Smooth Tensor Qatar Riyal Decomposition for Dynamic MRI Reconstruction
abstract
Dynamic magnetic resonance imaging (dMRI) speed and imaging quality have always been a crucial issue in medical imaging research. Most existing methods characterize the tensor rank-based minimization to reconstruct dMRI from sampling $\bf k$-$t$ space data. However, (1) these approaches that unfold the tensor along each dimension destroy the inherent structure of dMR images. (2) they focus on preserving global information only, while ignoring the local details reconstruction such as the spatial piece-wise smoothness and sharp boundaries. To overcome these obstacles, we suggest a novel low-rank tensor decomposition approach by integrating tensor Qatar Riyal (QR) decomposition, low-rank tensor nuclear norm, and asymmetric total variation to reconstruct dMRI, named TQRTV. Specifically, while preserving the tensor inherent structure by utilizing tensor nuclear norm minimization to approximate tensor rank, QR decomposition reduces the dimensions in the low-rank constraint term, thereby improving the reconstruction performance. TQRTV further exploits the asymmetric total variation regularizer to capture local details. Numerical experiments demonstrate that the proposed reconstruction approach is superior to the existing ones.
Yongyong Chen, Haijin Zeng, Jingyong Su
IEEE J. Biomed. Health Informatics3
2024 DiffSCI: Zero-Shot Snapshot Compressive Imaging via Iterative Spectral Diffusion Model
abstract
This paper endeavors to advance the precision of snap-shot compressive imaging (SCI) reconstruction for multi-spectral image (MSI). To achieve this, we integrate the ad-vantageous attributes of established SCI techniques and an image generative model, propose a novel structured zero-shot diffusion model, dubbed DiffSCI. DiffSCI leverages the structural insights from the deep prior and optimization-based methodologies, complemented by the generative ca-pabilities offered by the contemporary denoising diffusion model. Specifically, firstly, we employ a pre-trained diffusion model, which has been trained on a substantial corpus of RGB images, as the generative denoiser within the Plug-and-Play framework for the first time. This integration allows for the successful completion of SCI reconstruction, especially in the case that current methods struggle to address effectively. Secondly, we systematically account for spectral band correlations and introduce a robust methodology to mitigate wavelength mismatch, thus enabling seamless adaptation of the RGB diffusion model to MSIs. Thirdly, an accelerated algorithm is implemented to expedite the resolution of the data subproblem. This augmentation not only accelerates the convergence rate but also elevates the quality of the reconstruction process. We present extensive testing to show that DiffSCI exhibits discernible performance en-hancements over prevailing self-supervised and zero-shot approaches, surpassing even supervised transformer coun-terparts across both simulated and real datasets. Code is at https://github.com/PAN083/DiffSCI.
Zhenghao Pan, Haijin Zeng, Jiezhang Cao, Kai Zhang 0008, Yongyong Chen
CVPR2
2024 Unmixing Diffusion for Self-Supervised Hyperspectral Image Denoising
abstract
Hyperspectral images (HSIs) have extensive applications in various fields such as medicine, agriculture, and industry. Nevertheless, acquiring high signal-to-noise ratio HSI poses a challenge due to narrow-band spectral filtering. Consequently, the importance of HSI denoising is substantial, especially for snapshot hyperspectral imaging technology. While most previous HSI denoising methods are supervised, creating supervised training datasets for the diverse scenes, hyperspectral cameras, and scan parameters is impractical. In this work, we present Diff-Unmix, a self-supervised denoising method for HSI using diffusion denoising generative models. Specifically, Diff-Unmix addresses the challenge of recovering noise-degraded HSI through a fusion of Spectral Unmixing and conditional abundance generation. Firstly, it employs a learnable block-based spectral unmixing strategy, complemented by a pure transformer-based backbone. Then, we introduce a self-supervised generative diffusion network to enhance abundance maps from the spectral unmixing block. This network reconstructs noise-free Unmixing probability distributions, effectively mitigating noise-induced degradations within these components. Finally, the reconstructed HSI is reconstructed through unmixing reconstruction by blending the diffusion-adjusted abundance map with the spectral endmembers. Experimental results on both simulated and real-world noisy datasets show that Diff-Unmix achieves state-of-the-art performance.
Haijin Zeng, Jiezhang Cao, Kai Zhang 0008, Yongyong Chen, Hiêp Quang Luong, Wilfried Philips
CVPR1
2024 Dual Prior Unfolding for Snapshot Compressive Imaging
abstract
Recently, deep unfolding methods have achieved remarkable success in the realm of Snapshot Compressive Imaging (SCI) reconstruction. However, the existing methods all follow the iterative framework of a single image prior, which limits the efficiency of the unfolding methods and makes it a problem to use other priors simply and effectively. To break out of the box, we derive an effective Dual Prior Unfolding (DPU), which achieves the joint utilization of multiple deep priors and greatly improves iteration efficiency. Our unfolding method is implemented through two parts, i.e., Dual Prior Framework (DPF) and Focused Attention (FA). In brief, in addition to the normal image prior, DPF introduces a residual into the iteration formula and constructs a degraded prior for the residual by considering various degradations to establish the unfolding framework. To improve the effectiveness of the image prior based on self-attention, FA adopts a novel mechanism inspired by PCA denoising to scale and filter attention, which lets the attention focus more on effective features with little computation cost. Besides, an asymmetric backbone is proposed to further improve the efficiency of hierarchical self-attention. Remarkably, our 5-stage DPU achieves state-of-the-art (SOTA) performance with the least FLOPs and parameters compared to previous methods, while our 9-stage DPU significantly outperforms other unfolding methods with less computational requirement. https: / /gi thub. com/ZhangJC-2k/DPU
Jiancheng Zhang 0003, Haijin Zeng, Jiezhang Cao, Yongyong Chen, Dengxiu Yu, Yin-Ping Zhao
CVPR2
2024 Improving Spectral Snapshot Reconstruction with Spectral-Spatial Rectification
abstract
How to effectively utilize the spectral and spatial char-acteristics of Hyperspectral Image (HSI) is always a key problem in spectral snapshot reconstruction. Recently, the spectra-wise transformer has shown great potential in capturing inter-spectra similarities of HSI, but the classic design of the transformer, i.e., multi-head division in the spectral (channel) dimension hinders the modeling of global spectral information and results in mean effect. In addition, previous methods adopt the normal spatial priors without taking imaging processes into account and fail to address the unique spatial degradation in snapshot spectral reconstruction. In this paper, we analyze the influence of multi-head division and propose a novel Spectral-Spatial Recti-fication (SSR) method to enhance the utilization of spectral information and improve spatial degradation. Specifically, SSR includes two core parts: Window-based Spectra-wise Self-Attention (WSSA) and spAtial Rectification Block (ARB). WSSA is proposed to capture global spectral in-formation and account for local differences, whereas ARB aims to mitigate the spatial degradation using a spatial alignment strategy. The experimental results on simulation and real scenes demonstrate the effectiveness of the proposed modules, and we also provide models at multiple scales to demonstrate the superiority of our approach. https://github.com/ZhangJC-2k/SSR
Jiancheng Zhang 0003, Haijin Zeng, Yongyong Chen, Dengxiu Yu, Yin-Ping Zhao
CVPR2
2024 SAH-SCI: Self-supervised Adapter for Efficient Hyperspectral Snapshot Compressive Imaging
Haijin Zeng, Yongyong Chen, Youfa Liu, Chong Peng 0001, Jingyong Su
ECCV (64)1
2024 Wavelength-Embedding-Guided Filter-Array Transformer for Spectral Demosaicing
Haijin Zeng, Hiêp Quang Luong, Wilfried Philips
ECCV (14)1
2024 MambaSCI: Efficient Mamba-UNet for Quad-Bayer Patterned Video Snapshot Compressive Imaging
abstract
Color video snapshot compressive imaging (SCI) employs computational imaging techniques to capture multiple sequential video frames in a single Bayer-patterned measurement. With the increasing popularity of quad-Bayer pattern in mainstream smartphone cameras for capturing high-resolution videos, mobile photography has become more accessible to a wider audience. However, existing color video SCI reconstruction algorithms are designed based on the traditional Bayer pattern. When applied to videos captured by quad-Bayer cameras, these algorithms often result in color distortion and ineffective demosaicing, rendering them impractical for primary equipment. To address this challenge, we propose the MambaSCI method, which leverages the Mamba and UNet architectures for efficient reconstruction of quad-Bayer patterned color video SCI. To the best of our knowledge, our work presents the first algorithm for quad-Bayer patterned SCI reconstruction, and also the initial application of the Mamba model to this task. Specifically, we customize Residual-Mamba-Blocks, which residually connect the Spatial-Temporal Mamba (STMamba), Edge-Detail-Reconstruction (EDR) module, and Channel Attention (CA) module. Respectively, STMamba is used to model long-range spatial-temporal dependencies with linear complexity, EDR is for better edge-detail reconstruction, and CA is used to compensate for the missing channel information interaction in Mamba model. Experiments demonstrate that MambaSCI surpasses state-of-the-art methods with lower computational and memory costs. PyTorch style pseudo-code for the core modules is provided in the supplementary materials. Code is at https://github.com/PAN083/MambaSCI.
Zhenghao Pan, Haijin Zeng, Jiezhang Cao, Yongyong Chen, Kai Zhang 0008, Yong Xu 0001
NeurIPS2
2024 Inheriting Bayer's Legacy: Joint Remosaicing and Denoising for Quad Bayer Image Sensor
Haijin Zeng, Jiezhang Cao, Shaoguang Huang, Yongqiang Zhao 0001, Hiêp Quang Luong, Jan Aelterman, Wilfried Philips
Int. J. Comput. Vis.1
2024 Learnable Spatial-Spectral Transform-Based Tensor Nuclear Norm for Multi-Dimensional Visual Data Recovery
abstract
Recently, transform-based tensor nuclear norm (TNN) methods have received increasing attention as a powerful tool for multi-dimensional visual data (color images, videos, and multispectral images, etc.) recovery. Especially, the redundant transform-based TNN achieves satisfactory recovery results, where the redundant transform along spectral mode can remarkably enhance the low-rankness of tensors. However, it suffers from expensive computational cost induced by the redundant transform. In this paper, we propose a learnable spatial-spectral transform-based TNN model for multi-dimensional visual data recovery, which not only enjoys better low-rankness capability but also allows us to design fast algorithms accompanying it. More specifically, we first project the large-scale original tensor to the small-scale intrinsic tensor via the learnable semi-orthogonal transforms along the spatial modes. Here, the semi-orthogonal transforms, serving as the key building block, can boost the spatial low-rankness and lead to a small-scale problem, which paves the way for designing fast algorithms. Secondly, to further boost the low-rankness, we apply the learnable redundant transform along the spectral mode to the small-scale intrinsic tensor. To tackle the proposed model, we apply an efficient proximal alternating minimization-based algorithm, which enjoys a theoretical convergence guarantee. Extensive experimental results on real-world data (color images, videos, and multispectral images) demonstrate that the proposed method outperforms state-of-the-art competitors in terms of evaluation metrics and running time.
Sheng Liu 0033, Jinsong Leng, Xi-Le Zhao, Haijin Zeng, Yao Wang 0003
IEEE Trans. Circuits Syst. Video Technol.4
2024 Spatial and Cluster Structural Prior-Guided Subspace Clustering for Hyperspectral Image
abstract
Subspace clustering has achieved remarkable performance for hyperspectral image (HSI). However, existing methods are often computationally expensive and have limited ability to capture the intrinsic structural information of HSI. In this paper, we propose a structural prior-guided subspace clustering method, which simultaneously incorporates the local and non-local spatial information and the cluster prior information. Accordingly, three efficient regularizations are developed. Considering the local connectivity of pixels, we propose an ℓ2,1norm based constraint on the representation difference matrix to improve the homogeneity of clustering result. Next, to capture the non-local geometric structure of HSI, we propose a manifold-based regularization with an adaptively learned landmark graph. Furthermore, we explore the block-diagonal cluster structure of HSI and develop a landmark-based clustering constraint, which makes the representations more favorable for clustering. Our local constraint is imposed on all the data points due to its efficiency and the latter two are solely imposed on landmarks, leading to computationally efficient regularizations. Due to the local constraint, the manifold and cluster structure of the landmarks can be effectively propagated to all the data points. To make our model scalable to large-scale data, we learn a compact dictionary with an orthogonal constraint, significantly reducing the number of parameters. In addition, we propose a novel landmark selection method to support our landmark-based constraints using multi-scale super-pixel segmentation and clustering, which improves the uniformity and diversity of landmarks. We also develop an efficient algorithm to solve the proposed model. Experimental results demonstrate that our model outperforms the state-of-the-art.
Shaoguang Huang, Haijin Zeng, Hongyu Chen 0003, Hongyan Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Unsupervised Spectral Demosaicing With Lightweight Spectral Attention Networks
abstract
This paper presents a deep learning-based spectral demosaicing technique trained in an unsupervised manner. Many existing deep learning-based techniques relying on supervised learning with synthetic images, often underperform on real-world images, especially as the number of spectral bands increases. This paper presents a comprehensive unsupervised spectral demosaicing (USD) framework based on the characteristics of spectral mosaic images. This framework encompasses a training method, model structure, transformation strategy, and a well-fitted model selection strategy. To enable the network to dynamically model spectral correlation while maintaining a compact parameter space, we reduce the complexity and parameters of the spectral attention module. This is achieved by dividing the spectral attention tensor into spectral attention matrices in the spatial dimension and spectral attention vector in the channel dimension. This paper also presents Mosaic 25 , a real 25-band hyperspectral mosaic image dataset featuring various objects, illuminations, and materials for benchmarking purposes. Extensive experiments on both synthetic and real-world datasets demonstrate that the proposed method outperforms conventional unsupervised methods in terms of spatial distortion suppression, spectral fidelity, robustness, and computational cost. Our code and dataset are publicly available at https://github.com/polwork/Unsupervised-Spectral-Demosaicing.
Haijin Zeng, Yongqiang Zhao 0001, Seong G. Kong, Yuanyang Bu
IEEE Trans. Image Process.2
2024 Tensor Completion Using Bilayer Multimode Low-Rank Prior and Total Variation
abstract
In this article, we propose a novel bilayer low-rankness measure and two models based on it to recover a low-rank (LR) tensor. The global low rankness of underlying tensor is first encoded by LR matrix factorizations (MFs) to the all-mode matricizations, which can exploit multiorientational spectral low rankness. Presumably, the factor matrices of all-mode decomposition are LR, since local low-rankness property exists in within-mode correlation. In the decomposed subspace, to describe the refined local LR structures of factor/subspace, a new low-rankness insight of subspace: a double nuclear norm scheme is designed to explore the so-called second-layer low rankness. By simultaneously representing the bilayer low rankness of the all modes of the underlying tensor, the proposed methods aim to model multiorientational correlations for arbitrary N -way ( N ≥ 3 ) tensors. A block successive upper-bound minimization (BSUM) algorithm is designed to solve the optimization problem. Subsequence convergence of our algorithms can be established, and the iterates generated by our algorithms converge to the coordinatewise minimizers in some mild conditions. Experiments on several types of public datasets show that our algorithm can recover a variety of LR tensors from significantly fewer samples than its counterparts.
Haijin Zeng, Shaoguang Huang, Yongyong Chen, Sheng Liu 0033, Hiêp Quang Luong, Wilfried Philips
IEEE Trans. Neural Networks Learn. Syst.1
2023 Asymmetry total variation and framelet regularized nonconvex low-rank tensor completion
Yongyong Chen, Xiaojia Zhao, Haijin Zeng, Yanhui Xu, Junxing Chen
Signal Process.4
2023 Generalized Nonconvex Low-Rank Tensor Representation for Hyperspectral Anomaly Detection
abstract
Low-rank tensor representation (LRTR) methods have attracted great interest for their powerful ability to separate backgrounds and anomalies. However, most of the current LRTR models use the popular and convex surrogate tensor nuclear norm to solve optimization problems, which results in a loose approximation and suboptimal solver for the original problem. Besides, most existing methods solve the nonconvex optimization problems case-by-case, consequently losing one unified solver. To solve the above issues, we propose the Generalized Nonconvex Low-rank Tensor Representation (GNLTR) for hyperspectral anomaly detection (HAD), a unified solver not case-by-case one of existing nonconvex optimization problems. Compared to the tensor nuclear norm, GNLTR contains many popular nonconvex penalty functions as tighter regularizers of the tensor tubal rank to constrain the low rank of the background. Moreover, theL2,1norm has been integrated into the GNLTR model for the sparse anomalies. For the optimization problem, it is handled quickly and efficiently through a well-organized alternating direction method of multipliers (ADMM). The experiments on several real-world hyperspectral data sets demonstrate the superior performance of the GNLTR model in comparison with some state-of-the-art anomaly detection models.
Qiangqiang Shen, Haijin Zeng, Yongyong Chen, Guangming Lu 0002
IEEE Trans. Geosci. Remote. Sens.3
2023 Multimodal Core Tensor Factorization and its Applications to Low-Rank Tensor Completion
abstract
Low-rank tensor completion has been widely used in computer vision and machine learning. This paper develops a novel multimodal core tensor factorization (MCTF) method combined with a tensor low-rankness measure and a better nonconvex relaxation form of this measure (NC-MCTF). The proposed models encode low-rank insights for general tensors provided by Tucker and T-SVD and thus are expected to simultaneously model spectral low-rankness in multiple orientations and accurately restore the data of intrinsic low-rank structure based on few observed entries. Furthermore, we study the MCTF and NC-MCTF regularization minimization problem and design an effective block successive upper-bound minimization (BSUM) algorithm to solve them. Theoretically, we prove that the iterates generated by the proposed models converge to the set of coordinatewise minimizers. This efficient solver can extend MCTF to various tasks such as tensor completion. A series of experiments including hyperspectral image (HSI), video and MRI completion confirm the superior performance of the proposed method.
Haijin Zeng, Jize Xue, Hiêp Quang Luong, Wilfried Philips
IEEE Trans. Multim.1
2021 Hyperspectral image denoising via global spatial-spectral total variation regularized nonconvex local low-rank tensor approximation
Haijin Zeng, Xiaozhen Xie, Jifeng Ning
Signal Process.1
2021 Hyperspectral Image Restoration via Global L1-2 Spatial-Spectral Total Variation Regularized Local Low-Rank Tensor Recovery
abstract
Hyperspectral images (HSIs) are usually corrupted by various noises, e.g., Gaussian noise, impulse noise, stripes, dead lines, and many others. In this article, motivated by the good performance of the L1-2nonconvex metric in image sparse structure exploitation, we first develop a 3-D L1-2spatial-spectral total variation ( L1-2SSTV) regularization to globally represent the sparse prior in the gradient domain of HSIs. Then, we divide HSIs into local overlapping 3-D patches, and low-rank tensor recovery (LTR) is locally used to effectively separate the low-rank clean HSI patches from complex noise. The patchwise LTR can not only adapt to the local low-rank property of HSIs well but also significantly reduce the information loss caused by the global LTR. Finally, integrating the advantages of both the global L1-2SSTV regularization and local LTR model, we propose a L1-2SSTV regularized local LTR model for hyperspectral restoration. In the framework of the alternating direction method of multipliers, the difference of convex algorithm, the split Bregman iteration method, and tensor singular value decomposition method are adopted to solve the proposed model efficiently. Simulated and real HSI experiments show that the proposed model can reduce the dependence on noise independent and identical distribution hypotheses, and simultaneously remove various types of noise, even structure-related noise.
Haijin Zeng, Xiaozhen Xie, Haojie Cui, Hanping Yin, Jifeng Ning
IEEE Trans. Geosci. Remote. Sens.1
2020 Hyperspectral Image Restoration via Global Total Variation Regularized Local Nonconvex Low-Rank Matrix Approximation
abstract
Several bandwise total variation (TV) regularized low-rank (LR)-based models have been proposed to remove mixed noise in hyperspectral images (HSIs). Conventionally, the rank of LR matrix is approximated using nuclear norm (NN). The NN is defined by adding all singular values together, which is essentially a L1-norm of the singular values. It results in non-negligible approximation errors and thus the resulting matrix estimator can be significantly biased. Moreover, these bandwise TV-based methods exploit the spatial information in a separate manner. To cope with these problems, we propose a spatial-spectral TV (SSTV) regularized non-convex local LR matrix approximation (NonLLRTV) method to remove mixed noise in HSIs. From one aspect, local LR of HSIs is formulated using a non-convex Lγ-norm, which provides a closer approximation to the matrix rank than the traditional NN. From another aspect, HSIs are assumed to be piecewisely smooth in the global spatial domain. The TV regularization is effective in preserving the smoothness and removing Gaussian noise. These facts inspire the integration of the NonLLR with TV regularization. To address the limitations of bandwise TV, we use the SSTV regularization to simultaneously consider global spatial structure and spectral correlation of neighboring bands. Experiment results indicate that the use of local non-convex penalty and global SSTV can boost the preserving of spatial piecewise smoothness and overall structural information.
Haijin Zeng, Xiaozhen Xie, Jifeng Ning
IGARSS1
2020 Hyperspectral image restoration via CNN denoiser prior regularized low-rank tensor recovery
Haijin Zeng, Xiaozhen Xie, Haojie Cui, Yuan Zhao 0016, Jifeng Ning
Comput. Vis. Image Underst.1