Liheng Bian

dblp:139/0758 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
11since 2021 · last 2026
0000-0002-8016-0375ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HySaDe-Mamba: A Mamba-Based Network for Hyperspectral Salient Object Detection
abstract
Hyperspectral salient object detection (HSOD) aims to identify visually and spectrally distinctive regions in hyperspectral images (HSIs). However, existing HSOD methods often suffer from spectral redundancy and inefficient spatial-spectral modeling, which hinder their scalability and accuracy in complex scenes. To tackle these challenges, we propose HySaDe-Mamba, a novel HSOD framework built upon the Mamba architecture. Specifically, to address information redundancy in HSI, we design a spatial-enhanced spectral-embedding (SeSe) module, which maps high-dimensional data into a more compact but effective representation. On the compact SeSe representation features, we further propose a Bi-scale spatial and Bi-directional spectral (BsBd) Mamba module, performing the selective scanning mechanism in a spatial-spectral hybrid, end-to-end way, which not only facilitates comprehensive spatial structural interaction across both global and local scales, but also effectively exploits the underlying spectral semantic correlation. Extensive experiments on two public HSOD datasets demonstrate that our HySaDe-Mamba achieves state-of-the-art detection accuracy across seven metrics, while maintaining an efficient inference speed of 40.22 FPS. The source code is publicly available at https://github.com/Leezl/HySaDe-Mamba.
Lei Zhang 0109, Xiaoyan Luo, Liheng Bian, Xiantong Zhen
IEEE Trans. Circuits Syst. Video Technol.4
2025 Adaptive Dual-domain Learning for Underwater Image Enhancement
abstract
Recently, learning-based Underwater Image Enhancement (UIE) methods have demonstrated promising performance. However, existing learning-based methods still face two challenges. 1) They rarely consider the inconsistent degradation levels in different spatial regions and spectral bands simultaneously. 2) They treat all regions equally, ignoring that the regions with high-frequency details are more difficult to reconstruct. To address these challenges, we propose a novel UIE method based on spatial-spectral dual-domain adaptive learning, termed SS-UIE. Specifically, we first introduce a spatial-wise Multi-scale Cycle Selective Scan (MCSS) module and a Spectral-Wise Self-Attention (SWSA) module, both with linear complexity, and combine them in parallel to form a basic Spatial-Spectral block (SS-block). Benefiting from the global receptive field of MCSS and SWSA, SS-block can effectively model the degradation levels of different spatial regions and spectral bands, thereby enabling degradation level-based dual-domain adaptive UIE. By stacking multiple SS-blocks, we build our SS-UIE network. Additionally, a Frequency-Wise Loss (FWL) is introduced to narrow the frequency-wise discrepancy and reinforce the model's attention on the regions with high-frequency details. Extensive experiments validate that the SS-UIE technique outperforms state-of-the-art UIE methods while requiring cheaper computational and memory costs.
Lintao Peng, Liheng Bian
AAAI2
2025 S3A: A Self-Supervised Saliency Analysis Framework for Hyperspectral Image
abstract
Most existing visual saliency analysis methods are meticulously crafted for natural RGB images, which typically lean upon foundational spatial visual cues, such as contrast, structure and texture in scenes. However, these methods exhibit limitations in spectral saliency analysis due to the existence of numerous materials that may manifest identical RGB values while harboring disparate spectral characteristics. Consequently, it is an exigency to introduce saliency method specially for hyperspectral images (HSIs). In this paper, we establish a set of hyperspectral saliency principles that incorporate both spectral and spatial attributes, and accordingly present a novel self-supervised saliency analysis (S3A) framework for HSIs. Note that our S3A is designed as a discriminant architecture composed of three interconnected components. More specifically, an unsupervised HSI representive patch sampling (RPS) module is designed to pick up some representative pixels for self-supervised saliency analysis training. Subsequently, we construct a classification-based patch spectral discrimination network (CPSD-Net) to evaluate the HSI saliency. Finally, a patch-to-global spatial diffusion (P2G-SD) module is constructed to diffuse the saliency from few supervised samples to the other HSI pixels. Moreover, to demonstrate the performance of our saliency analysis framework, we apply the HSI saliency to some classic downstream tasks including band selection (BS) and HSI classification. On several popular HSI datasets, the satisfactory quantization results fully verify the rationality and effectiveness of our S3A framework, in terms of entropy value and mean spectral divergence (MSD) of the selected bands in BS task, as well as the accuracy in HSI classification task. The source code is publicly available at https://github.com/Lee-zl/S3A.
Xiaoyan Luo, Lei Zhang 0109, Peixin Gan, Liheng Bian
IEEE Trans. Geosci. Remote. Sens.5
2025 Uncertainty-Driven Parallel Transformer-Based Segmentation for Oral Disease Dataset
abstract
Accurate oral disease segmentation is a challenging task, for three major reasons: 1) The same type of oral disease has a diversity of size, color and texture; 2) The boundary between oral lesions and their surrounding mucosa is not sharp; 3) There is a lack of public large-scale oral disease segmentation datasets. To address these issues, we first report an oral disease segmentation network termed Oralformer, which enables to tackle multiple oral diseases. Specifically, we use a parallel design to combine local-window self-attention (LWSA) with channel-wise convolution (CWC), modeling cross-window connections to enlarge the receptive fields while maintaining linear complexity. Meanwhile, we connect these two branches with bi-directional interactions to form a basic parallel Transformer block namely LC-block. We insert the LC-block as the main building block in a U-shape encoder-decoder architecture to form Oralformer. Second, we introduce an uncertainty-driven self-adaptive loss function which can reinforce the network's attention on the lesion's edge regions that are easily confused, thus improving the segmentation accuracy of these regions. Third, we construct a large-scale oral disease segmentation (ODS) dataset containing 2602 image pairs. It covers three common oral diseases (including dental plaque, calculus and caries) and all age groups, which we hope will advance the field. Extensive experiments on six challenging datasets show that our Oralformer achieves state-of-the-art segmentation accuracy, and presents advantages in terms of generalizability and real-time segmentation efficiency (35fps). The code and ODS dataset will be publicly available at https://github.com/LintaoPeng/Oralformer.
Lintao Peng, Siyu Xie, Fei Xiao 0003, Liheng Bian
IEEE Trans. Image Process.7
2025 Efficient High-Fidelity Global Low-Rank Optimization for Multispectral Demosaicing
abstract
The nonlocal low-rank (NLR) optimization has shown promise for generalized multispectral filter array (MSFA) demosaicing. However, it faces challenges in balancing efficiency and accuracy. To tackle these challenges, we report here the multi-channel global low-rank optimization technique, achieving efficient high-fidelity MSFA demosaicing. Inspired by the cross-band correlations of natural multispectral images, we introduce the multi-channel matching and low-rank strategies that jointly optimize image patches of all channels, exhibiting higher efficiency and accuracy than existing approaches. Furthermore, we present global structural matching (GSM) which performs structure-aware multi-channel matching across the entire multispectral image. GSM extracts structurally important patches and efficiently searches their similar patches via parallel correlation, providing an order-of-magnitude improvement in efficiency. By combining the aforementioned techniques, we have achieved superior performance over the state-of-the-art NLR demosaicing technique, leading to up to 3.9 dB peak signal-to-noise ratio (PSNR) gain and over a 150-fold increase in computational speed. Experiments validated that the technique outperforms existing methods in reconstructing fine textures and details and exhibits superior robustness to noise.
Daoyu Li, Xin Yuan 0002, Liheng Bian
IEEE Trans. Multim.5
2024 Diffusion-based Blind Text Image Super-Resolution
abstract
Recovering degraded low-resolution text images is chal-lenging, especially for Chinese text images with complex strokes and severe degradation in real-world scenarios. En-suring both text fidelity and style realness is crucial for high-quality text image super-resolution. Recently, diffusion models have achieved great success in natural image synthesis and restoration due to their powerful data distribution modeling abilities and data generation capabili-ties. In this work, we propose an Image Diffusion Model (IDM) to restore text images with realistic styles. For diffusion models, they are not only suitable for modeling realis-tic image distribution but also appropriate for learning text distribution. Since text prior is important to guarantee the correctness of the restored text structure according to existing arts, we also propose a Text Diffusion Model (TDM) for text recognition which can guide IDM to generate text images with correct structures. We further propose a Mixture of Multi-modality module (MoM) to make these two diffusion models cooperate with each other in all the diffusion steps. Extensive experiments on synthetic and real-world datasets demonstrate that our Diffusion-based Blind Text Image Super-Resolution (DiffTSR) can restore text images with more accurate text structures as well as more realistic appearances simultaneously. Code is available at https://github.com/YuzheZhang-1999/DiffTSR.
Yuzhe Zhang 0004, Zhouxia Wang, Luwei Hou, Dongqing Zou, Liheng Bian
CVPR7
2024 Uncertainty-Driven Spectral Compressive Imaging with Spatial-Frequency Transformer
Lintao Peng, Siyu Xie, Liheng Bian
ECCV (6)3
2023 Generalized Imaging Augmentation via Linear Optimization of Neurons
Daoyu Li, Liheng Bian
BMVC4
2023 U-Shape Transformer for Underwater Image Enhancement
abstract
The light absorption and scattering of underwater impurities lead to poor underwater imaging quality. The existing data-driven based underwater image enhancement (UIE) techniques suffer from the lack of a large-scale dataset containing various underwater scenes and high-fidelity reference images. Besides, the inconsistent attenuation in different color channels and space areas is not fully considered for boosted enhancement. In this work, we built a large scale underwater image (LSUI) dataset, which covers more abundant underwater scenes and better visual quality reference images than existing underwater datasets. The dataset contains 4279 real-world underwater image groups, in which each raw image's clear reference images, semantic segmentation map and medium transmission map are paired correspondingly. We also reported an U-shape Transformer network where the transformer model is for the first time introduced to the UIE task. The U-shape Transformer is integrated with a channel-wise multi-scale feature fusion transformer (CMSFFT) module and a spatial-wise global feature modeling transformer (SGFMT) module specially designed for UIE task, which reinforce the network's attention to the color channels and space areas with more serious attenuation. Meanwhile, in order to further improve the contrast and saturation, a novel loss function combining RGB, LAB and LCH color spaces is designed following the human vision principle. The extensive experiments on available datasets validate the state-of-the-art performance of the reported technique with more than 2dB superiority. The dataset and demo code are available at https://bianlab.github.io/.
Lintao Peng, Chunli Zhu, Liheng Bian
IEEE Trans. Image Process.3
2023 INFWIDE: Image and Feature Space Wiener Deconvolution Network for Non-Blind Image Deblurring in Low-Light Conditions
abstract
Under low-light environment, handheld photography suffers from severe camera shake under long exposure settings. Although existing deblurring algorithms have shown promising performance on well-exposed blurry images, they still cannot cope with low-light snapshots. Sophisticated noise and saturation regions are two dominating challenges in practical low-light deblurring: the former violates the Gaussian or Poisson assumption widely used in most existing algorithms and thus degrades their performance badly, while the latter introduces non-linearity to the classical convolution-based blurring model and makes the deblurring task even challenging. In this work, we propose a novel non-blind deblurring method dubbed image and feature space Wiener deconvolution network (INFWIDE) to tackle these problems systematically. In terms of algorithm design, INFWIDE proposes a two-branch architecture, which explicitly removes noise and hallucinates saturated regions in the image space and suppresses ringing artifacts in the feature space, and integrates the two complementary outputs with a subtle multi-scale fusion network for high quality night photograph deblurring. For effective network training, we design a set of loss functions integrating a forward imaging model and backward reconstruction to form a close-loop regularization to secure good convergence of the deep neural network. Further, to optimize INFWIDE's applicability in real low-light conditions, a physical-process-based low-light noise model is employed to synthesize realistic noisy night photographs for model training. Taking advantage of the traditional Wiener deconvolution algorithm's physically driven characteristics and deep neural network's representation ability, INFWIDE can recover fine details while suppressing the unpleasant artifacts during deblurring. Extensive experiments on synthetic data and real data demonstrate the superior performance of the proposed approach.
Zhihong Zhang 0004, Yuxiao Cheng, Jin-Li Suo, Liheng Bian, Qionghai Dai
IEEE Trans. Image Process.4
2021 Generalized MSFA Engineering With Structural and Adaptive Nonlocal Demosaicing
abstract
The emerging multispectral-filter-array (MSFA) cameras require generalized demosaicing for MSFA engineering. The existing interpolation, compressive sensing and deep learning based methods suffer from either limited reconstruction accuracy or poor generalization. In this work, we report a generalized demosaicing method with structural and adaptive nonlocal optimization, enabling boosted reconstruction accuracy for different MSFAs. The advantages lie in the following three aspects. First, the nonlocal low-rank optimization is applied and extended to the multiple spatial-spectral-temporal dimensions to exploit more crucial details. Second, the block matching accuracy is promoted by employing a novel structural similarity metric instead of the conventional Euclidean distance. Third, the running efficiency is boosted by an adaptive iteration strategy. We built a prototype system to capture raw mosaic images under different MSFAs, and used the technique as an off-the-shelf tool to demonstrate MSFA engineering. The experiments show that the binary tree (BT) based filter array produces higher accuracy than the random and regular ones for different number of channels.
Liheng Bian, Yugang Wang
IEEE Trans. Image Process.1
2016 Signal-dependent noise removal for color videos using temporal and cross-channel priors
Jin-Li Suo, Liheng Bian, Feng Chen 0007, Qionghai Dai
J. Vis. Commun. Image Represent.2
2014 Automatic inpainting of linearly related video frames
abstract
This paper addresses automatic inpainting of a specific but common kind of videos captured by imaging a far or planar scene with a moving camera. The projective model tells that the frames of such videos can be approximately aligned by linear mappings except for some to-be-inpainted small regions. Mathematically, we treat inpainting as a global optimization with a linear system incorporating both the temporal consistency and the priors of the inpainting regions: (i) temporally registered frames form a low rank matrix; (ii) the pixels in the given inpainting regions destroy the low rank-ness with gross sparse errors. Besides, we also use a soft mask to ensure consistent global brightness before and after inpainting. Further, we propose a numerical solution to above optimization based on Augmented Lagrangian Method. The experiment results demonstrated our advantageous in both preserving thin scene structures and the details prone to be smoothed out by previous methods.
Yudong Xiao, Jin-Li Suo, Liheng Bian, Qionghai Dai
ICIP3
2014 Joint Non-Gaussian Denoising and Superresolving of Raw High Frame Rate Videos
abstract
High frame rate cameras capture sharp videos of highly dynamic scenes by trading off signal-noise-ratio and image resolution, so combinational super-resolving and denoising is crucial for enhancing high speed videos and extending their applications. The solution is nontrivial due to the fact that two deteriorations co-occur during capturing and noise is nonlinearly dependent on signal strength. To handle this problem, we propose conducting noise separation and super resolution under a unified optimization framework, which models both spatiotemporal priors of high quality videos and signal-dependent noise. Mathematically, we align the frames along temporal axis and pursue the solution under the following three criterion: 1) the sharp noise-free image stack is low rank with some missing pixels denoting occlusions; 2) the noise follows a given nonlinear noise model; and 3) the recovered sharp image can be reconstructed well with sparse coefficients and an over complete dictionary learned from high quality natural images. In computation aspects, we propose to obtain the final result by solving a convex optimization using the modern local linearization techniques. In the experiments, we validate the proposed approach in both synthetic and real captured data.
Jin-Li Suo, Yue Deng 0001, Liheng Bian, Qionghai Dai
IEEE Trans. Image Process.3