VLDB 2026 Research / reviewers in the wild / expert
Lihui Chen 0002
dblp:56/1277-2
· DBLP profile ↗
23ranked-venue papers
6as first author
21since 2021 · last 2026
0000-0002-0948-1600ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 2 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Remote sensing optical image matching through neighborhood-aware global propagation in graph neural networks
Yanchun Liu, Gemine Vivone, Jing Nie 0001, Haijun Liu 0001, Xichuan Zhou, Lihui Chen 0002 |
Eng. Appl. Artif. Intell. | 6 |
| 2026 | Multiscale wavelet-based spatial-spectral compression network for hyperspectral image
Mingyang Wan, Aibin Peng, Xiangfei Shen, Rulong He, Lihui Chen 0002, Haijun Liu 0001, Xichuan Zhou |
Eng. Appl. Artif. Intell. | 7 |
| 2026 | Sparse gain adaptation with dual-domain fusion network for multimodal object detection
Xichuan Zhou, Boya Wei, Cong Mao, Lihui Chen 0002, Haijun Liu 0001, Jin Xie 0005, Jing Nie 0001 |
Neurocomputing | 6 |
| 2026 | RA-PTQ: Reparameterization-Aware Post-Training Quantization for accurate vision transformers in low-bit scenarios
Rui Ding 0009, Sihuan Zhao, Jing Nie 0001, Lihui Chen 0002, Haijun Liu 0001, Xichuan Zhou |
Knowl. Based Syst. | 5 |
| 2026 | Multimodality Image Registration With Modality DistillationabstractMultimodal image registration aims to spatially align images from different modalities at the pixel level. However, due to the nonlinear relationship of radiation intensities caused by different imaging modalities, achieving high accuracy in multimodal image registration presents a significant challenge. Additionally, the presence of both global transformations (i.e., large-scale rigid affine transformations) and local distortions (i.e., small-scale nonrigid deformations) between paired images further complicates the registration process. This article addressed the challenge resulting from modality differences through modality distillation. Specifically, a teacher (i.e., a homomodal image registration model) is trained to guide the student (i.e., a multimodal image registration model). Besides, this article simultaneously aligned large-scale rigid and small-scale nonrigid deformations by predicting deformation flow from both global and local features, thereby achieving high-precision registration. Furthermore, this proposed method incorporated a deformation mask during training to mitigate the negative impact of black edges in the obtained registration results on model performance. Experimental results demonstrate that the proposed method delivers state-of-the-art registration accuracy across various multimodal datasets, with ablation studies confirming the effectiveness of each component. The codes will be available at https://github.com/2351056918/Multimodality-Image-Registration-with-Modailty-Distillation. Xichuan Zhou, Jicheng Zhao, Lihui Chen 0002, Gemine Vivone, Yanchun Liu, Jing Nie 0001, Haijun Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | EigenSR: Eigenimage-Bridged Pre-Trained RGB Learners for Single Hyperspectral Image Super-ResolutionabstractSingle hyperspectral image super-resolution (single-HSI-SR) aims to improve the resolution of a single input low-resolution HSI. Due to the bottleneck of data scarcity, the development of single-HSI-SR lags far behind that of RGB natural images. In recent years, research on RGB SR has shown that models pre-trained on large-scale benchmark datasets can greatly improve performance on unseen data, which may stand as a remedy for HSI. But how can we transfer the pre-trained RGB model to HSI, to overcome the data-scarcity bottleneck? Because of the significant difference in the channels between the pre-trained RGB model and the HSI, the model cannot focus on the correlation along the spectral dimension, thus limiting its ability to utilize on HSI. Inspired by the HSI spatial-spectral decoupling, we propose a new framework that first fine-tunes the pre-trained model with the spatial components (known as eigenimages), and then infers on unseen HSI using an iterative spectral regularization (ISR) to maintain the spectral correlation. The advantages of our method lie in: 1) we effectively inject the spatial texture processing capabilities of the pre-trained RGB model into HSI while keeping spectral fidelity, 2) learning in the spectral-decorrelated domain can improve the generalizability to spectral-agnostic data, and 3) our inference in the eigenimage domain naturally exploits the spectral low-rank property of HSI, thereby reducing the complexity. This work bridges the gap between pre-trained RGB models and HSI via eigenimages, addressing the issue of limited HSI training data, hence the name EigenSR. Extensive experiments show that EigenSR outperforms the state-of-the-art (SOTA) methods in both spatial and spectral metrics. Xi Su, Xiangfei Shen, Mingyang Wan, Jing Nie 0001, Lihui Chen 0002, Haijun Liu 0001, Xichuan Zhou |
AAAI | 5 |
| 2025 | Hybrid cross-modality fusion network for medical image segmentation with contrastive learning
Xichuan Zhou, Jing Nie 0001, Haijun Liu 0001, Fu Liang, Lihui Chen 0002, Jin Xie 0005 |
Eng. Appl. Artif. Intell. | 7 |
| 2025 | Progressive fine-to-coarse reconstruction for accurate low-bit post-training quantization in vision transformers
Rui Ding 0009, Liang Yong, Sihuan Zhao, Jing Nie 0001, Lihui Chen 0002, Haijun Liu 0001, Xichuan Zhou |
Neural Networks | 5 |
| 2025 | High-Fidelity Pansharpening via Trigeminal Pyramid Decoding of CNN-Transformer Encoded FeaturesabstractSpectral and spatial fidelity remains a longstanding challenge in the field of pansharpening, which aims to generate high-resolution multispectral (HRMS) images by integrating high-resolution panchromatic (PAN) images with low-resolution multispectral (LRMS) images. This study proposes a high-fidelity pansharpening network that utilizes bidirectional trigeminal pyramid decoding of features encoded by a CNN-Transformer architecture. Specifically, local and global features at multiple scales are initially extracted using a CNN-Transformer encoder to facilitate multi-scale feature fusion. Subsequently, we design a decoder based on bidirectional trigeminal pyramids to achieve a high-fidelity fusion output. One reverse decoding pyramid decodes the fused features of LRMS and PAN images from the encoder. One spectral feature pyramid is employed to enhance the spectral information of the reverse decoding pyramid, while the last spatial feature pyramid is utilized to enrich the spatial information, thereby improving the overall spectral and spatial fidelity of the fused output. Furthermore, content-guided attention (CGA) is incorporated to adaptively integrate the spectral and spatial feature pyramids into the reverse decoding pyramid. Extensive experiments demonstrate that our network surpasses the comparative state-of-the-art (SOTA) methods in both qualitative and quantitative evaluations. The code is available at https://github.com/songvvvv/pansharpening. Lihui Chen 0002, Tianxin Song, Lihua Jian, Di Zhang 0002, Gemine Vivone, Xichuan Zhou |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Bi-SSFormer: An Ultralightweight Binary Spectral-Spatial Transformer for Hyperspectral Image Classification
Rui Ding 0009, Yanchun Liu, Baoliang Wang, Lihui Chen 0002, Haijun Liu 0001, Gemine Vivone, Xichuan Zhou |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | MIT-SAM: Medical Image-Text SAM With Mutually Enhanced Heterogeneous Features Fusion for Medical Image SegmentationabstractIn recent times, leveraging lesion text as supplementary data to enhance the performance of medical image segmentation models has garnered attention. Previous approaches only used attention mechanisms to integrate image and text features, while not effectively utilizing the highly condensed textual semantic information in improving the fused features, resulting in inaccurate lesion segmentation. This paper introduces a novel approach, the Medical Image-Text Segment Anything Model (MIT-SAM), for text-assisted medical image segmentation. Specifically, we introduce the SAM-enhanced image encoder and a Bert-based text encoder to extract heterogeneous features. To better leverage the highly condensed textual semantic information for heterogeneous feature fusion, such as crucial details like position and quantity, we propose the image-text interactive fusion (ITIF) block and self-supervised text reconstruction (SSTR) method. The ITIF block facilitates the mutual enhancement of homogeneous information among heterogeneous features and the SSTR method empowers the model to capture crucial details concerning lesion text, including location, quantity, and other key aspects. Experimental results demonstrate that our proposed model achieves state-of-the-art performance on the QaTa-COV19 and MosMedData+ datasets. Xichuan Zhou, Lingfeng Yan, Rui Ding 0009, Chukwuemeka Clinton Atabansi, Jing Nie 0001, Lihui Chen 0002, Haijun Liu 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | MeSAM: Multiscale Enhanced Segment Anything Model for Optical Remote Sensing ImagesabstractSegment anything model (SAM) has been widely applied to various downstream tasks for its excellent performance and generalization capability. However, SAM exhibits three limitations related to remote sensing semantic segmentation task: 1) the image encoders excessively lose high-frequency information, such as object boundaries and textures, resulting in rough segmentation masks; 2) due to being trained on natural images, SAM faces difficulty in accurately recognizing objects with large-scale variations and uneven distribution in remote sensing images; 3) the output tokens used for mask prediction are trained on natural images and not applicable to remote sensing image segmentation. In this paper, we explore an efficient paradigm for applying SAM to the semantic segmentation of remote sensing images. Furthermore, we propose MeSAM, a new SAM fine-tuning method more suitable for remote sensing images to adapt it to semantic segmentation tasks. Our method first introduces an inception mixer into the image encoder to effectively preserve high-frequency features. Secondly, by designing a mask decoder with remote-sensing correction and incorporating multiscale connections, we make up the difference in SAM from natural images to remote sensing images. Experimental results demonstrated that our method significantly improves the segmentation accuracy of SAM for remote sensing images, outperforming some state-of-the-art methods. The code will be available at https://github.com/Magic-lem/MeSAM. Xichuan Zhou, Fu Liang, Lihui Chen 0002, Haijun Liu 0001, Gemine Vivone, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Spectral-Spatial Transformer for Hyperspectral Image SharpeningabstractConvolutional neural networks (CNNs) have recently achieved outstanding performance for hyperspectral (HS) and multispectral (MS) image fusion. However, CNNs cannot explore the long-range dependence for HS and MS image fusion because of their local receptive fields. To overcome this limitation, a transformer is proposed to leverage the long-range dependence from the network inputs. Because of the ability of long-range modeling, the transformer overcomes the sole CNN on many tasks, whereas its use for HS and MS image fusion is still unexplored. In this article, we propose a spectral-spatial transformer (SST) to show the potentiality of transformers for HS and MS image fusion. We devise first two branches to extract spectral and spatial features in the HS and MS images by SST blocks, which can explore the spectral and spatial long-range dependence, respectively. Afterward, spectral and spatial features are fused feeding the result back to spectral and spatial branches for information interaction. Finally, the high-resolution (HR) HS image is reconstructed by dense links from all the fused features to make full use of them. The experimental analysis demonstrates the high performance of the proposed approach compared with some state-of-the-art (SOTA) methods. Lihui Chen 0002, Gemine Vivone, Jiayi Qin, Jocelyn Chanussot, Xiaomin Yang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Spatial-temporal feature refine network for single image super-resolution
Jiayi Qin, Lihui Chen 0002, Kai Liu 0012, Gwanggil Jeon, Xiaomin Yang |
Appl. Intell. | 2 |
| 2023 | Spatial Data Augmentation: Improving the Generalization of Neural Networks for PansharpeningabstractDeep learning (DL) methods have achieved impressive performance for pansharpening in recent years. However, because of poor generalization, most DL methods achieve unsatisfactory performance for data acquired by sensors not considered during the training phase and decreased performance for samples at full resolution. To solve this issue, we propose a data augmentation framework for pansharpening neural networks. Specifically, we introduce first a random spatial degradation based on anisotropic Gaussian-shaped modulation transfer functions (MTFs) to increase the generalization with respect to different spatial models and sensors. Then, considering that various sensors have different ground sampling distances (GSDs), we randomly rescale the GSD of the training samples to improve the generalization with respect to spatial resolution. Thanks to this module, the generalization to tests from different sensors and samples at full resolution can easily be achieved. Experimental results demonstrate the effectiveness of the proposed approach with better performance when data for training are decoupled with the ones for testing and comparable performance when training and testing are coupled (i.e., data acquired by the same sensor are considered in the two phases). Besides, performance at full resolution for pansharpening neural networks is improved by the proposed approach. The proposed approach has been integrated into existing pansharpening neural networks showing satisfactory performance for widely used sensors, including, GaoFen-1, QuickBird, WorldView-2, WorldView-3, IKONOS, Spot-7, GeoEye, and PHR1A. Lihui Chen 0002, Gemine Vivone, Zihao Nie, Jocelyn Chanussot, Xiaomin Yang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Efficient Hyperspectral Sparse Regression Unmixing With MultilayersabstractThe sparse regression method is known for its ability to unmix hyperspectral data, but it can be computationally expensive and accurately insufficient due to the large scale and high coherence of the spectral library. To address this issue, a new approach called layered sparse regression unmixing (termed LSU) has been proposed in this paper. This method involves breaking down the sparse unmixing process into multilayers, each of which interactively learns a row-sparsity-promoting abundance matrix and fine-tunes active library atoms based on measured activeness. By doing so, LSU outputs both a learned abundance matrix and an optimal library that can best model each mixed pixel in the scene. The proposed LSU can be efficiently solved by the alternating direction method of the multipliers framework. Experimental results obtained from simulated and real hyperspectral images demonstrate the effectiveness of LSU. The demo of the proposed LSU will be publicly available at https://github.com/XiangfeiShen/Layered_Sparse_Regression_Unmixing. Xiangfei Shen, Lihui Chen 0002, Haijun Liu 0001, Xi Su, Wenjia Wei, Xichuan Zhou |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Progressive Interaction-Learning Network for Lightweight Single-Image Super-Resolution in Industrial ApplicationsabstractRecently, deep learning (DL)-based industrial applications have attracted broad attention due to their advanced performance. However, the limited computational resource in portable devices always makes big DL models inapplicable in the industry. DL-based single-image super-resolution also encounters this problem because of its large computations. Besides, most lightweight convolutional-neural-network-based methods utilize features insufficiently, which restricts their capability for industrial reconstruction. To alleviate this problem, we present a progress interaction-learning network (PILN) to refine features at different levels: at the global level, we employ a progressive interaction-learning strategy to integrate hierarchical features in temporal and spatial dimensions; at the mediate level, enhanced interaction-learning units, adopting the enhanced interactive study, significantly boost the reconstruction performance; at the local level, employing pixelwise learning, residual cells are raised to search for an optimal information flow by weight distribution. Extensive experiments demonstrate that the PILN outperforms other state-of-the-art methods. Jiayi Qin, Lihui Chen 0002, Seunggil Jeon, Xiaomin Yang |
IEEE Trans. Ind. Informatics | 2 |
| 2022 | Spectral-Spatial Transformer for Hyperspectral Image SharpeningabstractConvolutional neural networks (CNNs) have achieved impressive performance for hyperspectral (HS) and multispectral (MS) image fusion in recent years. They extract features by local filters, which is limited to explore long-range dependency in input images. However, long-range dependence is an import cue for HS and MS image fusion, as it contributes to exploration of spatial self-similarity and spectral dependence. To take advantage of long-range dependence, we propose a spectral-spatial transformer (SST) for MS and HS image fusion. The experimental results demonstrate the high performance of the proposed approach compared to some state-of-the-art methods. Lihui Chen 0002, Gemine Vivone, Jiayi Qin, Jocelyn Chanussot, Xiaomin Yang |
IGARSS | 1 |
| 2022 | Medical image super-resolution with laplacian dense network
Lihui Chen 0002, Rongzhu Zhang, Awais Ahmad 0001, Marcelo Keese Albertini, Xiaomin Yang |
Multim. Tools Appl. | 2 |
| 2022 | Aerial image super-resolution based on deep recursive dense network for disaster area surveillance
Feiqiang Liu, Lihui Chen 0002, Gwanggil Jeon, Marcelo Keese Albertini, Xiaomin Yang |
Pers. Ubiquitous Comput. | 3 |
| 2022 | ArbRPN: A Bidirectional Recurrent Pansharpening Network for Multispectral Images With Arbitrary Numbers of BandsabstractAlthough the performance of pansharpening has been significantly improved by advanced deep-learning (DL) technologies in recent years, most DL-based methods fail to process multispectral (MS) images with arbitrary numbers of bands by a single model. Consequently, it is inevitable to train separate models for MS images with different numbers of bands, which is time- and storage-consuming as well as inefficient in practice. To tackle the above problem, we propose a bidirectional recurrent pansharpening network (named ArbRPN) for MS images with arbitrary numbers of bands. Our ArbRPN can dynamically reconstruct high-resolution (HR) MS images with different numbers of bands by adaptively changing the number of recurrence to the number of bands of the low-resolution (LR) MS images. Leveraging on the ability of the ArbRPN to process MS images with any number of bands, one can even customize the bands to be pansharpened. Moreover, to achieve superior performance, spectral discrepancy and dependence are considered in the ArbRPN. Details from the panchromatic (PAN) image are adaptively injected into the fused product according to the captured spectral dependence. Furthermore, training strategies of existing DL-based pansharpening methods can only group MS images with a constant number of bands into mini-batches. Therefore, we present a mask-based training method (called mask-training) to solve this problem. Benefiting from the mask-training, our ArbRPN can achieve superior performance and robustness during pansharpening. Extensive experiments show the superior performance of our ArbRPN with respect to the state-of-the-art (SOTA) methods applied to MS images with different numbers of bands. The code of our ArbRPN is available onhttps://github.com/Lihui-Chen/ArbRPN.git. Lihui Chen 0002, Zhibing Lai, Gemine Vivone, Gwanggil Jeon, Jocelyn Chanussot, Xiaomin Yang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | A trusted medical image super-resolution method based on feedback adaptive weighted dense network
Lihui Chen 0002, Xiaomin Yang, Gwanggil Jeon, Marco Anisetti, Kai Liu 0012 |
Artif. Intell. Medicine | 1 |
| 2020 | Medical image fusion method by using Laplacian pyramid and convolutional sparse representationabstractSummary Medical image fusion is a technology of combining multi‐modal images to generate a composite image, which is favorable to improve the capability of doctors in diagnosis and treatment of the disease. In order to achieve good performance, a fusion method by combining Laplacian pyramid (LP) and convolutional sparse representation (CSR) is proposed. In the proposed fusion method, LP transform is performed on each pair of pre‐registered computed tomography image and magnetic resonance image to obtain their detail layers and base layer. Then, the base layer is fused with a CSR‐based approach, whereas the detail layers are merged using the popular “max‐absolute” rule. Finally, the fused image is reconstructed by performing the inverse LP transform over the fused base layer and detail layers. The advantages of our method are that the texture detail information contained in source images can be fully extracted and the overall contrast of the final fused image will not be decreased. Experimental results demonstrate the superiority of the proposed method. Feiqiang Liu, Lihui Chen 0002, Lu Lu 0005, Awais Ahmad 0001, Gwanggil Jeon, Xiaomin Yang |
Concurr. Comput. Pract. Exp. | 2 |