EDBT 2026 Demo / reviewers in the wild / expert
Huan Liu 0015
dblp:92/309-15
· DBLP profile ↗
18ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0003-4642-4485ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LiteMFT: Lightweight Multi-Modal Fine-Tuning for Semantic SegmentationabstractMulti-modal image segmentation has recently attracted considerable attention due to its ability to integrate complementary information from diverse sensors, thereby enabling more accurate semantic predictions in complex or specialized scenarios. However, as data volume and model capacity continue to grow, many existing methods suffer substantial increases in parameters and computational costs, particularly with the widespread adoption of Vision Foundation Models (VFMs). To address these challenges, we introduce a Lightweight Multi-modal Fine-Tuning framework (LiteMFT) designed for efficient and generalizable adaptation of RGB-pretrained VFMs to multi-modal semantic segmentation. By incorporating only a small number of trainable parameters, LiteMFT enables effective extension of existing models to handle multi-modal image fusion tasks. The framework centers around two key components: the Modality Local Competition (MLC) module, which dynamically and efficiently fuses complementary features across modalities, and the Gated Low-Rank Adapter (GLR), which improves the backbone's adaptability to multi-modal data through content-aware low-rank transformation. Extensive experiments on both bi-modal and tri-modal segmentation tasks demonstrate that LiteMFT not only achieves competitive or superior performance but also exhibits strong scalability for additional modalities, underscoring its practicality and broad applicability in multi-modal semantic segmentation. Chengwang Guo, Yuxiang Zhang 0005, Mengmeng Zhang 0005, Huan Liu 0015, Wei Li 0032 |
IEEE Trans. Image Process. | 4 |
| 2026 | Data-Driven Band Optimization and Frequency-Aware Modeling in Medical Hyperspectral Image SegmentationabstractHyperspectral imaging delivers high-resolution spectral-spatial information to support molecular tissue characterization, but its clinical utility is far from being fully realized. Existing segmentation techniques are constrained by fixed or suboptimal band selection strategies and insufficient frequency-domain modeling, which limit their ability to fully exploit discriminative spectral cues and subtle tissue structures. To address these challenges, we propose AMBS-SF2Net, a unified framework that enhances spectral representation and hierarchical frequency modeling for accurate and efficient segmentation. Specifically, the Adaptive Mask-based Band Selection (AMBS) module dynamically identifies informative spectral channels, the Adaptive Spectral-Frequency Integration (ASFI) module fuses multi-scale spatial edges and frequency-aware spectral features, and the Multi-Axis Frequency Enhanced (MAFE) module captures complementary spectral and spatial frequency patterns along different tensor dimensions. To rigorously evaluate the method’s generalizability, we conduct extensive experiments on datasets spanning distinct imaging scales, comprising a microscopic cholangiocarcinoma pathology dataset and two macroscopic tissue datasets of pig abdominal organs and human placenta. Results demonstrate that AMBS-SF2Net significantly outperforms state-of-the-art methods, exhibiting superior robustness across varying spectral resolutions and spatial modalities, thereby validating its strong potential for diverse clinical applications. Wei Li 0032, Geng Qin, Huan Liu 0015, Xueyu Zhang, Haihao Zhang, Xiang-Gen Xia 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2025 | A Virtual Domain Collaborative Learning Framework for Semi-supervised Microscopic Hyperspectral Image Segmentation
Geng Qin, Huan Liu 0015, Wei Li 0032, Haihao Zhang, Yuxing Guo |
MICCAI (16) | 2 |
| 2025 | MGCD: Change Detection Focuses on Multiscale Style Domain GeneralizationabstractChange detection (CD) plays a crucial role in remote sensing (RS) applications. Although deep learning (DL)-based CD methods have achieved impressive performance, they typically rely on large amounts of well-annotated data, which is often scarce in real-world scenarios. This data scarcity leads to overfitting and limited generalization ability in conventional CD models. To address these challenges, this article proposes a novel few-sample CD framework named multiscale style DG method for change detection (MGCD). The core idea is to enhance the model’s ability to learn domain-invariant features by increasing data diversity across multiple style scales. Specifically, MGCD introduces global style diversity by incorporating out-of-domain natural images, and enriches local structural styles through unsupervised clustering and randomization. Additionally, a semantic consistency supervision (SCS) strategy is designed to guide multitemporal feature learning, enabling the network to better capture changes across diverse styles and scales. Extensive experiments conducted on three benchmark datasets demonstrate the effectiveness and robustness of the proposed MGCD framework in few-sample CD tasks. Chengwang Guo, Mengmeng Zhang 0005, Yuxiang Zhang 0005, Huan Liu 0015, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | LHAS: A Lightweight Network Based on Hierarchical Attention for Hyperspectral Image SegmentationabstractDeep learning has garnered extensive attention in hyperspectral image (HSI) processing. However, its application in HSI semantic segmentation tasks has been relatively limited. Although segmentation methods can often interpret images up to two orders of magnitude faster than classification methods when interpreting images of the same scene, the segmentation task requires the training data to be fully labeled, i.e., each pixel has a corresponding label. Such data are scarce in HSI data. To address this problem, this article proposes a lightweight segmentation network based on a hierarchical attention segmentation network (LHAS), in which a generalized data augmentation (GDA) method is utilized to acquire relatively sufficient data for semantic segmentation. Specifically, the hierarchical attention module is designed to extract global and local information on HSI patches from different layers. A prototype auxiliary module (PAM) of cluster contrast has also been developed to enhance feature discrimination. Across two different datasets in various scenarios, the proposed LHAS demonstrates superior segmentation performance compared to existing methods, affirming its effectiveness. In addition, experiments conducted on embedded devices validate the efficacy of LHAS. Lujie Song, Yunhao Gao, Yuanyuan Gui, Daguang Jiang, Mengmeng Zhang 0005, Huan Liu 0015, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Attention Multiscale Network for Semantic Segmentation of Multimodal Remote Sensing ImagesabstractDue to recent advancements in deep learning, techniques for urban structure extraction and semantic segmentation of multimodal remote sensing images have significant improvements. However, the challenge arises from the variable color intensity and complex texture of urban structures in optical images, particularly in buildings and roads. Fortunately, the light detection and ranging (LiDAR) images promote the task of developing an optimal multimodal fusion network that effectively leverages information from different modalities. In this article, we propose an attention multiscale network (AMSNet) for binary semantic segmentation tasks focused on building extraction, as well as multiclass semantic segmentation tasks, by integrating optical and LiDAR remote sensing images. AMSNet introduces two feature fusion modules—spatial scale adaptive fusion (S2AF) and semantic guided fusion (SGF). S2AF facilitates feature fusion between optical and LiDAR images within the same layer. This module contains a spatial scale selection strategy and an adaptive weight learning strategy, which enables the network to adaptively extract and intentionally select multiscale features from multimodal data. SGF addresses the semantic gap between different layered block features through semantic feature guidance strategy while achieving feature fusion. Furthermore, we introduce robust feature learning (RFL) to ensure the network robustness in rotation and variation in objects, making it resilient to images captured from different viewpoints and sensors. RFL incorporates point-to-point similarity learning strategy and multiscale feature reuse strategy. Experimental results on publicly available datasets demonstrate that AMSNet outperforms other state-of-the-art models. Extensive ablation studies further confirm the significance of all key components in the proposed approach. The source code of this method is available athttps://github.com/B-LG-J/AMSNet.git. Zhen Ye 0007, Yuan Li 0037, Zhen Li 0063, Huan Liu 0015, Yuxiang Zhang 0005, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | UFPF: A Universal Feature Perception Framework for Microscopic Hyperspectral ImagesabstractIn recent years, deep learning has shown immense promise in advancing medical hyperspectral imaging diagnostics at the microscopic level. Despite this progress, most existing research models remain constrained to single-task or single-scene applications, lacking robust collaborative interpretation of microscopic hyperspectral features and spatial information, thereby failing to fully explore the clinical value of hyperspectral data. In this paper, we propose a microscopic hyperspectral universal feature perception framework (UFPF), which extracts high-quality spatial-spectral features of hyperspectral data, providing a robust feature foundation for downstream tasks. Specifically, this innovative framework captures different sequential spatial nearest-neighbor relationships through a hierarchical corner-to-center mamba structure. It incorporates the concept of "progressive focus towards the center", starting by emphasizing edge information and gradually refining attention from the edges towards the center. This approach effectively integrates richer spatial-spectral information, boosting the model's feature extraction capability. On this basis, a dual-path spatial-spectral joint perception module is developed to achieve the complementarity of spatial and spectral information and fully explore the potential patterns in the data. In addition, a Mamba-attention Mix-alignment is designed to enhance the optimized alignment of deep semantic features. The experimental results on multiple datasets have shown that this framework significantly improves classification and segmentation performance, supporting the clinical application of medical hyperspectral data. The code is available at: https://github.com/Qugeryolo/UFPF. Geng Qin, Huan Liu 0015, Wei Li 0032, Xueyu Zhang, Yuxing Guo, Xiang-Gen Xia 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | Semi-Supervised Medical Hyperspectral Image Segmentation Using Adversarial Consistency Constraint Learning and Cross Indication NetworkabstractHyperspectral imaging technology is considered a new paradigm for high-precision pathological image segmentation due to its ability to obtain spatial and spectral information of the detected object simultaneously. However, due to the time-consuming and laborious manual annotation, precise annotation of medical hyperspectral images is difficult to obtain. Therefore, there is an urgent need for a semi-supervised learning framework that can fully utilize unlabeled data for medical hyperspectral image segmentation. In this work, we propose an adversarial consistency constraint learning cross indication network (ACCL-CINet), which achieves accurate pathological image segmentation through adversarial consistency constraint learning training strategies. The ACCL-CINet comprises a contextual and structural encoder to form the spatial-spectral feature encoding part. The contextual and structural indications are aggregated into features through a cross indication attention module and finally decoded by a pixel decoder to generate prediction results. For the semi-supervised training strategy, a pixel perceptual consistency module encourages the two models to generate consistent and low-entropy predictions. Secondly, a pixel maximum neighborhood probability adversarial constraint strategy is designed, which produces high-quality pseudo labels for cross supervision training. The proposed ACCL-CINet has been rigorously evaluated on both public and private datasets, with experimental results demonstrating that it outperforms state-of-the-art semi-supervised methods. The code is available at: https://github.com/Qugeryolo/ACCL-CINet. Geng Qin, Huan Liu 0015, Xueyu Zhang, Wei Li 0032, Yuxing Guo, Chuanbin Guo |
IEEE Trans. Image Process. | 2 |
| 2024 | Clusterformer for Pine Tree Disease Identification Based on UAV Remote Sensing Image SegmentationabstractPine wilt disease (PWD) is one of the most prevalent pine trees diseases, resulting in both ecological and economic havoc. UAV remote sensing segmentation plays a crucial role in early identifying and preventing PWD. However, deep learning segmentation models customized for PWD identification in scenarios with complex backgrounds have not received extensive exploration. In this paper, we propose a novel UAV remote sensing segmentation model called Clusterformer with a conventional encoder-decoder structure. The encoder is comprised of the specially designed Cluster Transformer, which includes a cluster token mixer and a spatial-channel feed-forward network (SC-FFN). The cluster token mixer utilizes constructed clusters from the feature maps to represent pixels, thereby reducing redundant and interfering information. The SC-FFN extracts multi-scale spatial information through depth-wise convolutions and channel information through a multilayer perceptron in sequence. The decoder primarily consists of the specially designed D-Cluster Transformer. The token mixer of the D-Cluster Transformer employs constructed clusters from high-level decoded tokens to represent low-level encoded tokens without relying on traditional upsampling methods such as interpolation, transpose convolution, or patch expansion. Consequently, more robust and less redundant features from high-level decoded feature maps are transferred to low-level encoded feature maps. Experimental results on two PWD datasets demonstrate that Clusterformer outperforms existing state-of-the-art segmentation models. This confirms the effectiveness and efficiency of Clusterformer in PWD identification. Code is available at https://github.com/huanliu233/Clusterformer. Huan Liu 0015, Wei Li 0032, Wen Jia, Mengmeng Zhang 0005, Lujie Song, Yuanyuan Gui |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Self-Supervised Learning With Multiscale Densely Connected Network for Hyperspectral Image ClassificationabstractIn recent years, deep learning-based methods have exhibited remarkable performance in the field of hyperspectral image (HSI) classification. However, conventional supervised methods heavily rely on a substantial number of labeled samples. Self-supervised learning, as a prominent unsupervised representation learning technique, offers the potential to extract valuable information from unlabeled data. In this article, we introduce a novel unsupervised approach called self-supervised learning with the multiscale densely connected network (SS-MSDCNet) to make full use of unlabeled samples for HSI classification. First, a two-stream structure was designed to generate more positive pairs, which enables the contrast self-supervised training to learn more useful information from unlabeled data. Subsequently, a data augmentation technique based on spectral splitting was proposed to coordinate the two-stream structure of SS-MSDCNet, enhancing spectral information expression. The backbone of the proposed approach is the multiscale densely connected network (MSDCNet), which obtains input HSIs with various spatial scales by removing peripheral pixels from the original input and subsequently extracts multiscale spatial-spectral features using 3-D densely connected modules and 3-D spatial attention modules. The 3-D densely connected module effectively harnesses multiscale features extracted by various convolutional layers, while the 3-D spatial attention module enhances the network’s focus on features conducive to accurate classification. To validate the efficacy of our approach, we conducted extensive experiments using four distinct HSI datasets. The results unequivocally demonstrate that SS-MSDCNet outperforms several well-established supervised and unsupervised classification methods. Furthermore, we designed a transfer experiment to confirm SS-MSDCNet’s robust generalization capabilities. The code is available athttps://github.com/mrblank99/SS-MSDCNet. Zhen Ye 0007, Zhan Cao, Huan Liu 0015, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | GCCD: A Generative Cross-Domain Change Detection NetworkabstractChange detection (CD) in hyperspectral image (HSI) is of great importance in the remote sensing area. The HSI-CD method based on deep learning (DL) has shown significant progress in achieving precise detection performance. However, many existing methods overlook cross-domain challenges in the CD task. In addition, the scarcity of annotated samples makes the DL models prone to overfitting. To address these issues, a generative cross-domain CD (GCCD) network based on a domain generalization (DG) technique is proposed. GCCD consists of a generator and a discriminator. The generator, with a Morph encoder (ME) and a Semantic encoder (SE), preserves fundamental structural information while introducing randomization to style and content. The discriminator extracts change information for discrimination through dual-temporal images and their difference map. Supervised adversarial learning between the generator and discriminator enhances the model’s ability to extract domain-invariant information. Extensive experiments on various datasets demonstrate the superior performance of the proposed method. Mengmeng Zhang 0005, Chengwang Guo, Yuxiang Zhang 0005, Huan Liu 0015, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | SegHSI: Semantic Segmentation of Hyperspectral Images With Limited Labeled PixelsabstractHyperspectral images (HSIs), with hundreds of narrow spectral bands, are increasingly used for ground object classification in remote sensing. However, many HSI classification models operate pixel-by-pixel, limiting the utilization of spatial information and resulting in increased inference time for the whole image. This paper proposes SegHSI, an effective and efficient end-to-end HSI segmentation model, alongside a novel training strategy. SegHSI adopts a head-free structure with cluster attention modules and spatial-aware feedforward networks (SA-FFN) for multiscale spatial encoding. Cluster attention encodes pixels through constructed clusters within the HSI, while SA-FFN integrates depth-wise convolution to enhance spatial context. Our training strategy utilizes a student-teacher model framework that combines labeled pixel class information with consistency learning on unlabeled pixels. Experiments on three public HSI datasets demonstrate that SegHSI not only surpasses other state-of-the-art models in segmentation accuracy but also achieves inference time at the scale of seconds, even reaching sub-second speeds for full-image classification. Code is available at https://github.com/huanliu233/SegHSI. Huan Liu 0015, Wei Li 0032, Xiang-Gen Xia 0001, Mengmeng Zhang 0005, Zhengqi Guo, Lujie Song |
IEEE Trans. Image Process. | 1 |
| 2023 | Multiarea Target Attention for Hyperspectral Image ClassificationabstractIn hyperspectral image (HSI) classification, objects corresponding to pixels of different classes exhibit varying size characteristics, which causes a challenge for effective pixelwise feature extraction and classification. In this article, we propose a novel multiscale model, called multiarea target attention (MATA). The proposed MATA uses an architecture that includes a shared feature extractor (FE) and classifier to capture multiscale spectral–spatial information effectively and efficiently. The FE uses a multiscale target attention module (MSTAM) to extract spectral–spatial information from target pixels and their similar pixels across multiscale areas, while$L_{2}$-normalization is used to address discrepancies between features of different scales. The classifier adopts a classwise decision weighting strategy to account for the varying sizes of different classes and the different contributions of semantic features at each scale to each class. Experimental results on five public HSI datasets demonstrate that the proposed MATA outperforms existing state-of-the-art single- and multiscale models, confirming its effectiveness and efficiency in HSI classification. Code is available athttps://github.com/huanliu233/MATA. Huan Liu 0015, Wei Li 0032, Xiang-Gen Xia 0001, Mengmeng Zhang 0005, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Mask-Reconstruction-Based Decoupled Convolution Network for Hyperspectral Imagery ClassificationabstractDeep learning has attracted much attention in hyperspectral image(HSI) classification. However, most deep learning methods ignore the information loss during spatial-spectral feature extraction, which potentially affects the classification performance. In this article, Mask Reconstruction-Based Decoupled Convolution Network(MrDCN) is proposed, which including the decoupled feature extraction module (DFEM) to extract spectral information and spatial information of target HSI patch respectively. The reconstruction modules are designed to maintain the feature extraction ability of DFEM and ensure that discriminative information in high-dimensional features and low-dimensional features is preserved. MrDCN outperforms state-of-the-art methods in classification on three datasets of various scenarios, which indicates its effectiveness, and experiments on embedded devices are executed to affirm the efficiency of MrDCN. Lujie Song, Mengmeng Zhang 0005, Wei Li 0032, Daguang Jiang, Huan Liu 0015, Yuxiang Zhang 0005 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Adaptive Domain-Adversarial Few-Shot Learning for Cross-Domain Hyperspectral Image ClassificationabstractThe process of annotating hyperspectral image (HSI) data is characterized by its time-consuming and labor-intensive nature. To address this challenge, researchers often employ a meta-learning paradigm known as few-shot learning (FSL), which leverages source domains containing a substantial number of labeled samples to assist in the classification of target domains with limited labeled samples. Many existing FSL methods rely on a conditional domain-adversarial strategy to mitigate the domain shift between source and target domains. However, these methods overlook the fact that the degrees of conditional distribution discrepancies between the two domains can vary significantly across different classes, leading to suboptimal conditional distribution alignment. To address this problem, we propose a framework called Adaptive Domain-Adversarial Few-Shot Learning (ADAFSL). Overall, the proposed ADAFSL employs an adaptive strategy that assigns varying weights to the conditional adversarial losses for different classes based on their respective degrees of discrepancies, thereby achieving global conditional distribution alignment. Specifically, a local alignment score map is constructed by measuring the similarity between labeled and unlabeled samples using both Euclidean and class-covariance metrics. This map is then multiplied with the conditional adversarial loss map, thus allocating more emphasis to the classes exhibiting greater discrepancies between the two domains. Moreover, to enhance cross-domain FSL, we design a multi-scale spectral-spatial feature extraction (MSFE) module, which incorporates cascaded multi-scale dilated convolutions. Experimental results on four public HSI datasets demonstrate that the proposed ADAFSL outperforms other state-of-the-art methods. The source code of this method can be found at https://github.com/JieW-ww/ADAFSL. Zhen Ye 0007, Jie Wang 0135, Huan Liu 0015, Yu Zhang 0200, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Morphological Transformation and Spatial-Logical Aggregation for Tree Species Classification Using Hyperspectral ImageryabstractHyperspectral image (HSI) consists of abundant spectral and spatial characteristics, which contribute to a more accurate identification of materials and land covers. However, most existing methods of hyperspectral image analysis primarily focus on spectral knowledge or coarse-grained spatial information while neglecting the fine-grained morphological structures. In the classification task of complex objects, spatial morphological differences can help to search for the boundary of fine-grained classes, e.g., forestry tree species. Focusing on subtle traits extraction, a spatial-logical aggregation network (SLA-NET) is proposed with morphological transformation for tree species classification. The morphological operators are effectively embedded with the trainable structuring elements, which contributes to distinctive morphological representations. We evaluate the classification performance of the proposed method on two tree species datasets, and the results demonstrate that the proposed SLA-NET significantly outperforms the other state-of-the-art classifiers. Mengmeng Zhang 0005, Wei Li 0032, Xudong Zhao 0003, Huan Liu 0015, Ran Tao 0003, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Central Attention Network for Hyperspectral Imagery ClassificationabstractIn this article, the intrinsic properties of hyperspectral imagery (HSI) are analyzed, and two principles for spectral-spatial feature extraction of HSI are built, including the foundation of pixel-level HSI classification and the definition of spatial information. Based on the two principles, scaled dot-product central attention (SDPCA) tailored for HSI is designed to extract spectral-spatial information from a central pixel (i.e., a query pixel to be classified) and pixels that are similar to the central pixel on an HSI patch. Then, employed with the HSI-tailored SDPCA module, a central attention network (CAN) is proposed by combining HSI-tailored dense connections of the features of the hidden layers and the spectral information of the query pixel. MiniCAN as a simplified version of CAN is also investigated. Superior classification performance of CAN and miniCAN on three datasets of different scenarios demonstrates their effectiveness and benefits compared with state-of-the-art methods. Huan Liu 0015, Wei Li 0032, Xiang-Gen Xia 0001, Mengmeng Zhang 0005, Chenzhong Gao, Ran Tao 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | Variation of a signal in Schwarzschild spacetime
Huan Liu 0015, Xiang-Gen Xia 0001, Ran Tao 0003 |
Sci. China Inf. Sci. | 1 |