EDBT 2026 Demo / reviewers in the wild / expert
Xianju Li
dblp:196/9514
· DBLP profile ↗
21ranked-venue papers
2as first author
19since 2021 · last 2026
0000-0001-7785-2541ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 9 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards knowledge-infused seabed sediment mapping: A semi-supervised framework integrating large language models and knowledge graphs for multibeam data
Haoyi Wang, Weitao Chen 0001, Xianju Li, Gaodian Zhou, Qianyong Liang, Jun Li 0009, Mercedes Eugenia Paoletti, Juan Mario Haut |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Dual-Stream Global-Local Feature Collaborative Representation Network for Scene Classification of Mining AreaabstractThe scene classification of mining areas provides accurate foundational data to support geological environment monitoring and resource development planning. This study fuses multi-source data to construct a multi-modal mine land cover scene classification dataset. A significant challenge in mining area classification lies in the complex spatial layout and multi-scale characteristics of these regions. By extracting global and local features, it becomes possible to comprehensively reflect the spatial distribution and overall arrangement of different landforms, thereby enabling a more accurate capture of the holistic characteristics of mining scenes. We propose a dual-branch fusion model utilizing collaborative representation to decompose global features into a set of key semantic vectors. This model comprises three key components: (1) Multi-scale Global Transformer Branch: This branch leverages adjacent large-scale features to generate global channel attention features for small-scale features, effectively capturing the multi-scale feature relationships inherent in mining areas. (2) Local Enhancement Collaborative Representation Branch: This branch refines the attention weights by leveraging local features and reconstructed key semantic sets, ensuring that the local context and detailed characteristics of the mining area are effectively integrated. This enhances the model’s sensitivity to fine-grained spatial variations within the mining environment. (3) Dual-Branch Deep Feature Fusion Module: This module fuses the complementary features of the two branches to incorporate more scene information. This fusion strengthens the model’s ability to distinguish and classify complex mining landscapes. Finally, this study employs multi-loss computation to ensure a balanced integration of the modules. The overall accuracy of this model is 83.63%, which outperforms other comparative models. Additionally, it achieves the best performance across all other evaluation metrics. The experimental results demonstrate the effectiveness of the proposed dataset and model for classifying mining areas. Shuqi Fan, Haoyi Wang, Xianju Li |
IJCNN | 3 |
| 2025 | Dual Feature Enhancement and Adaptive Attention Fusion for Cross-Modal Scene Classification of Mining LandabstractMining area scene classification is crucial for deposit evaluation and environmental monitoring. However, existing methods struggle with homogeneous and heterogeneous spectral spatial and topographic features of mining areas, large intra-class variations, and small target sizes. To overcome these limitations, this study integrates RGB and SAR data to construct a multi-modal dataset and proposes an RGB-SAR mining scene classification model with dual feature enhancement and adaptive cross-modal attention interaction. The model includes: (1) Dual feature enhancement module that suppresses irrelevant features and enhances discriminative multi-scale representations of mining targets; (2) BifocalNet based feature extraction module using a CNN-Transformer hybrid architecture to capture local textures and model global context; (3) Attention based adaptive cross-modal interaction module that achieves deep spectral geometric feature complementarity through the fusion of RGB and SAR modalities. Experiments show the model achieves an OA of 84.58%, outperforming other models and ranking first or second in most evaluation metrics. The proposed dataset and model thus advance mining scene classification. Jiangyuan Wang, Xianju Li |
SMC | 3 |
| 2025 | Diversity Learning Guided Dual Graph Autoencoder for Unsupervised Hyperspectral Band SelectionabstractHyperspectral band selection, aimed at identifying key spectral bands from the original image, is crucial for reducing dimensionality and enhancing computational efficiency in hyperspectral image (HSI) analysis. Graph learning-based methods have attracted considerable attention due to their efficiency in representing structural correlations between bands and their powerful capability to extract features. However, existing methods have limitations in utilizing spatial relationships among bands and learning their discriminative characteristics. To address these limitations, we propose a Diversity Learning Guided Dual Graph Autoencoder (DLG-DGAE) for unsupervised hyperspectral band selection. In our framework, we integrate a Dual Graph Autoencoder (DGAE) module designed to extract information from both the spatial and spectral relationships among bands, thus fully capturing the structural similarity of the bands. Additionally, we introduce a Spectral Diversity Learning (SDL) strategy to reduce redundant information in the latent representation and enhance the discriminative properties of each band. In the final step, we proceed to cluster the fused latent embeddings. Within each cluster, we select the band exhibiting the highest information entropy as the representative band. Through extensive experimentation on three publicly available datasets, our results consistently indicate that the proposed method surpasses other state-of-the-art techniques. The code is available athttps://github.com/fengwe1/DLG-DGAE. Chang Tang, Xinwang Liu 0002, Junjun Jiang, Xianju Li, Xinzhong Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Pixel-Superpixel Contrastive Learning and Pseudo-Label Correction for Hyperspectral Image ClusteringabstractHyperspectral image (HSI) clustering is gaining considerable attention owing to recent methods that overcome the inefficiency and misleading results from the absence of supervised information. Contrastive learning methods excel at existing pixel-level and superpixel-level HSI clustering tasks. The pixel-level contrastive learning method can effectively improve the ability of the model to capture fine features of HSI but requires a large time overhead. The superpixel-level contrastive learning method utilizes the homogeneity of HSI and reduces computing resources; however, it yields rough classification results. To exploit the strengths of both methods, we present a pixel–superpixel contrastive learning and pseudo-label correction (PSCPC) method for the HSI clustering. PSCPC can reasonably capture domain-specific and fine-grained features through superpixels and the comparative learning of a small number of pixels within the superpixels. To improve the clustering performance of superpixels, this paper proposes a pseudo-label correction module that aligns the clustering pseudo-labels of pixels and superpixels. In addition, pixel-level clustering results are used to supervise superpixel-level clustering, improving the generalization ability of the model. Extensive experiments demonstrate the effectiveness and efficiency of PSCPC. Renxiang Guan, Xianju Li, Chang Tang |
ICASSP | 3 |
| 2024 | Superpixel-Based Dual-Neighborhood Contrastive Graph Autoencoder for Deep Subspace Clustering of Hyperspectral Image
Renxiang Guan, Yaowen Hu, Xianju Li |
ICIC (6) | 8 |
| 2024 | S2RC-GCN: A Spatial-Spectral Reliable Contrastive Graph Convolutional Network for Complex Land Cover Classification Using Hyperspectral ImagesabstractSpatial correlations between different ground objects are an important feature of mining land cover research. Graph Convolutional Networks (GCNs) can effectively capture such spatial feature representations and have demonstrated promising results in performing hyperspectral imagery (HSI) classification tasks of complex land. However, the existing GCN-based HSI classification methods are prone to interference from redundant information when extracting complex features. To classify complex scenes more effectively, this study proposes a novel spatial-spectral reliable contrastive graph convolutional classification framework named S2RC-GCN. Specifically, we fused the spectral and spatial features extracted by the 1D- and 2D-encoder, and the 2D-encoder includes an attention model to automatically extract important information. We then leveraged the fused high-level features to construct graphs and fed the resulting graphs into the GCNs to determine more effective graph representations. Furthermore, a novel reliable contrastive graph convolution was proposed for reliable contrastive learning to learn and fuse robust features. Finally, to test the performance of the model on complex object classification, we used imagery taken by Gaofen-5 in the Jiang Xia and Xin Jiang area to construct complex land cover datasets. The test results show that compared with other models, our model achieved the best results and effectively improved the classification performance of complex remote sensing imagery. Renxiang Guan, Chujia Song, Xianju Li, Ruyi Feng |
IJCNN | 5 |
| 2024 | Fusion of Attention-Based Cascaded CNN and Label Dependency-Based GCN for Multi-label Scene Classification of Mining LandabstractMulti-label scene classification (MLSC) of mining land (ML), that used to investigate whether there are some MLs, is of great significance for mine environmental monitoring and sustainable development. The ML’s characteristics of homogeneity and heterogeneity of spectral-spatial and topographic feature, large scale difference, and complexity of spatial co-occurrence greatly limit the accuracy of MLSC. This study has constructed a multi-modal dataset for MLSC of ML, by incorporating multispectral, synthetic aperture radar, and topographic data. And a novel fusion model that integrates attention-based cascaded convolution neural network (CNN) with label dependency-based graph convolution network (GCN) was proposed. The model consists of three main components. (1) Attention enhanced multi-scale feature cascade fusion, employed to extract crucial multi-scale features and reduce feature redundancy. (2) Label dependency-based GCN, utilizing multiple layers of GCN to extract spatial dependencies of ML from the label co-occurrence probability matrix. (3) Multi-modal feature fusion and multi-label classification, integrating the image features extracted by CNN with the spatial co-occurrence features extracted by GCN to obtain multi-label classification results. The mAP of the proposed model is 68.91%, outperforming other comparative models. Most of the other evaluation metrics also rank as either optimal or suboptimal for the proposed model. In summary, the dataset and model proposed in this study are beneficial for MLSC of ML. Xianju Li, Wenxi He, Weitao Chen 0001 |
IJCNN | 1 |
| 2024 | Multi-level Graph Subspace Contrastive Learning for Hyperspectral Image ClusteringabstractHyperspectral image (HSI) clustering is a challenging task due to its high complexity. Despite subspace clustering shows impressive performance for HSI, traditional methods tend to ignore the global-local interaction in HSI data. In this study, we proposed a multi-level graph subspace contrastive learning (MLGSC) for HSI clustering. The model is divided into the following main parts. Graph convolution subspace construction: utilizing HSI’s spectral and texture feautures to construct two graph convolution views. Local-global graph representation: local graph representations were obtained by step-by-step convolutions and a more representative global graph representation was obtained using an attention-based pooling strategy. Multi-level graph subspace contrastive learning: multi-level contrastive learning was conducted to obtain local-global joint graph representations, to improve the consistency of the positive samples between views, and to obtain more robust graph embeddings. Specifically, graph-level contrastive learning is used to better learn global representations of HSI data. Node-level intra-view and inter-view contrastive learning is designed to learn joint representations of local regions of HSI. The proposed model is evaluated on four popular HSI datasets: Indian Pines, Pavia University, Houston, and Xu Zhou. The overall accuracies are 97.75%, 99.96%, 92.28%, and 95.73%, which significantly outperforms the current state-of-the-art clustering methods. Renxiang Guan, Kainan Gao, Xianju Li, Chang Tang |
IJCNN | 6 |
| 2024 | Spectral-Spatial Blockwise Masked Transformer With Contrastive Multi-View Learning for Hyperspectral Image Classification
Zhenhui Liu, Ziqing Xu, Haoyi Wang, Xianju Li, Jianyi Peng |
PRCV (4) | 5 |
| 2024 | Feature Exchange and Distribution-Based Mining Land Detection Method by Multispectral Imagery
Haoyi Wang, Xianju Li, Huijun Ding, Yiran Chang, Jianyi Peng |
PRCV (13) | 3 |
| 2024 | Hyperspectral band selection via region-wise latent feature fusion and graph filter embedded subspace clustering
Minhui Wang, Chang Tang, Weiying Xie, Xianju Li, Jiangfeng Xu |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | Contrastive Multiview Subspace Clustering of Hyperspectral Images Based on Graph Convolutional NetworksabstractHigh-dimensional and complex spectral structures make the clustering of hyperspectral images (HSI) a challenging task. Subspace clustering is an effective approach for addressing this problem. However, current subspace clustering algorithms are primarily designed for a single view and do not fully exploit the spatial or textural feature information in HSI. In this study, contrastive multi-view subspace clustering of HSI was proposed based on graph convolutional networks. Pixel neighbor textural and spatial-spectral information were sent to construct two graph convolutional subspaces to learn their affinity matrices. To maximize the interaction between different views, a contrastive learning algorithm was introduced to promote the consistency of positive samples and assist the model in extracting robust features. An attention-based fusion module was used to adaptively integrate these affinity matrices, constructing a more discriminative affinity matrix. The model was evaluated using four popular HSI datasets: Indian Pines, Pavia University, Houston, and Xu Zhou. It achieved overall accuracies of 97.61%, 96.69%, 87.21%, and 97.65%, respectively, and significantly outperformed state-of-the-art clustering methods. In conclusion, the proposed model effectively improves the clustering accuracy of HSI. Our implementation is available at https://github.com/GuanRX/CMSCGC. Renxiang Guan, Wenxuan Tu, Jun Wang 0118, Yue Liu 0008, Xianju Li, Chang Tang, Ruyi Feng |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | MS-Former: Memory-Supported Transformer for Weakly Supervised Change Detection With Patch-Level AnnotationsabstractFully supervised change detection methods have achieved significant advancements in performance, yet they depend severely on acquiring costly pixel-level labels. Considering that the patch-level annotations also contain abundant information corresponding to both changed and unchanged objects in bi-temporal images, an intuitive solution is to segment the changes with patch-level annotations. How to capture the semantic variations associated with the changed and unchanged regions from the patch-level annotations to obtain promising change results is the critical challenge for the weakly supervised change detection task. In this paper, we propose a memory-supported transformer (MS-Former), a novel framework consisting of a bi-directional attention block (BAB) and a patch-level supervision scheme (PSS) tailored for weakly supervised change detection with patch-level annotations. More specifically, the BAB captures contexts associated with the changed and unchanged regions from the temporal difference features to construct informative prototypes stored in the memory bank. On the other hand, the BAB extracts useful information from the prototypes as supplementary contexts to enhance the temporal difference features, thereby better distinguishing changed and unchanged regions. After that, the PSS guides the network learning valuable knowledge from the patch-level annotations, thus further elevating the performance. Experimental results on three benchmark datasets demonstrate the effectiveness of our proposed method in the change detection task. The demo code for our work will be publicly available at https://github.com/guanyuezhen/MS-Former. Zhenglai Li, Chang Tang, Xinwang Liu 0002, Changdong Li, Xianju Li, Wei Zhang 0049 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Spatial and Spectral Structure Preserved Self-Representation for Unsupervised Hyperspectral Band SelectionabstractAs an effective manner to reduce data redundancy and processing inconvenience, hyperspectral band selection aims to select a subset of informative and discriminative bands from the original data cube. Although a large number of approaches have been proposed and obtained great success, they still face at least two issues. Firstly, most of the previous methods only consider the redundancy between neighbor bands, while the global information has been ignored. Secondly, each band is often treated as a whole and reshaped to a feature vector without considering the spatial structure of different regions. In this paper, in order to address these issues, we propose a spatial and spectral structure preserved self-representation model for unsupervised hyperspectral band selection without using any label information, referred to as S4P briefly. Different from previous methods that stretch each band into a feature vector, the first principal component of the original hyperspectral cube is segmented into different superpixels, which can reflect the spatial structure of homogeneous regions. Then each band can be represented by a superpixel level feature vector and the self-representation model is utilized to learn the spectral correlation of different bands. In addition, an adaptive and weighted multiple graph fusion term is designed to generate a unified similarity graph between different superpixels, which is used to capture the spatial structure in the self-representation space. Finally, anl2,1-norm is imposed on the self-representation coefficient matrix to measure the band importance. We design an alternative update scheme to optimize the resultant problem, the self-representation coefficient matrix and the superpixel-wise similarity graph can boost each other during the updating process to obtain optimal results. Extensive experiments with detailed analysis of three public datasets are conducted to validate the superiority of the proposed S4P when compared with other state-of-the-art competitors. Chang Tang, Jun Wang 0118, Xinwang Liu 0002, Weiying Xie, Xianju Li, Xinzhong Zhu |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Deep Feature Enhancement Method for Land Cover With Irregular and Sparse Spatial Distribution Features: A Case Study on Open-Pit MiningabstractLand cover classification in mining areas (LCMA) is essential for the environmental assessment of mines and plays a crucial role in their sustainable development. The shapes of mine land occupation elements are irregular, and the overall proportion of their area is relatively small. Therefore, their features may be easily lost during feature extraction, which limits the interpretation accuracy in mining areas. This study attempts to address these issues. We propose a model named EG-UNet to enhance the features of elements with few samples and to capture long-range information. The proposed EG-UNet includes two main modules. First, the edge feature enhancement module, the edges of elements of mine land occupation contain more information than other spatial locations. Hence, during the feature extraction of elements, a Sobel operator is used to extract the object boundary, which increases the weight of these features before the pooling operation for their preservation. Second, the long-range information extraction module, long-range information helps extract tiny objects, such as dumping grounds in the mining area. We present a graph convolutional network (GCN) to capture the long-range features and apply convolutional neural networks to learn the graph construction. A total of ten deep-learning networks were compared using the LCMA semantic segmentation dataset. Our model exhibited the best performance, especially in classifying classes with few samples. Furthermore, to evaluate the general ability of EG-UNet, a benchmark-Gaofen Image Dataset (GID) was used, and the result still reflected the superiority of our method. Gaodian Zhou, Weitao Chen 0001, Xianju Li, Jun Li 0009, Lizhe Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | NIGAN: A Framework for Mountain Road Extraction Integrating Remote Sensing Road-Scene Neighborhood Probability Enhancements and Improved Conditional Generative Adversarial NetworkabstractMountain roads are a source of important basic geographic data used in various fields. The automatic extraction of road images through high-resolution remote sensing imagery using deep learning has attracted considerable attention. But the interference of context information limited extraction accuracy, especially for roads in mountain area. Furthermore, when pursuing research in a new district, many algorithms are difficult to train due to a lack of data. To address these issues, a framework based on remote sensing road-scene neighborhood probability enhancement and improved conditional generative adversarial network (NIGAN) is proposed in this article. This framework can be divided into two sections: 1) road scenes classification section. A remote sensing road-scene neighborhood confidence enhancement method was designed for classifying road scenes of the study area to reduce the impact of nonroad information on subsequent fine-road segmentation and 2) fine-road segmentation section. An improved dilated convolution module, which is helpful in extracting small objects such as road, was added into the conditional generative adversarial network (CGAN) to increase the receptive field and pay attention to global information, and segment roads from the results of road scenes classification section. To validate the NIGAN framework, new mountain road-scene and label datasets were constructed, and diverse comparison experiments were performed. The results indicate that the NIGAN framework can improve the integrity and accuracy of mountain road-scene extraction in diverse and complex conditions. The results further confirm the validity of the NIGAN framework in small samples. In addition, the mountain road-scene datasets can serve as benchmark datasets for studying mountain road extraction. Weitao Chen 0001, Gaodian Zhou, Zhuoyue Liu, Xianju Li, Xiongwei Zheng, Lizhe Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | A Fine-Grained Genetic Landform Classification Network Based on Multimodal Feature Extraction and Regional Geological ContextabstractDeep learning networks have facilitated the automated scene recognition of landforms based on geomorphogenesis. However, current genetic landform classification methods do not consider regional geological context, which can more accurately reflect the formation and evolution mechanism of geomorphic landforms than local ones. Therefore, this study proposes a multimodal, deep learning landform recognition framework based on a joint contextual geological and channel attention module (GCMENET). First, the multibranch feature extraction network of DenseNet121 is used to extract the respective features from the target scene and the contextual geological scene. Second, the features similar to the landform features of the target scene are extracted from the contextual geological features based on the cosine method and then combined with geomorphic features of the target scene. Third, channel attention mechanism is used to reduce the interference caused by redundant contextual geological information after fusion of data. To measure the classification accuracy of GCMENET, we establish a fine geomorphogenic dataset consisting of remote sensing images of six landform types with a$64\times64$-pixel size and 10-m resolution (JOS10m). During the training process of two geomorphogenic datasets, the feature extraction network without batchnorm2d (batch normalization) could preserve the distribution and spatial alignment of data from the components. Using different training-to-validation data ratios and combinations of input components, the results of the GCMENET supplemented with the joint contextual geological and channel attention module exhibited greater accuracy than those obtained without the module. This observation confirms the importance of contextual geological information in automated geomorphogenic landforms. Shubing Ouyang, Weitao Chen 0001, Yusen Dong, Xianju Li, Jun Li 0009 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Split Depth-Wise Separable Graph-Convolution Network for Road Extraction in Complex Environments From High-Resolution Remote-Sensing ImagesabstractRoad information from high-resolution remote-sensing images is widely used in various fields, and deep-learning-based methods have effectively shown high road-extraction performance. However, for the detection of roads sealed with tarmac, or covered by trees in high-resolution remote-sensing images, some challenges still limit the accuracy of extraction: 1) large intraclass differences between roads and unclear interclass differences between urban objects, especially roads and buildings; 2) roads occluded by trees, shadows, and buildings are difficult to extract; and 3) lack of high-precision remote-sensing datasets for roads. To increase the accuracy of road extraction from high-resolution remote-sensing images, we propose a split depth-wise (DW) separable graph convolutional network (SGCN). First, we split DW-separable convolution to obtain channel and spatial features, to enhance the expression ability of road features. Thereafter, we present a graph convolutional network to capture global contextual road information in channel and spatial features. The Sobel gradient operator is used to construct an adjacency matrix of the feature graph. A total of 13 deep-learning networks were used on the Massachusetts roads dataset and nine on our self-constructed mountain road dataset, for comparison with our proposed SGCN. Our model achieved a mean intersection over union (mIOU) of 81.65% with an F1-score of 78.99% for the Massachusetts roads dataset, and an mIOU of 62.45% with an F1-score of 45.06% for our proposed dataset. The visualization results showed that SGCN performs better in extracting covered and tiny roads and is able to effectively extract roads from high-resolution remote-sensing images. Gaodian Zhou, Weitao Chen 0001, Qianshan Gui, Xianju Li, Lizhe Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2018 | Rolling Guidance Based Scaled-Aware Spatial Sparse Unmixing for Hyperspectral Remote Sensing ImageryabstractSpatial regularization based sparse unmixing has been attracted much attention and has achieved improved fractional abundance results. However, the traditional approach to spatial consideration can only suppress discrete wrong unmixing points and smooth an abundance map with low-contrast changes, and it has no concept of scale difference. As the different levels of structures and edges in remote sensing have different meanings and importance, to better extract the different levels of spatial details, rolling guidance based scale-aware spatial sparse unmixing (RGSU), is proposed in this paper to extract and recover the different levels important structures and details in the hyperspectral remote sensing image unmixing procedure. Differing from the existing spatial regularization based sparse unmixing approaches, the proposed method considers the different levels of edges by combining a Gaussian filter-like method to realize small-scale structure removal with a joint bilateral filtering process to account for the spatial domain and range domain correlations. The experimental results obtained with both simulated and real hyperspectral images show that the proposed method achieves a better performance and produces more accurate abundance maps, as well as higher quantitative results, when compared to the current state-of-the-art sparse unmixing algorithms Ruyi Feng, Tian Tian 0007, Xianju Li, Kun Sun 0002 |
IGARSS | 3 |
| 2017 | Comparison and integration of feature reduction methods for land cover classification with RapidEye imagery
Xianju Li, Weitao Chen 0001, Xinwen Cheng, Yiwei Liao |
Multim. Tools Appl. | 1 |