Xichun Li

dblp:203/3830 · DBLP profile ↗
← Back
17ranked-venue papers
0as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Prototype similarity-constraint enhancement network: A few-Shot class-Incremental learning for hyperspectral image classification
Yifeng Tan, Lianhui Liang, Huafu Xu, Thomas Wu 0001, Xichun Li, Yuan Yan Tang
Expert Syst. Appl.7
2026 Unlocking the Potential of Auxiliary Captions via Dual-Branch Multi-Scale Network for Composed Image Retrieval
abstract
Composed image retrieval (CIR) aims to retrieve target images by combining a reference image with a modification text. Traditional CIR methods often struggle with feature-level multimodal fusion, leading to deviations from the original embedding space. To address this, we propose a Dual-Branch Multi-Scale Network (DMN) that integrates a combining branch and a complete text branch. To enhance the use of captions generated by advanced image captioning models for CIR, the DMN leverages an attribute-driven disentanglement layer to separate features into distinct latent factors and employs a dual-path multimodal fusion module for effective feature integration. Additionally, a multi-scale matching module incorporating both global and local matching strategies is introduced to enhance fine-grained feature discrimination. Experimental results on the FashionIQ, Shoes, and CIRR datasets demonstrate that our DMN model consistently outperforms state-of-the-art methods, achieving improvements of up to 1.43% in mean recall metrics.
Jinhong Xu, Xichun Li, Thomas Wu 0001, Yuan Yan Tang, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.3
2026 Similarity-Guided Denoising Reconstruction for Unsupervised Image Captioning
abstract
Image captioning aims to generate natural and accurate textual descriptions of given images. Although significant progress has been made in image captioning models in recent years, most existing approaches heavily rely on high-quality image-text paired datasets that require expensive human annotation, thus limiting model scalability. Current unsupervised image captioning methods primarily focus on leveraging zero-shot learning capabilities of large pre-trained models (e.g. CLIP, GPT-2), yet still face persistent challenges including modality gaps, inefficient inference, and excessive noise incorporation, which constrain model accuracy and generalization capabilities. To address these limitations, we propose SGDR-Cap ( Similarity- Guided Denoising Reconstruction for Captioning), a novel unsupervised image captioning method that bridges the vision-language modality gap through a similarity-guided denoising reconstruction module. Our method leverages similarity information to guide the reconstruction of authentic text features during caption generation while simultaneously forcing the model to learn how to extract crucial image-relevant features and filter out unnecessary noise information. This enhances both coarse- and fine-grained cross-modal alignment. Furthermore, our approach jointly optimizes denoising reconstruction loss and language modeling loss, ensuring accuracy and fluency, and promoting greater diversity. Extensive evaluations on the MSCOCO and Flickr30K benchmarks demonstrate that our method achieves state-of-the-art results across all major metrics, with the most notable gain on the CIDEr score, improving from 101.1 to 104.4.
Dongnan Yang, Thomas Wu 0001, Xichun Li, Yuan Yan Tang, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.4
2026 Enhanced small object detection in aerial imagery through context-aware and localization-optimized deep learning
Qinghua Lai, Xichun Li, Qizong Lu, Jiangtao Peng
Vis. Comput.4
2026 Enhancing video captioning with contextual anchor-guided semantic modeling
Xichun Li, Thomas Wu 0001, Yifeng Tan
Vis. Comput.3
2025 MSSCFormer: Multigranularity Spatial-Spectral Convolution Transformer Network for Hyperspectral Image Classification
abstract
Recently, Convolutional Neural Networks (CNNs) and Transformer have achieved considerable success in Hyperspectral Image (HSI) classification tasks. However, existing methods not only lack the study of spectral variability of samples from the same land class, but also struggle to mine local-global spectral information and spatial structure information of HSI at different granularities effectively. To mitigate these limitations, this article proposes a multigranularity spatial-spectral convolution Transformer network (MSSCFormer), which can reduce the intra-class spectral differences of samples and extract multigranularity spatial-spectral features from a local-global-local perspective. Specifically, MSSCFormer consists of three components: intra-class spectral attention (ICSA), spatial-spectral feature extractor (SSFE), and global-local convolutional Transformer (GLCT). Firstly, ICSA redistributes the spectral weights of samples of the same land class by establishing an attention mapping between spectral channels within the class to reduce the intra-class spectral differences. Second, SSFE extracts shallow spatial features and multigranularity spectral features of HSI with reassigned spectral weights samples from a local perspective. Finally, GLCT takes advantage of CNNs and Transformer, it uses MHSA to model global spectral features and utilizes local context feature block (LCFB) to capture local spatial features in a multigranularity way. Experimental results on three benchmark datasets show that MSSCFormer exhibits excellent performance on the HSI datasets and outperforms state-of-the-art HSI classification algorithms.
Thomas Wu 0001, Lianhui Liang, Xichun Li, Jun Li 0009
IEEE Trans. Geosci. Remote. Sens.5
2024 LCSTR: Scene Text Recognition with Large Convolutional Kernels
abstract
The task of scene text recognition involves processing information from two modalities: images and text, thereby requiring models to have the ability to extract features from images and model sequences simultaneously. Although linguistic knowledge greatly aids scene text recognition tasks, the extensive use of language models in sequence modeling and model prediction stages in recent years has made model architectures increasingly complex and inefficient. In this paper, we propose LCSTR, a pure convolutional visual model that can complete text recognition without the need for attention mechanisms or language models. This approach applies large kernels to text recognition tasks for the first time, extracting word-level text information through large text-aware blocks, capturing long-range dependencies between characters, and using small text-aware blocks to obtain local features within characters. Experiments show that this model strikes a good trade-off between accuracy and speed, achieving notable results on seven public benchmarks, validating the generalizability and effectiveness of this method. Furthermore, owing to the absence of a language module, this model demonstrates remarkable accuracy even in limited sample scenarios, and the lightweight and low computational overhead features make it suitable for engineering applications.
Jing Wang 0233, Patrick Shen-Pei Wang, Xichun Li, Huiwu Luo, Huafu Xu
Int. J. Pattern Recognit. Artif. Intell.7
2023 Constraint-Based Adversarial Networks for Unsupervised Abstract Text Summarization
abstract
Abstract text summarization is a classic sequence-to-sequence natural language generation task. In order to improve the quality of unsupervised abstract text summarization in unsupervised mode, we propose two constraints for training text summarization model, embedding space constraint and information ratio constraint. We construct a generative adversarial network with two discriminators based on these two constraints (TC-SUM-GAN). We use unsupervised and supervised methods to train the model in the experiment. Experimental results show that the ROUGE-1 value of the unsupervised TC-SUM-GAN increases by [Formula: see text] points compared with the basic model and at least 1.96 points compared with other comparative models. The ROUGE scores of the supervised TC-SUM-GAN are also improved. TC-SUM-GAN achieves very competitive results for the metrics of ROUGE-1 and ROUGE-2. In addition, the abstracts generated by our model are closer to those generated manually.
Liwei Jing, Yujian Yuan, Zuqiang Meng, Yifeng Tan, Patrick Shen-Pei Wang, Xichun Li
Int. J. Pattern Recognit. Artif. Intell.7
2023 Edge-Labeled and Node-Aggregated Graph Neural Networks for Few-Shot Relation Classification
abstract
Relation classification as a core technique for building knowledge graphs becomes a critical task in natural language processing. The fact that humans can learn by summarizing and generalizing limited knowledge motivates scholars to explore few-shot learning. Graph neural networks provide a method to measure the distance between nodes, which improves the model effect in the problem of few-shot relation classification. However, graph neural network methods focus only on node information and ignore edge information which implies inter-class and intra-class relations. This paper proposes edge-labeled and node-aggregated graph neural networks (ENGNNs) for few-shot relation classification: edge labels are encoded and used for node information aggregation. In addition, a process of semi-supervised learning is designed to discover a better solution for one-shot learning. Compared with previous methods, experimental results show that the proposed ENGNN model improves the performance of the graph neural network on the FewRel dataset.
Xichun Li, Patrick Shen-Pei Wang, Zuqiang Meng
Int. J. Pattern Recognit. Artif. Intell.3
2022 FFP: joint Fast Fourier transform and fractal dimension in amino acid property-aware phylogenetic analysis
abstract
BACKGROUND: Amino acid property-aware phylogenetic analysis (APPA) refers to the phylogenetic analysis method based on amino acid property encoding, which is used for understanding and inferring evolutionary relationships between species from the molecular perspective. Fast Fourier transform (FFT) and Higuchi's fractal dimension (HFD) have excellent performance in describing sequences' structural and complexity information for APPA. However, with the exponential growth of protein sequence data, it is very important to develop a reliable APPA method for protein sequence analysis. RESULTS: Consequently, we propose a new method named FFP, it joints FFT and HFD. Firstly, FFP is used to encode protein sequences on the basis of the important physicochemical properties of amino acids, the dissociation constant, which determines acidity and basicity of protein molecules. Secondly, FFT and HFD are used to generate the feature vectors of encoded sequences, whereafter, the distance matrix is calculated from the cosine function, which describes the degree of similarity between species. The smaller the distance between them, the more similar they are. Finally, the phylogenetic tree is constructed. When FFP is tested for phylogenetic analysis on four groups of protein sequences, the results are obviously better than other comparisons, with the highest accuracy up to more than 97%. CONCLUSION: FFP has higher accuracy in APPA and multi-sequence alignment. It also can measure the protein sequence similarity effectively. And it is hoped to play a role in APPA's related research.
Yujian Yuan, Xichun Li, Zuqiang Meng
BMC Bioinform.5
2022 Phylogenetic Analysis: A Novel Method of Protein Sequence Similarity Analysis
abstract
Protein sequence similarity analysis (PSSA) is a significant task in bioinformatics, which can obtain information about unknown sequences such as protein structures and homology relationships. Protein sequence refers to the series of amino acids with rich physical and chemical properties, namely the basic structure of proteins. However, sequence similarity analysis and phylogenetic analysis between different species which have complex amino acid sequences is a challenging problem. In this paper, nine properties of amino acids were considered and the sequence was converted into numerical values by principal component analysis (PCA); with Haar Wavelet Transform, and Higuchi fractal dimension (HFD), a new feature vector is constructed to represent the sequence; Spearman distance was selected to calculate the distance matrix and the phylogenetic tree was constructed. In this paper, two representative protein sequences (9 ND5 (NADH dehydrogenase 5) and 8 ND6 (NADH dehydrogenase 6)) were selected for similarity analysis and phylogenetic analysis, and compared with MEGA software and other existing methods. The extensive results show that our method is outperforming and results consistent with the known facts.
Wei Li 0330, Zuqiang Meng, Patrick Shen-Pei Wang, Xichun Li
Int. J. Pattern Recognit. Artif. Intell.6
2022 Improving Utterance Rewriter Based on MMI and Text Data Augmentation
abstract
In multi-round dialogue tasks, how to maintain the consistency of model answers is a major research challenge. Every answer to the model should be time dependent, causal, and logical. In order to maintain the consistency of the personality, dialogue style, and context of the model, it is necessary to retain the key information in the historical dialogue as much as possible so that the model can generate more accurate answers. Utterance rewriting is a technique that replenishes the information of the current sentence by analyzing the historical dialogue, so as to retain the key information. This paper mainly uses text augmentation, Maximum Mutual Information (MMI) method and character correction method based on Knuth–Morria–Pratt (KMP) algorithm to improve the effect of utterance rewriting generation. The number of original statement rewriting datasets is limited, and the cost of manual manufacturing is too high. By using the method of text data augmentation based on coreference resolution, the positive dataset that is missing from the statement rewriting dataset is repaired. At the same time, the existing datasets are expanded to increase the number of data. The generated results are optimized by using the MMI method, and the KMP character correction method is used to modify the wrong characters to improve the overall accuracy.
Wei Li 0330, Zuqiang Meng, Patrick Shen-Pei Wang, Xichun Li, Huiwu Luo
Int. J. Pattern Recognit. Artif. Intell.6
2022 Automatic Detection of Bridge Surface Crack Using Improved YOLOv5s
abstract
Bridge crack detection is a key task in the structural health monitoring of Civil Engineering. In the traditional bridge crack detection methods, there exist some problems such as high cost, low speed, and complex structure. This paper developed a bridge surface crack detection system based on improved YOLOv5s. The GhostBottleneck module was employed to replace the classic C3 module of the YOLOv5s backbone network, meanwhile the channel attention module namely ECA-Net was also added to the network, which not only reduced the amount of calculation, but also enhanced the ability of the network in extracting cross-channel information features. The adaptive spatial feature fusion (ASFF) was introduced to address the conflict problem caused by the inconsistency of feature scale in the network feature fusion stage, and the transfer learning was utilized to train the network. The experimental results showed that the improved YOLOv5s performed better than Faster R-CNN, SSD, YOLOv3, and YOLOv5s, with the Precision of 93.6%, Recall of 95.4%, and mAP of 98.4%. Further, the improved YOLOv5s was deployed in PyQt5 to realize the real-time detection of bridge cracks. This research showed that the proposed model not only provides a novel solution for bridge surface crack detection, but also has certain industrial application value.
Thomas Wu 0001, Zuqiang Meng, Youju Huang, Patrick Shen-Pei Wang, Xichun Li
Int. J. Pattern Recognit. Artif. Intell.8
2022 Multi-Content Merging Network Based on Focal Loss and Convolutional Block Attention in Hyperspectral Image Classification
abstract
Simultaneous extraction of spectral and spatial features and their fusion is currently a popular solution in hyperspectral image (HSI) classification. It has achieved satisfactory results in some research. Because the scales of objects are often different in HSI, it is necessary to extract multi-scale features. However, this aspect was not taken into account in many spectral-spatial feature fusion methods. This causes the model to be unable to get sufficient features on scales with a large difference range. The model (MCMN: Multi-Content Merging Network) proposed in this paper designs a multi-branch fusion structure to extract multi-scale spatial features by using multiple dilated convolution kernels. Considering the interference of the surrounding heterogeneous objects, the useful information from different directions is also fused together to realize the merging of multiple regional features. MCMN introduces a convolution block attention mechanism, which fully extracts attention features in both spatial and spectral directions, so that the network can focus on more useful parts, which can effectively improve the performance of the model. In addition, since the number of objects in each class is often discrepant, it will have some impact on the training process. We apply the focal loss function to eliminate the negative factor. The experimental results of MCMN on three data sets have a breakthrough compared with the other comparison models, which highlights the role of MCMN structure.
Fengqi Zhang, Patrick Shen-Pei Wang, Xichun Li, Huiwu Luo
Int. J. Pattern Recognit. Artif. Intell.4
2022 Multi-scale spatial-spectral fusion based on multi-input fusion calculation and coordinate attention for hyperspectral image classification
Fengqi Zhang, Patrick Shen-Pei Wang, Xichun Li, Zuqiang Meng
Pattern Recognit.4
2021 Smart Home Privacy Protection Based on the Improved LSB Information Hiding
abstract
Smart home is an emerging form of the Internet of Things (IoT), enabling people to enjoy a convenient and intelligent life. The data generated by smart home devices are transmitted through the public channel, which is not secure enough, so the secret data in smart home are easily intercepted by malicious adversaries. In order to solve this problem, this paper proposes a smart home privacy protection method combining DES encryption and the improved Least Significant Bit (LSB) information hiding algorithm, changing the practice of directly exposing smart home secret information to the Internet, first, using Data Encryption Standard (DES) encryption to encrypt the smart home information and second, the improved LSB information hiding algorithm is used to hide the ciphertext, so that the adversary cannot detect the smart home secret information. The goal of the scheme is to provide a double protection for the secure transmission of the smart home secret information. If an attacker wants to carry out an attack, it has to break through at least two defense lines, which seems impossible to do. Experiment results show that the improved LSB algorithm is more robust than the existing algorithms, and it is very safe. Therefore, the scheme proposed in this paper is very practical for protecting the smart home secret information.
Haiyu Deng, Ren Ping Liu 0001, Patrick Shen-Pei Wang, Xiaocui Dang, Yuan Yan Tang, Xichun Li
Int. J. Pattern Recognit. Artif. Intell.7
2021 PseKNC and Adaboost-Based Method for DNA-Binding Proteins Recognition
abstract
DNA-binding proteins are an essential part of the DNA. It also an integral component during life processes of various organisms, for instance, DNA recombination, replication, and so on. Recognition of such proteins helps medical researchers pinpoint the cause of disease. Traditional techniques of identifying DNA-binding proteins are expensive and time-consuming. Machine learning methods can identify these proteins quickly and efficiently. However, the accuracies of the existing related methods were not high enough. In this paper, we propose a framework to identify DNA-binding proteins. The proposed framework first uses PseKNC (ps), MomoKGap (mo), and MomoDiKGap (md) methods to combine three algorithms to extract features. Further, we apply Adaboost weight ranking to select optimal feature subsets from the above three types of features. Based on the selected features, three algorithms (k-nearest neighbor (knn), Support Vector Machine (SVM), and Random Forest (RF)) are applied to classify it. Finally, three predictors for identifying DNA-binding proteins are established, including [Formula: see text], [Formula: see text], [Formula: see text]. We utilize benchmark and independent datasets to train and evaluate the proposed framework. Three tests are performed, including Jackknife test, 10-fold cross-validation and independent test. Among them, the accuracy of ps+md is the highest. We named the model with the best result as psmdDBPs and applied it to identify DNA-binding proteins.
Patrick Shen-Pei Wang, Xichun Li
Int. J. Pattern Recognit. Artif. Intell.5