Xiaoli Lin

dblp:88/3623 · DBLP profile ↗
← Back
71ranked-venue papers
21as first author
44since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 63 · 18 first-author · 37 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Side Effect Aware Multi-Source Representation Learning for Drug-Drug Interaction Prediction
Tongtong Xie, Youkang Jiang, Zhuoya Hu, Xiaoli Lin
ICIC (30)4
2026 A multi-target drug design method based on target feature fusion
abstract
BACKGROUND: Targeted drugs are medications designed to treat diseases by targeting specific sites on cancerous or diseased cells. Multi-target drugs can target multiple protein sites to treat diseases, improving therapeutic efficiency, but are more challenging to design. Computer-aided targeted drug design can reduce costs and shorten development time, with most drugs being single-target. Recent research on multi-target drug design has focused on optimizing single-target drugs into multi-target drugs, but this approach has limitations. This study proposes a multi-target drug design method based on protein feature fusion, which encodes and integrates features based on the target's sequence characteristics, enabling the design of multi-target drugs without prior knowledge of the targeted drug. The target protein sequences are embedded to extract features. Each target's features are independently encoded into latent vectors, while the features of multiple targets are encoded into similarity latent vectors. By leveraging both individual target features and the similarity features among targets, multi-target drugs can be efficiently designed. RESULTS: We validated the proposed multi-target drug design method on three groups of targets: the 3CLpro and PLpro targets for COVID-19, the TAAR1 and DRD2 targets for schizophrenia, and the MEK1 and mTOR targets for tumors. The designed multi-target drugs can be docked with target proteins possessing unique molecular structures, tailored to the specific requirements of different target pocket structures. The excellent fit between the molecular structures of the multi-target drugs and the protein structures of multiple targets validates the performance of the proposed method. CONCLUSIONS: The proposed method can efficiently design multi-target drugs with stronger predicted binding affinities than those reported in previous studies. These drugs are capable of adapting to multiple targets based on the features of the target proteins. Additionally, the model demonstrates excellent generalization ability for untrained multiple targets.
Xiaoli Lin
BMC Bioinform.2
2026 Multi-scale hierarchical causality-inspired graph network for interpretable anomaly detection in industrial time series
Yushan Fang, Yu Yao 0002, Wei Yang 0044, Xiaoli Lin, Chuan Sheng
Eng. Appl. Artif. Intell.4
2026 A Real-Time Hacker Group Identification Method for Industrial Control Systems Based on Gaussian Self-Organizing Incremental Neural Network
abstract
Several methods have been proposed to identify hacker groups targeting Industrial Control Systems (ICS); however, real-time identification of complex hacker groups with new data distributions, different attack patterns and Internet Protocol (IP) prefixes remains to be accomplished. This paper utilises the honeynet data to propose a hacker group identification method, Hacker-Identifier, which focuses on identifying hacker groups targeting ICS. To effectively distinguish attack patterns with different temporal features, we present a novel attack pattern temporal feature modelling method, which combines temporal feature extraction and learning. On this basis, we propose an improved genetic algorithm-based feature selection method to further improve the accuracy of attack pattern classification. To accurately identify attack patterns and hacker groups with new data distributions in real time, we propose a Gaussian self-organising incremental neural network, which can not only effectively distinguish highly overlapping classes, but also prevent excessive segmentation of homogeneous classes. To accurately identify hacker groups with different attack patterns and IP prefixes, we propose an improved hacker group identification method that effectively fuses the attribute and structural features of hackers through heterogeneous graph node embedding. Take the Modbus protocol as an example, we implement a prototype of Hacker-Identifier, and use 2.5 years of real-world honeynet dataset to evaluate its performance. The experimental results demonstrate its effectiveness and superiority compared with existing state-of-the-art hacker group identification methods. It achieves weighted F1-score of 98.86% and accuracy of 99.12%, outperforming the state-of-the-art work by 8.12% and 6.86% respectively, while reducing the total error rate by 17.11%.
Xiaoli Lin, Yu Yao 0002, Yushan Fang, Boxue Song, Wei Yang 0044
IEEE Internet Things J.1
2025 Drug-Drug Interaction Prediction Based on Multi-Scale Feature Fusion
abstract
Although deep learning models based on multi-source biomedical data have made progress in DDI prediction, existing approaches often suffer from shallow cross-modal interactions when integrating heterogeneous data modalities, limiting their ability to fully capture high-order, nonlinear relationships across modalities. Therefore, we propose a new model MSFF-DDI. The model integrates TransE and a graph attention network (GAT) to extract semantic features from knowledge graphs, and uses a lightweight graph convolutional network to capture DDI network topology. It employs GIN, Word2Vec, and ESM2 to represent molecular graphs, SMILES, and protein sequences, enabling multi-granularity drug representation. A master-slave dual-channel interaction module with gating fuses cross-modal features dynamically. Experiments show MSFF-DDI outperforms state-of-the-art methods across four metrics, with ablation studies confirming the contribution of each component.
Yunhao Du, Xiaoli Lin
BIBM2
2025 A Hybrid Architecture for 3D Abdominal Medical Images Based on Mamba
Baitao Li, Xiaoli Lin, He Deng
ICIC (25)3
2025 MMF2Drug: A Multi-modal Feature Fusion Method for Improving Targeted Drug Design
Xiongwei Liao, Xiaoli Lin
ICIC (26)2
2025 Combining Improved Relational Embedding and Multi-Hop Sampling for Temporal Knowledge Graph Reasoning
Yuheng Guo, Xiaoli Lin, Mengxiang Wang
ICIC (8)3
2025 DrugGAN-MSM: A Generative Adversarial Approach to Molecular Design Integrating Masked Modeling and Multi-objective Optimization
Zihang Xie, Xiaoli Lin
ICIC (25)3
2025 A 3D Liver and Tumor Segmentation Method Based on U-Mamba and Efficient Paired Attention
Shili Yang, Xiaoli Lin, He Deng
ICIC (27)3
2025 BIOFUSE-DDI: A Dual-Source Transformer Framework for Drug-Drug Interaction Prediction
Hengpeng Zhao, Shuoyu Cui, Xiaoli Lin
ICIC (26)3
2025 PAMol: Pocket-Aware Drug Design Method with Hypergraph Representation of Protein Pocket Structure and Feature Fusion
abstract
Efficient generation of targeted drug molecules is crucial in the field of drug discovery. Most existing methods neglect the high-order information in the structure of protein pockets, limiting the performance of generated drug molecules. This paper proposes a pocket-aware drug design framework, namely PAMol, constructing the hypergraph to represent the spatial structure of protein pockets, effectively capturing high-order relations and neighborhood information within the pocket structures. This framework also fuses different modal embeddings from proteins and molecules, to generate high-quality molecules. In addition, a conditional molecule generation module uses the high-order structural information in protein pockets as constraints to more accurately generate molecules for specific targets. The performance of PAMol has been assessed by analyzing generated molecules in terms of vina score, high affinity, QED, SA, LogP, Lipinski, diversity, and time. Experimental results demonstrate the potential of PAMol for targeted drug design. The source code is available at https://github.com/YICHUANSYQ/PAMol.git.
Xiaoli Lin, Xiongwei Liao
IJCAI1
2025 A Generative Strategy for Target-Oriented Molecular Design: Integrating Docking Scores with Drug-Likeness Constraints
abstract
With the growing application of deep learning in drug design, efficiently generating molecules with desirable drug-like properties and target specificity remains a key challenge. Existing methods often rely on complex fusion of molecular and protein features, leading to high data demands and training costs. To address this, a novel molecule generation model based on multi-constraint optimization is proposed, incorporating docking scores, Lipinski’s Rule of Five, and other pharmacological constraints to guide generation. The model is built on a Bidirectional Gated Recurrent Unit (BiGRU) and improves training efficiency through cross-entropy loss optimization. Experimental results show that the proposed approach achieves superior performance in molecular validity, novelty, and uniqueness, and generates compounds with lower binding energies to target proteins, significantly enhancing the practicality and scalability of molecular design.
Xiaolong Zhang 0002, Xiaoli Lin
SMC3
2025 A Real-Time Anomaly Detection Method for Industrial Control Systems Based on Long-Short Period Deterministic Finite Automaton
abstract
Anomaly detection has proven effective in detecting cyber-attacks in industrial control systems (ICS). However, most existing anomaly detection methods suffer from low accuracy because they ignore the effects of packet loss and network delay on time features, the sequential nature of transition time, masquerade transitions, and system recovery. Meanwhile, current cyber-physical model (CPM) construction methods struggle to effectively address the state explosion problem and properly balance the removal and retention of low frequency states (LFS). In this article, we propose a novel baseline model for ICS to detect anomalies through learning device-level polling time patterns and system-level CPM. The polling time pattern learning method reduces the effects of packet loss and network delay on time features by extracting only matching packets and replacing outliers. The CPM construction method mitigates state explosion through mixed-event discretization, reduces the effects of network delay on transition/action times through outlier replacement, and captures the sequential nature of transition times with circular permutation sets (CPSets). CPM model optimization uses a post-pruning algorithm to balance the removal and retention of LFSs, and a CPM periodicity detection method that mitigates the effects of network delay to ensure that all industrial process periods are detected. A real-time anomaly detection method with a two-layer defence mechanism is proposed using the baseline model. Experimental results from two lab-scale ICSs with six process-related attacks confirm the effectiveness and superiority of the proposed method. It achieves average F1 scores of 98.81% and accuracy of 99.24%, outperforming the state-of-the-art work by 18.51% and 13.96%, respectively.
Xiaoli Lin, Yu Yao 0002, Wei Yang 0044, Xiaoming Zhou, Guangao Li
IEEE Internet Things J.1
2025 An Efficient Targeted Drug Design Method Using Framework of Multiscale Encoder-Decoder
abstract
This paper describes a targeted drug design method based on a framework of multiscale encoder-decoder. Encoders are used to encode target gene and protein features. A decoder is used to design drugs based on target features. This method fuses target gene and protein information for targeted drug design, and invokes effective feature extraction strategies. A multilevel gene feature extraction (MGFE) is proposed to extract multilevel target features by extracting base and codon features in gene expression. The process of extracting features from nucleotide sequences by MGFE based gene encoder simulates the process of gene transcription and translation. Meanwhile, a multi-embedding protein feature extraction (MPFE) is proposed to extract target protein features from amino acid sequences. The MPFE based protein encoder includes three embedding layers which provides a unique linear layer for each amino acid. According to structural characteristics of proteins, amino acids with different positions but the same type are embedded into the same embedding vector without location encoding. Finally, a gated recurrent unit based drug decoder is used to decode gene and protein features, and creates new targeted drugs. The experiments adminstrate that the proposed method outperforms the previous ones in terms of validity, novelty and binding affinity.
Xiaoli Lin, Jing Hu 0003, Jun Pang 0002, Xiaolong Zhang 0002
IEEE Trans. Comput. Biol. Bioinform.2
2025 InSyfer: Industrial Control Protocols Syntax Inference via Graph Representation Learning
abstract
Industrial control protocols (ICPs) play a significant role in ensuring dependable interconnection among devices in industrial environments. Protocol reverse engineering (PRE) techniques are commonly used to analyze a large number of agnostic and proprietary protocols based on network traffic traces or programs. However, conventional PRE methods face several challenges in reversing ICPs with complex data representations that contain rich structural features. In this work, we present a new perspective on message representation using the graph, and design a syntax inference framework for ICPs reverse analysis (InSyfer). Specifically, we propose a novel method to construct a single message graph for entire traces, automatically extracting syntactical similarity features. We also design an adaptive message clustering model that abstracts the clustering problem into a binary pairwise-classification framework to judge whether pairs of messages belong to the same groups and jointly optimizes it with feature extraction. The above design enables InSyfer to accurately identify message types and greatly improves the correctness of protocol format inference. We conduct extensive experiments to verify the effectiveness of InSyfer. Evaluations of four standard ICPs and two unknown protocols demonstrate that InSyfer outperforms the state-of-the-art PRE methods.
Daoqing Yang, Yu Yao 0002, Yao Shan, Xiaoli Lin, Wei Yang 0044, Licheng Yang 0002
IEEE Trans. Dependable Secur. Comput.5
2024 LKC-NET: A Liver and Tumor Segmentation Method based on Large Kernel Parallel Dilated Convolution
abstract
Accurate drug dosage in cancer treatments depends on precise liver and tumor size estimation, but tumors pose challenges for automatic segmentation due to their small, scattered volumes, complex structures, and low contrast against surrounding organs. Conventional U-Net models struggle to capture these microstructural details. To address this, we propose LKC-Net, a liver and tumor segmentation method leveraging large kernel parallel dilated convolutions. The encoder utilizes depth-wise separable large kernels combined with parallel dilated convolutions and squeeze-and-excitation (SE) modules. This architecture enhances the receptive field and improves the capture of small-scale patterns. SE modules filter redundant information, focusing the model on key regions. Experiments on ATLAS, LiTS, and private datasets demonstrate LKC-Net’s superior performance in liver and tumor segmentation tasks.
Baitao Li, Xiaolong Zhang 0002, Xiaoli Lin, He Deng
BIBM3
2024 Detecting DTI Using Graph Embedding and Multi-Head Attention Mechanism
abstract
Drug-target interaction (DTI) prediction has an important role in drug discovery, significantly expediting the drug design and development process. This paper proposes a new model, ERW_BiAN, that seamlessly integrates graph embedding representations with deep learning network. Specifically, we utilize an improved graph embedding model, ERW, to generate comprehensive feature vectors for each node within the knowledge graph. Subsequently, these feature vectors are fed into the BiAN model, which leverages an attention mechanism to assign precise weights to sequences. Experimental results demonstrate the superior accuracy and predictive prowess of ERW_BiAN in DTI prediction tasks. Furthermore, the case about COVID-19 shows that ERW_BiAN has the better application potential for predicting drug-target interactions.
Xiaoli Lin, Yaoyu Chen, Haiping Yu, Zimo Wu, Xiaolong Zhang 0002
BIBM1
2024 ResPocket: A Multi-Scale Feature Fusion Method for Improving Protein Binding Site Detection
abstract
Accurate detection of protein binding sites is critical for facilitating drug design. Most existing models rely on features at a single scale, leading to the loss of crucial global or local structural information. Therefore, this paper proposes a novel 3D model (ResPocket), to fully capture the multi-scale features of proteins. First, Residual Block (ResB) based on U-Net is used to capture features from different scales, including global and local structural information. Then, an Enhanced Attention Module (EAM) is designed to encode semantic dependencies by effectively capturing global context information. In addition, ResPocket focuses on learning local information about protein pockets with Self-Guided Mask Block (SGMB), to increase the robustness of the overall model and reduce its dependence on external mask guidance. Experimental results show that ResPocket improves by an average of 4.89% on DCA Top-n, by an average of 1.48% and 2.25% on DCC and DVO metrics, respectively.
Xiaoli Lin, Yaoyu Chen, Xiongwei Liao, Zimo Wu, Xiaolong Zhang 0002
BIBM1
2024 A Multi Target Drug Design Method Based on Target Protein Sequence and Feature Similarity
abstract
This paper describes a multi target drug design method based on features of the target proteins. The multi target drug which inhibit multiple proteins have prospective applications, but design is difficult. In this study, the target protein sequences are embedded to obtain target features. The features of each target are independently encoded as target latent vectors, and the features of multiple targets are jointly encoded as target similarity latent vectors. Based on the features of targets and the similarity features among the targets, multi target drugs are efficiently designed. In the experiments, the designed multi target drugs can be docked with target proteins with a proprietary molecular structure according to the requirements of different target pocket structures. The excellent fitness among the molecular structure of multi target drugs and the protein structure of multiple targets confirms the performance of the proposed method. In the molecular docking part of the experi¬ments, the binding affinity of designed multi target drugs is far better than that in the previous studies, which administrates the better performance of this work.
Xiaoli Lin, Jing Hu 0003, Xiaolong Zhang 0002
BIBM2
2024 Knowledge Completion Method Based on Relational Embedding with GNN
Zhuang Yin, Honghong Tan, Xiaoli Lin
ICIC (13)4
2024 A 3D Liver Semantic Segmentation Method Based on U-shaped Feature Fusion Enhancement
Daoran Jiang, Xiaoli Lin, He Deng
ICIC (2)3
2024 A feature selection based on genetic algorithm for intrusion detection of industrial control systems
Yushan Fang, Yu Yao 0002, Xiaoli Lin
Comput. Secur.3
2024 KGRLFF: Detecting Drug-Drug Interactions Based on Knowledge Graph Representation Learning and Feature Fusion
abstract
Accurate prediction of drug-drug interactions (DDIs) plays an important role in improving the efficiency of drug development and ensuring the safety of combination therapy. Most existing models rely on a single source of information to predict DDIs, and few models can perform tasks on biomedical knowledge graphs. This paper proposes a new hybrid method, namely Knowledge Graph Representation Learning and Feature Fusion (KGRLFF), to fully exploit the information from the biomedical knowledge graph and molecular structure of drugs to better predict DDIs. KGRLFF first uses a Bidirectional Random Walk sampling method based on the PageRank algorithm (BRWP) to obtain higher-order neighborhood information of drugs in the knowledge graph, including neighboring nodes, semantic relations, and higher-order information associated with triple facts. Then, an embedded representation learning model named Knowledge Graph-based Cyclic Recursive Aggregation (KGCRA) is used to learn the embedded representations of drugs by recursively propagating and aggregating messages with drugs as both the source and destination. In addition, the model learns the molecular structures of the drugs to obtain the structured features. Finally, a Feature Representation Fusion Strategy (FRFS) was developed to integrate embedded representations and structured feature representations. Experimental results showed that KGRLFF is feasible for predicting potential DDIs.
Xiaoli Lin, Zhuang Yin, Xiaolong Zhang 0002, Jing Hu 0003
IEEE ACM Trans. Comput. Biol. Bioinform.1
2023 Drug-Target Interaction Prediction Based on Drug Subgraph Fingerprint Extraction Strategy and Subgraph Attention Mechanism
Xiaolong Zhang 0002, Xiaoli Lin, Jing Hu 0003
ADMA (3)3
2023 Anti-3CLpro Molecular Design Based on the Model Constrained by Specific DTIs
abstract
Computer-aided drug design and artificial intelligence-driven drug design have accelerated drug discovery. However, how to design effective drugs that have strong interaction ability with target proteins to further improve the efficacy of drugs in treating diseases remains a key issue. This paper proposes a new target-specific drug generative model 3CLpro2mol to generate new drug molecules, which uses features of drug-target interactions (DTIs) to constrain the correlation between the drug and the target protein. To obtain as many drug-target interaction features as possible from a small amount of data, a small molecule extraction strategy is proposed to ensure the diversity of small molecules in the training samples. To improve the efficiency and accuracy of the generative model, a TOP K sampling strategy is used to generate tokens, which can improve the rationality and diversity of the generated molecules. The experimental results show that the proposed model has the potential to generate small molecules that interact better with the target protein.
Xiaoli Lin, Xiaolong Zhang 0002
BIBM1
2023 Generating Molecules Conditional on 3D Protein Pockets with HGAF
abstract
Most current representations of protein pockets are the atom-pair graph, which ignore the global structural information of amino acids. Therefore, we propose a new molecular generation model, which uses the hypergraph to represent protein pocket structure, and combines the structural features obtained by atom-pair graph representation. These two levels of graphs are more capable of representing the complex structural information of protein pockets. Then, the graphs of the two levels of the protein pockets are input into the improved network model, which is named Hypergraph Graph Attention Fusion (HGAF), to obtain the embedding representation of the protein pockets, which is used as a condition to constrain the molecule generation. The molecules sampled by HGAF are subjected to quality assessment and docking targeting validation. Experimental results show that the molecules generated by proposed method can achieve better results in both of these assessment approaches.
Qinglian Zhu, Xiaoli Lin, Xiaolong Zhang 0002
BIBM2
2023 DU-DANet: Efficient 3D Automatic Brain Tumor Segmentation Based on Dual Attention
Zhenhua Cai, Xiaoli Lin, Xiaolong Zhang 0002, Jing Hu 0003
ICIC (3)2
2023 An Efficient Drug Design Method Based on Drug-Target Affinity
Xiaolong Zhang 0002, Xiaoli Lin, Jing Hu 0003
ICIC (3)3
2023 Drug-Target Interaction Prediction via Graph Auto-Encoder and Multi-Subspace Deep Neural Networks
abstract
Computational prediction of drug-target interaction (DTI) is important for the new drug discovery. Currently, the deep neural network (DNN) has been widely used in DTI prediction. However, parameters of the DNN could be insufficiently trained and features of the data could be insufficiently utilized, because the DTI data is limited and its dimension is very high. To deal with the above problems, in this paper, a graph auto-encoder and multi-subspace deep neural network (GAEMSDNN) is designed. GAEMSDNN enhances its learning ability with a graph auto-encoder, a subspace layer and an ensemble layer. The graph auto-encoder can preserve the reconstruction information. The subspace layer can obtain different strong feature subsets. The ensemble layer in the GAEMSDNN can comprehensively utilize these strong feature subsets in a unified optimization framework. As a result, more features can be extracted from the network input and the DNN network can be better trained. In experiments, the results of GAEMSDNN are significantly improved compared to the previous methods, which validates the effectiveness of our strategies.
Xiaolong Zhang 0002, Xiaoli Lin
IEEE ACM Trans. Comput. Biol. Bioinform.3
2022 Towards DDIs Identification by Knowledge Graph with BiRW and Back Aggregation
abstract
Effective identification of potential drug-drug interactions (DDIs) can prevent adverse effects caused by DDIs to a certain extent. This paper proposes the new hybrid method for predicting DDIs, which is called BiRW-KGBAN that combines bidirectional random walk and back aggregation based on knowledge graph. BiRW-KGBAN first constructs two knowledge graph KG-DrugBank and KG-KEGG that integrate data related to drugs. Then, an improved sampling method based on bidirectional random walk (BiRW) is used to sample the neighborhood information of drugs, including directly connected entities, related semantic relations and potential information. In addition, the path with drug as the center node is also extracted. Finally, two improved back aggregation model KGBAN1 and KGBAN2 are used to obtain the final embedded representation of the drugs. The experiments show that, compared with other existing methods, our method could obtain the higher-order topological information and the deeper potential neighborhood information of drugs, and has great improvements in DDIs prediction. The case study also shows that the proposed method has the potential for actual DDIs prediction.
Zhuang Yin, Xiaoli Lin, Xiaolong Zhang 0002
BIBM2
2022 A Targeted Drug Design Method Based on GRU and TopP Sampling Strategies
Jinglu Tao, Xiaolong Zhang 0002, Xiaoli Lin
ICIC (2)3
2022 KGAT: Predicting Drug-Target Interaction Based on Knowledge Graph Attention Network
Xiaolong Zhang 0002, Xiaoli Lin
ICIC (2)3
2022 Drug-Target Binding Affinity Prediction Based on Graph Neural Networks and Word2vec
Minghao Xia, Jing Hu 0003, Xiaolong Zhang 0002, Xiaoli Lin
ICIC (2)4
2022 An Optimization Method for Drug-Target Interaction Prediction Based on RandSAS Strategy
Huimin Xiang, Aoxing Li, Xiaoli Lin
ICIC (2)3
2022 Unsupervised Prediction Method for Drug-Target Interactions Based on Structural Similarity
Xiaoli Lin, Jing Hu 0003, Wenquan Ding
ICIC (2)2
2022 Drug-target interaction prediction via multiple classification strategies
abstract
BACKGROUND: Computational prediction of the interaction between drugs and protein targets is very important for the new drug discovery, as the experimental determination of drug-target interaction (DTI) is expensive and time-consuming. However, different protein targets are with very different numbers of interactions. Specifically, most interactions focus on only a few targets. As a result, targets with larger numbers of interactions could own enough positive samples for predicting their interactions but the positive samples for targets with smaller numbers of interactions could be not enough. Only using a classification strategy may not be able to deal with the above two cases at the same time. To overcome the above problem, in this paper, a drug-target interaction prediction method based on multiple classification strategies (MCSDTI) is proposed. In MCSDTI, targets are firstly divided into two parts according to the number of interactions of the targets, where one part contains targets with smaller numbers of interactions (TWSNI) and another part contains targets with larger numbers of interactions (TWLNI). And then different classification strategies are respectively designed for TWSNI and TWLNI to predict the interaction. Furthermore, TWSNI and TWLNI are evaluated independently, which can overcome the problem that result could be mainly determined by targets with large numbers of interactions when all targets are evaluated together. RESULTS: We propose a new drug-target interaction (MCSDTI) prediction method, which uses multiple classification strategies. MCSDTI is tested on five DTI datasets, such as nuclear receptors (NR), ion channels (IC), G protein coupled receptors (GPCR), enzymes (E), and drug bank (DB). Experiments show that the AUCs of our method are respectively 3.31%, 1.27%, 2.02%, 2.02% and 1.04% higher than that of the second best methods on NR, IC, GPCR and E for TWLNI; And AUCs of our method are respectively 1.00%, 3.20% and 2.70% higher than the second best methods on NR, IC, and E for TWSNI. CONCLUSION: MCSDTI is a competitive method compared to the previous methods for all target parts on most datasets, which administrates that different classification strategies for different target parts is an effective way to improve the effectiveness of DTI prediction.
Xiaolong Zhang 0002, Xiaoli Lin
BMC Bioinform.3
2021 Inferring DTIs Based on Similarity Clustering and CaGCN-DTI Model from Heterogeneous Network
abstract
Although much progress has been made in new drug development, it is still a costly, complicated and less efficient process. Therefore, drug repositioning studies are highly desirable. So far, many methods based on known drug-target interactions (DTIs) have been designed to detect potential DTIs, but there are many challenges for improving the performance of prediction DTIs. This paper proposes a new method (CaGCN-DTI) for DTIs prediction, which uses a heterogeneous network that incorporates a wide variety of biological data (drug, target, disease, and side effect) to discover potential interactions between drugs and targets. First, the drug-protein similarity network is preprocessed by Spectral clustering, then the graph convolutional network (GCN) combined with attention mechanism and random walk with restart (RWR) is used to aggregate message transmission to the network. Finally, the embedding vectors are used to discover potential DTIs by matrix decomposition. Experiments are performed based on ten-fold cross-validation, and the results show that CaGCN-DTI outperforms other previous methods in terms of AUC and AUPR.
Aoxing Li, Xiaoli Lin, Haiping Yu
BIBM2
2021 Prediction of Drug-Target Interactions Using Molecular Graph and GDNet-DTI Model
abstract
The prediction of drug-target interactions (DTIs) is of great significance to the fields of drug design and drug development. However, traditional biological experiments are time-consuming and cost-effective, which has prompted more people to turn their attention to the use of computers to assist in predicting DTIs. This paper proposes an improved prediction model based on multiple graph representation methods, which is GDNet-DTI that combined GCN and DeepWalk. First, a molecular map with atoms as nodes and chemical bonds as edges is generated using the SMILE sequence of drugs, and then GIN is used to extract the features of molecular map for better obtaining the complex interactions between atoms. For target proteins, the protein sequence is first represented by a word vector, and then the one-dimensional convolution is used to extract features for extracting the different levels of features. Then, based on obtained drug features and target features, a DTI-graph is generated, in which drugs and targets are represented as nodes and interactions are represented as edges. Finally, GDNet-DTI are used to obtain node neighborhood information and graph topology information of the DTI-graph. Compared with other advanced models, the results show that GDNet-DTI combined with multiple graph features can predict DTIs more accurately and effectively with DrugBank and four benchmark datasets. In addition, a case study with COVID-19 data is presented, which shows that the proposed method has the potential to predict the actual DTIs and can contribute to the development of drug discovery.
Xiaoli Lin, Haiping Yu
BIBM2
2021 De Novo Drug Design via Multi-Label Learning and Adversarial Autoencoder
abstract
generating new molecules is very important for drug design. Currently, many deep generative models have been designed, such as variational autoencoder (VAE), adversarial autoencoder (AAE), and reinforcement learning (RL). However, many problems are also existed in these models. Firstly, the information among molecules could be not utilized in optimizing these models. Secondly, some useful molecule information could be not used, such as fingerprint. Thirdly, the information contained in different molecule representations could be not used together. To overcome the above problems, in this paper, a multi-label learning and adversarial autoencoder (MLAAE) based de novo drug design method is designed. MLAAE enhances its learning ability with a new designed multi-label classifier and a new designed double AAEs collaborative optimization framework. These new designs can learn a better latent space whose global distribution is similar with the random distribution, but local distribution contains much information. As a result, the generator can be trained by more information, and the input of generator in testing is similar with that in training. The conducted experiments validate the effectiveness of our MLAAE.
Xiaolong Zhang 0002, Xiaoli Lin
BIBM3
2021 Discovering DTI and DDI by Knowledge Graph with MHRW and Improved Neural Network
abstract
Drug discovery is of great significance in medical and biological research, while the study of Drug-Target Interaction (DTI) and Drug-Drug Interaction (DDI) can help accelerate drug discovery progress. This paper proposes a new hybrid method for DTI prediction and DDI prediction, which is called MHRW2Vec-TBAN that combines graph representation learning and neural network. MHRW2VecTBAN first constructs knowledge graph KG-DTI and KG-DDI that integrate data related to drugs and targets. Then, an improved graph representation learning model, MHRW2Vec model, is used to obtain feature vectors of reflecting the network structure information for improving the performance of representation learning. Finally, the feature vectors obtained are input to the improved neural network model TextCNN-BiLSTM-Attention Network (TBAN). The experimental results show that, compared with other existing methods, our method could discover more deeper the relationship between drugs and their potential neighborhoods, and has great improvements in DTI prediction and DDI prediction. In addition, the case study of prediction COVID-19 DTI also shows that the proposed model has the potential for actual drug discovery.
Xiaoli Lin, Xiaolong Zhang 0002
BIBM2
2021 Drug-Target Interactions Prediction with Feature Extraction Strategy Based on Graph Neural Network
Aoxing Li, Xiaoli Lin, Minqi Xu, Haiping Yu
ICIC (3)2
2021 Drug-Target Interaction Prediction via Multiple Output Graph Convolutional Networks
Xiaolong Zhang 0002, Xiaoli Lin
ICIC (3)3
2021 A Diabetic Retinopathy Classification Method Based on Novel Attention Mechanism
Jinfan Zou, Xiaolong Zhang 0002, Xiaoli Lin
ICIC (1)3
2020 Inferring Drug-Target Interactions Using Graph Isomorphic Network and Word Vector Matrix
abstract
Computer-aided drug discovery can efficiently predict drug-target interactions, which helps biological researchers to narrow down the search space and reduce experimental consumption. However, the accuracy of existing prediction methods still needs to be further improved. This paper proposes a deep learning model that predicts drug-target interactions through effective strategies: the graphical representation of SMILES based on Graph Isomorphic Network (GIN) and word vector matrices of amino acid sequences based on one-dimensional convolution, which is used to extract features of drug-target pairs for classification prediction. Compared with the previous methods, the proposed method outperforms the existing work with AUC. A case study of predicting COVID-19 DTIs also shows that the proposed method can be invoked to be use for practical drug-target prediction.
Minqi Xu, Xiaolong Zhang 0002, Xiaoli Lin
BIBM3
2020 Drug-target Interaction Prediction via Multiple Output Deep Learning
abstract
Computational prediction of drug-target interaction (DTI) is very important for the new drug discovery. However, by connecting drugs and targets to form drug target pairs, the number of interactions is limit, most interactions focus on only a few targets or a few drugs, and the number of drug target pairs is far more than the number of interactions, which causes to be over fitting. To overcome the above problem, in this paper, a multiple output deep neural network (MODNN) based DTI prediction is designed. MODNN enhances its learning ability with a kind of auxiliary classifier layers. The parameters used in the training process are elaborated from the auxiliary and main classifier layers, which can increase the gradient signal that gets propagated back, utilize multi-level features to train the model, and use the features produced by the higher, middle or lower layers in a unified framework. The conducted experiments validate the effectiveness of our MODNN.
Xiaolong Zhang 0002, Xiaoli Lin
BIBM3
2020 Prediction of Drug-Target Interactions with CNNs and Random Forest
Xiaoli Lin, Minqi Xu, Haiping Yu
ICIC (2)1
2020 Efficient Classification of Hot Spots and Hub Protein Interfaces by Recursive Feature Elimination and Gradient Boosting
abstract
Proteins are not isolated biological molecules, which have the specific three-dimensional structures and interact with other proteins to perform functions. A small number of residues (hot spots) in protein-protein interactions (PPIs) play the vital role in bioinformatics to influence and control of biological processes. This paper uses the boosting algorithm and gradient boosting algorithm based on two feature selection strategies to classify hot spots with three common datasets and two hub protein datasets. First, the correlation-based feature selection is used to remove the highly related features for improving accuracy of prediction. Then, the recursive feature elimination based on support vector machine (SVM-RFE) is adopted to select the optimal feature subset to improve the training performance. Finally, boosting and gradient boosting (G-boosting) methods are invoked to generate classification results. Gradient boosting is capable of obtaining an excellent model by reducing the loss function in the gradient direction to avoid overfitting. Five datasets from different protein databases are used to verify our models in the experiments. Experimental results show that our proposed classification models have the competitive performance compared with existing classification methods.
Xiaoli Lin, Xiaolong Zhang 0002, Xin Xu 0007
IEEE ACM Trans. Comput. Biol. Bioinform.1
2019 Effective Analysis of Hot Spots in Hub Protein Interfaces Based on Random Forest
Xiaoli Lin, Fengli Zhou
ICIC (2)1
2019 Research on HP Model Optimization Method Based on Reinforcement Learning
Fengli Zhou, Xiaoli Lin
ICIC (2)2
2019 Resource Efficiency Optimization for Big Data Mining Algorithm with Multi-MapReduce Collaboration Scenario
Fengli Zhou, Xiaoli Lin
ICIC (3)2
2019 Efficiently Predicting Hot Spots in PPIs by Combining Random Forest and Synthetic Minority Over-Sampling Technique
abstract
Hot spot residues bring into play the vital function in bioinformatics to find new medications such as drug design. However, current datasets are predominately composed of non-hot spots with merely a tiny percentage of hot spots. Conventional hot spots prediction methods may face great challenges towards the problem of imbalance training samples. This paper presents a classification method combining with random forest classification and oversampling strategy to improve the training performance. A strategy with an oversampling ability is used to generate hot spots data to balance the given training set. Random forest classification is then invoked to generate a set of forest trees for this oversampled training set. The final prediction performance can be computed recursively after the oversampling and training process. This proposed method is capable of randomly selecting features and constructing a robust random forest to avoid overfitting the training set. Experimental results from three data sets indicate that the performance of hot spots prediction has been significantly improved compared with existing classification methods.
Xiaolong Zhang 0002, Xiaoli Lin, Jiafu Zhao, Xin Xu 0007
IEEE ACM Trans. Comput. Biol. Bioinform.2
2018 Identification of Hotspots in Protein-Protein Interactions Based on Recursive Feature Elimination
Xiaoli Lin, Xiaolong Zhang 0002, Fengli Zhou
ICIC (1)1
2018 A Novel Image Denoising Algorithm Based on Non-subsampled Contourlet Transform and Modified NLM
Huayong Yang, Xiaoli Lin
ICIC (3)2
2018 Frequent Sequence Pattern Mining with Differential Privacy
Fengli Zhou, Xiaoli Lin
ICIC (1)2
2018 New Particle Swarm Optimization Based on Sub-groups Mutation
Fengli Zhou, Xiaoli Lin
ICIC (3)2
2018 Discrete stationary wavelet transform based saliency information fusion from frequency and spatial domain in low contrast images
Nan Mu, Xin Xu 0007, Xiaolong Zhang 0002, Xiaoli Lin
Pattern Recognit. Lett.4
2018 Prediction of Hot Regions in PPIs Based on Improved Local Community Structure Detecting
abstract
The hot regions in PPIs are some assembly regions which are composed of the tightly packed HotSpots. The discovery of hot regions helps to understand life activities and has very important value for biological applications. The identification of hot regions is the basis for protein design and cancer prevention. The existing algorithms of predicting hot regions often have some defects, such as low accuracy and unstability. This paper proposes a novel hot region prediction method based on diverse biological characteristics. First, feature evaluation is employed by using an impoved mRMR method. Then, SVM is adopted to create cassification model based on the features selected. In addition, a new clustering algorithm, namely LCSD (Local community structure detecting), is developed to detect and analyze the conformation of hot regions. In the clustering process, the link similarity of protein residues is introduced to handle the boundary nodes. This algorithm can effectively deal with the missing residue nodes and control the local community boundaries. The results indicate that the spatial structure of hot regions can be obtained more effectively, and that our method is more effective than previous methods for precise identification of hot regions.
Xiaoli Lin, Xiaolong Zhang 0002
IEEE ACM Trans. Comput. Biol. Bioinform.1
2017 Protein Hot Regions Feature Research Based on Evolutionary Conservation
Jing Hu 0003, Xiaoli Lin, Xiaolong Zhang 0002
ICIC (2)2
2017 Effective Identification of Hot Spots in PPIs Based on Ensemble Learning
Xiaoli Lin, Fengli Zhou
ICIC (2)1
2017 Classification of Hub Protein and Analysis of Hot Regions in Protein-Protein Interactions
Xiaoli Lin, Xiaolong Zhang 0002, Jing Hu 0003
ICIC (2)1
2017 Similarity Comparison of 3D Protein Structure Based on Riemannian Manifold
Fengli Zhou, Xiaoli Lin
ICIC (2)2
2016 Prediction and analysis of hot region in protein-protein interactions
abstract
Proteins play a crucial role in every organism, which perform a vast amount of functions. The hot regions in protein-protein interactions consist of hot spot residues in protein-protein binding sites which are called interfaces, can help proteins to perform their biological function. Residue based computational prediction of hot regions might be useful to understand the molecular mechanism and is crucial in drug design and protein design. However, it is very challenging to identify the hot regions in protein-proteins. In this paper, we have proposed a support vector machine based on ensemble learning system for predicting hot spot residues, and predicted hot regions in protein-protein interactions. The efficiency of our method is analyzed in identifying hot spots and hot regions in protein-protein interactions and the results obtained are compared with the existing techniques. The results demonstrate that the proposed method is superior to identify the hot spots and hot regions in the protein interfaces.
Xiaoli Lin, Xiaolong Zhang 0002
BIBM1
2016 Identification of Hot Regions in Protein-Protein Interactions Based on Detecting Local Community Structure
Xiaoli Lin, Xiaolong Zhang 0002
ICIC (1)1
2016 Effective Protein Structure Prediction with the Improved LAPSO Algorithm in the AB Off-Lattice Model
Xiaoli Lin, Fengli Zhou, Huayong Yang
ICIC (1)1
2016 Tourism Network Comments Sentiment Analysis and Early Warning System Based on Ontology
Yanxia Yang, Xiaoli Lin
ICIC (2)2
2015 Identification of Hot Regions in Protein-Protein Interactions Based on SVM and DBSCAN
Xiaoli Lin, Huayong Yang
ICIC (2)1
2014 Protein folding structure optimization based on GAPSO algorithm in the off-lattice model
abstract
Predicting the spacial folding structure of a protein, given its sequence of amino acids, is one of the central problems in computational biology field. This paper studies the AB off-lattice model with two species of monomers, called hydrophobic (A) and hydrophilic (B). Based on this simplified model, the low energy configurations are searched by using the GAPSO. A kind of optimization about the mutation mechanism and the Euclidean interference mechanism are presented, where a novel local adjustment strategy is also used to enhance the searching ability of the global minimum within the AB off-lattice model. Starting from random conformations, the GAPSO method can find the low-energy conformation of the Fibonacci sequences and the real protein sequences. Compared with other optimization methods, the proposed novel method could converge to the lower energy folds. It appears that the proposed method can used for solving protein folding problem, which is based on the thermodynamic hypothesis.
Xiaoli Lin, Xiaolong Zhang 0002
BIBM1
2014 Research of Training Feedforward Neural Networks Based on Hybrid Chaos Particle Swarm Optimization-Back-Propagation
Fengli Zhou, Xiaoli Lin
ICIC (3)2
2014 Protein structure prediction with local adjust tabu search algorithm
abstract
BACKGROUND: Protein folding structure prediction is one of the most challenging problems in the bioinformatics domain. Because of the complexity of the realistic protein structure, the simplified structure model and the computational method should be adopted in the research. The AB off-lattice model is one of the simplification models, which only considers two classes of amino acids, hydrophobic (A) residues and hydrophilic (B) residues. RESULTS: The main work of this paper is to discuss how to optimize the lowest energy configurations in 2D off-lattice model and 3D off-lattice model by using Fibonacci sequences and real protein sequences. In order to avoid falling into local minimum and faster convergence to the global minimum, we introduce a novel method (SATS) to the protein structure problem, which combines simulated annealing algorithm and tabu search algorithm. Various strategies, such as the new encoding strategy, the adaptive neighborhood generation strategy and the local adjustment strategy, are adopted successfully for high-speed searching the optimal conformation corresponds to the lowest energy of the protein sequences. Experimental results show that some of the results obtained by the improved SATS are better than those reported in previous literatures, and we can sure that the lowest energy folding state for short Fibonacci sequences have been found. CONCLUSIONS: Although the off-lattice models is not very realistic, they can reflect some important characteristics of the realistic protein. It can be found that 3D off-lattice model is more like native folding structure of the realistic protein than 2D off-lattice model. In addition, compared with some previous researches, the proposed hybrid algorithm can more effectively and more quickly search the spatial folding structure of a protein chain.
Xiaoli Lin, Xiaolong Zhang 0002, Fengli Zhou
BMC Bioinform.1
2013 3D Protein Structure Prediction with Local Adjust Tabu Search Algorithm
Xiaoli Lin, Fengli Zhou
ICIC (3)1