EDBT 2026 Demo / reviewers in the wild / expert
Xiangrong Liu
dblp:84/4552
· DBLP profile ↗
74ranked-venue papers
8as first author
50since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 50 · 5 first-author · 36 since 2021Artificial intelligence and machine learning · 13 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Probing conflict resolution between contextual and parametric knowledge in medical language models
Zhengqiu Yu, Yueping Ding, Xiangrong Liu |
Expert Syst. Appl. | 3 |
| 2026 | Decoupled Spatiotemporal Forecasting from Extreme Sparse Observations via Quantized Latent SpaceabstractPredicting spatiotemporal fields governed by partial differential equations (PDEs) from sparse sensor data is a critical and long-standing challenge in science and engineering. Recent deep learning approaches, particularly neural operators, have shown considerable promise in solving PDEs. However, their performance degrades significantly in the demanding regime of extreme sparsity, characterized by spatial sensor coverage of less than 1% and limited temporal observations. To overcome this limitation, we propose a novel framework that decouples the task into two stages: spatial reconstruction and temporal extrapolation. In the first stage, rather than reconstructing the high-dimensional physical field directly, our model learns to reconstruct the complete latent features from sparse observations—features that would otherwise be extracted from a dense field. This process is stabilized by a Vector Quantization (VQ) bottleneck, which discretizes the latent space. In the second stage, a decoder-only Transformer performs temporal extrapolation by autoregressively predicting the future sequence of these discrete latent indices. This design inherently allows the model to generalize to new initial conditions and varying forecast horizons, akin to standard autoregressive models. We validate our framework on three challenging benchmarks, achieving state-of-the-art (SOTA) performance under severe sparsity constraints. Furthermore, we introduce a challenging benchmark dataset based on fire dynamics simulations. On this benchmark, our model successfully forecasts the field's evolution 30 frames into the future from a single timeframe with less than 0.1% spatial observations—a result that pushes well beyond the capabilities of existing methods. Zhongnan Weng, Jiayi Que, Xiangrong Liu |
AAAI | 6 |
| 2026 | Heterogeneous Property-Cross-Aware Relation Networks for Molecular Representation LearningabstractAccurate prediction of molecular properties is a central task in drug discovery and materials science, yet it is often constrained by the scarcity of labeled samples, leading to the challenging problem of few-shot molecular property prediction (FS-MPP). To address this issue, existing meta-learning approaches have made notable progress but still suffer from two critical limitations: (i) the reliance on homogeneous graph representations, which neglect higher-level chemical semantics such as pharmacophores; and (ii) the lack of adaptive relational reasoning tailored to different molecules and tasks. In this work, we propose a novel framework termed Heterogeneous Property-Cross-Aware Relation Network (HPCA). HPCA constructs a unified heterogeneous graph strictly limited to the training set, in which atoms, globally shared pharmacophores, and molecular properties are modeled as distinct types of nodes to prevent information leakage. Building upon this representation, HPCA incorporates an adaptive relational reasoning module and a cross-layer attention mechanism, enabling dynamic learning of critical interactions that determine molecular properties. Extensive experiments conducted on four public benchmarks—TOX21, SIDER, MUV, and ToxCast—under the standard episodic evaluation protocol of the few-shot learning community demonstrate that HPCA achieves statistically significant improvements over strong meta-learning baselines such as PAR and Meta-MGNN across diverse few-shot settings. Hongpeng Qiu, Yinghui Jiang, Lian Shen, Zheyi Cai, Xiangrong Liu |
ICIC | 5 |
| 2026 | HKD-CPI: high-order knowledge distillation enhanced inductive compound-protein interaction predictionabstractMOTIVATION: Accurately identifying compound-protein interactions (CPIs) is critical for accelerating drug discovery. Recent deep learning methods have achieved impressive results, yet they primarily focus on local structures and neighborhood information, often overlooking high-order interaction patterns shared among similar molecules. RESULTS: In this paper, we propose HKD-CPI, a high-order knowledge-enhanced inductive framework designed to improve generalization to unseen compound-protein pairs. Specifically, HKD-CPI introduces a molecular graph tokenization mechanism that aligns compound molecular graph features with token embeddings from sequence-pretrained large language models (LLMs), effectively infusing sequence-derived semantics into structural representations. To capture shared interaction patterns among functionally similar biomolecules, we construct a hypergraph-based representation to model high-order relationships between feature-similar compound/protein groups and their binding partners. Furthermore, a knowledge distillation strategy is further adopted to transfer high-order interaction knowledge from the hypergraph to a lightweight student model, enabling efficient and robust CPI prediction. Extensive experiments demonstrate that HKD-CPI outperforms existing state-of-the-art methods in inductive CPI prediction tasks. In particular, it achieves an average improvement of 4.94% in AUROC and 3.64% in AUPRC over the best-performing baseline across five benchmark datasets. AVAILABILITY AND IMPLEMENTATION: Our code and data are available at https://github.com/Hezy618/HKD-CPI. Zhongyu He, Xiangrong Liu, Yinghui Jiang, Junlin Xu, Shuting Jin, Leyi Wei, Youyu Wang |
Bioinform. | 2 |
| 2026 | SurfFold: a unified model for protein inverse folding by integrating surface and structural informationabstractMOTIVATION: Proteins play a crucial role in biological systems, and accurate protein sequence prediction is essential for applications such as drug discovery. Existing inverse folding models primarily rely on protein backbone structure information, overlooking the biochemical properties embedded in protein surface data that constrain its functionality, leading to limited prediction accuracy. RESULTS: We propose a novel inverse folding framework, SurfFold, which integrates both protein backbone structure and surface information for sequence prediction. Additionally, it incorporates side-chain structural information and its interaction with surface information. Then, we introduce a Representation alignment module to better fuze structure and surface Representations. Experimental results demonstrate that SurfFold achieves state-of-the-art performance on the CATH4.2 dataset, and additional experiments validate the effectiveness of the proposed modules. Moreover, the homologous structure inverse folding experiment also demonstrates that SurfFold possesses excellent capability in homologous protein design. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at https://github.com/jiudizhengf/SurfFold. Darong Li, Lian Shen, Meijia Song, Deyi Li, Xiangrong Liu |
Bioinform. | 6 |
| 2026 | Adaptive open-world learning for analyzing wafer map variations
Yinghang Wu, Lian Shen, Xiangrong Liu |
Expert Syst. Appl. | 4 |
| 2026 | Real-time hallucination detection and intervention in medical LLMs via calibrated hidden-state probes
Zhengqiu Yu, Xiangrong Liu |
J. Biomed. Informatics | 2 |
| 2026 | FKSUDDAPre: A drug-disease association prediction framework based on F-TEST feature selection and AMDKSU resampling with interpretability analysisabstractIn drug discovery and therapeutic research, the prediction of drug-disease associations (DDAs) holds significant scientific and clinical value. Drug molecules exert their effects by precisely identifying disease-related biological targets, systematically modulating the entire pharmacological process from absorption, distribution, and metabolism to final efficacy. Accurate prediction of drug-disease associations not only facilitates an in-depth understanding of molecular mechanisms of drug action but also provides critical theoretical foundations for drug repositioning and personalized medicine. While traditional prediction methods based on in vitro experiments and clinical statistics yield reliable results, they suffer from inherent drawbacks such as long development cycles, substantial resource consumption, and low throughput. In contrast, emerging machine learning techniques offer a promising solution to these bottlenecks, enabling the intelligent and efficient discovery of potential drug-disease association networks and significantly improving drug development efficiency. However, it is noteworthy that existing machine learning methods still face significant challenges in practical applications: the complexity of feature construction raises the threshold for data processing; data sparsity constrains the depth of information mining; and the pervasive issue of sample imbalance poses a severe challenge to the model's predictive accuracy and generalization performance. In this study, we developed an efficient and accurate framework for drug-disease association prediction named FKSUDDAPre. The model employs a multi-modal feature fusion strategy: on one hand, it leverages an ensemble of Mol2vec and K- BERT to deeply capture the semantic features of drug molecular fingerprints; on the other hand, it integrates Medical Subject Headings (MeSH) with DeepWalk to effectively reduce the dimensionality of disease features while preserving their relational structure. To address the class imbalance problem, FKSUDDAPre designed an optimization algorithm called AMDKSU, which combined clustering with an improved distance metric strategy, significantly enhancing the discriminative power of the sample set. For data processing, F-test was employed for feature importance ranking, effectively reducing data dimensionality and improving model generalization. For the predictive architecture, FKSUDDAPre proposed a novel ensemble framework composed of XGBoost, Decision Tree, Random Forest, and HyperFast. By employing a dynamic weight allocation strategy, this ensemble effectively harnesses the complementary strengths of these models to achieve significantly enhanced predictive performance. Rigorous validation demonstrated the system's outstanding performance across multiple evaluation metrics, with an average AUC of 0.9725, improving the AUC by approximately 3.88% compared to the best-performing baseline model. In the prediction of Alzheimer's disease and Parkinson's disease, 80% and 60% of the top 10 candidate drugs recommended by FKSUDDAPre, respectively, had been confirmed by literature, demonstrating the model's good practical application potential. Furthermore, we conducted a LIME-based feature importance analysis on the model's predictions, visualizing the correlations between features and the target variable to demonstrate the model's interpretability. A cross-platform, user-friendly visualization tool had also been developed using the PyQt5 framework. Yun Zuo 0001, Ge Hua, Xiangrong Liu, Xiangxiang Zeng, Zhaohong Deng |
PLoS Comput. Biol. | 5 |
| 2026 | A Robotic Laser Welding Seam Tracking Method for Complex 3-D Weld Trajectories Based on Geometric Structure AnalysisabstractSeam sampling with line-structured-light sensors in laser welding is prone to projection distortion due to the complex 3-D geometry of sheet-metal workpieces, which leads to viewpoint-dependent variations in the sampled seam profile and affects laser incidence direction planning. Moreover, stable target tracking control becomes challenging when the 3-D seam trajectory contains small curvature radii or right-angle transitions. To address these issues, a method is proposed that compensates for projection distortion via deep-learning-based local surface-normal estimation and achieves robust viewpoint control along complex 3-D seam trajectories using prior geometric approximation of the seam trajectory. For normal estimation, seam features are encoded as “point-vector pair” superpixels and extracted by a single-stage deep convolutional network, enabling efficient, robust detection under welding glare, spatter, smoke, and reflections while maintaining boundary stability. Using the “point-vector pair” representation, the laser incidence direction is adaptively adjusted with reference to estimated surface normals to reduce pose uncertainty caused by projection distortion. For viewpoint control, a feedforward leading-direction control algorithm, derived from geometric–numerical analysis of predefined trajectory, ensures stable perception and high-precision continuous tracking along tight curves and right-angle turns. Across two experiments involving different workpiece geometries, the method achieves mean absolute laser focal position errors of 0.16 and 0.20 mm and mean absolute 3-D laser incident-direction errors of 0.99$^\circ$and 1.58$^\circ$, validating its efficiency, stability, and accuracy for sheet-metal laser seam tracking. Zhaoqi Chu, Xiangrong Liu, Xuhui Que, Yawei Hu, Juan Liu 0003 |
IEEE Trans. Ind. Informatics | 2 |
| 2025 | LOHA: Direct Graph Spectral Contrastive Learning Between Low-Pass and High-Pass ViewsabstractSpectral Graph Neural Networks effectively handle graphs with different homophily levels, with low-pass filter mining feature smoothness and high-pass filter capturing differences. When these distinct filters could naturally form two opposite views for self-supervised learning, the commonalities between these counterparts for the same node remain unexplored, leading to suboptimal performance. In this paper, a simple yet effective self-supervised contrastive framework, LOHA, is proposed to address this gap. LOHA optimally leverages low-pass and high-pass views by embracing "harmony in diversity". Rather than solely maximizing the difference between these distinct views, which may lead to feature separation, LOHA harmonizes the diversity by treating the propagation of graph signals from both views as a composite feature. Specifically, a novel high-dimensional feature named spectral signal trend is proposed to serve as the basis for the composite feature, which remains relatively unaffected by changing filters and focuses solely on original feature differences. LOHA achieves an average performance improvement of 2.8% over runner-up models on 9 real-world datasets with varying homophily levels. Notably, LOHA even surpasses fully-supervised models on several datasets, which underscores the potential of LOHA in advancing the efficacy of spectral GNNs for diverse graph structures. Ziyun Zou, Yinghui Jiang, Lian Shen, Xiangrong Liu |
AAAI | 5 |
| 2025 | RETAIN: Reliable Topology Augmentation for both Heterophilic and Homophilic GraphsabstractCurrent graph topology augmentation methods are mostly static and heavily rely on the assumption of homophily, where connected nodes are presumed to share the same labels by default. Due to the complexity of real-world graphs, the underlying assumption is often disrupted, thus performance declines, demonstrating their limited adaptability. Although learnable methods flexibly change augmentation strategies based on data, ignorance of balancing consistency and diversity leads to suboptimal performance. This gap highlights the need for universally applicable graph augmentation strategies that ensure these two aspects, thereby enhancing model robustness. To address these challenges, we propose the RETAIN framework as an adaptive data augmentation method for graphs with different homophily levels. This method can dynamically adjust the graph structure based on a learned conditional distribution with the aid of a graph explainer, thus balancing the consistency and diversity of the augmented data. By framing graph topology modification and model refinement as a joint optimization problem, RETAIN facilitates concurrent learning from augmented data and model optimization. Empirical evaluations across diverse benchmarks on node classification tasks reveal that RETAIN can be effectively combined with other methods in a plug-and-play manner and consistently yields performance improvement across a diverse set of benchmarks for both homophilic and heterophilic graphs. Ziyun Zou, Lian Shen, Yanhao Li, Xiangrong Liu |
ICASSP | 6 |
| 2025 | MoHyperSynergy: A multimodal representation fusion method using hypergraph neural networks for drug synergy predictionabstractCombination therapy has been widely applied in the treatment of complex diseases due to its ability to improve efficacy and reduce drug resistance. In recent years, deep learning techniques have significantly enhanced the accuracy of computational models for predicting drug interactions. However, these methods either focus solely on the molecular structure of drugs and the gene expression of cell lines or consider only the interaction networks related to drugs and cell lines. To effectively explore the combined impact of the structural information of drugs and cell lines and their interaction information in networks on drug synergy prediction, we propose a novel multimodal representation fusion model using hypergraph neural networks, called MoHyperSynergy. MoHyperSynergy constructs a hypergraph based on multimodal representations using the K-nearest neighbors approach. It interacts multimodal representations at the local feature level through an outer product embedding layer and fuses multimodal representations at the global semantic level using a hypergraph neural network, effectively utilizing information from different modalities of drugs and cell lines. MoHyperSynergy simultaneously extracts initial representations from drug molecular structures, cell line gene expressions, and the embeddings of drugs and cell lines in biomedical knowledge graphs as different modalities, and pretrains the model’s representation extractors in a self-supervised manner. We conducted experiments on real-world public datasets and evaluated our model under various scenarios that may arise in practical applications. The experimental results show that MoHyperSynergy achieves excellent performance in predicting drug synergy and significantly outperforms the current state-of-the-art models. Zelong Chen, Yanni Xu, Menglong Zhang, Ruichao Hu, Juan Liu 0003, Xiangrong Liu |
IJCNN | 8 |
| 2025 | PipeQS: Pipeline-Based Adaptive Quantization and Staleness-Aware Distributed GNN Training System
Donghang Wu, Lian Shen, Changzhi Jiang, Yanhao Li, Xiangrong Liu |
ECML/PKDD (2) | 5 |
| 2025 | NABP-BERT: NANOBODY®-antigen binding prediction based on bidirectional encoder representations from transformers (BERT) architectureabstractAntibody-mediated immunity is crucial in the vertebrate immune system. Nanobodies, also known as VHH or single-domain antibodies (sdAbs), are emerging as promising alternatives to full-length antibodies due to their compact size, precise target selectivity, and stability. However, the limited availability of nanobodies (Nbs) for numerous antigens (Ags) presents a significant obstacle to their widespread application. Understanding the interactions between Nbs and Ags is essential for enhancing their binding affinities and specificities. Experimental identification of these interactions is often costly and time-intensive. To address this issue, we introduce NABP-BERT, a deep-learning model based on the BERT architecture, designed to predict NANOBODY®-Ag binding solely from sequence information. Furthermore, we have developed a general pretrained model with transfer capabilities suitable for protein-related tasks, including protein-protein interaction tasks. NABP-BERT focuses on the surrounding amino acid contexts and outperforms existing methods, achieving an AUROC of 0.986 and an AUPR of 0.985. Fatma S. Ahmed, Saleh K. H. Aly, Xiangrong Liu |
Briefings Bioinform. | 3 |
| 2025 | MlyPredCSED: based on extreme point deviation compensated clustering combined with cross-scale convolutional neural networks to predict multiple lysine sites in humanabstractIn post-translational modification, covalent bonds on lysine and attached chemical groups significantly change proteins' physical and chemical properties. They shape protein structures, enhance function and stability, and are vital for physiological processes, affecting health and disease through mechanisms like gene expression, signal transduction, protein degradation, and cell metabolism. Although lysine (K) modification sites are considered among the most common types of post-translational modifications in proteins, research on K-PTMs has largely overlooked the synergistic effects between different modifications and lacked the techniques to address the problem of sample imbalance. Based on this, the Extreme Point Deviation Compensated Clustering (EPDCC) Undersampling algorithm was proposed in this study and combined with Cross-Scale Convolutional Neural Networks (CSCNNs) to develop a novel computational tool, MlyPredCSED, for simultaneously predicting multiple lysine modification sites. MlyPredCSED employs Multi-Label Position-Specific Triad Amino Acid Propensity and the physicochemical properties of amino acids to enhance the richness of sequence information. To address the challenge of sample imbalance, the innovative EPDCC Undersampling technique was introduced to adjust the majority class samples. The model's training and testing phase relies on the advanced CSCNN framework. MlyPredCSED, through cross-validation and testing, outperformed existing models, especially in complex categories with multiple modification sites. This research not only provides an efficient method for the identification of lysine modification sites but also demonstrates its value in biological research and drug development. To facilitate efficient use of MlyPredCSED by researchers, we have specifically developed an accessible free web tool: http://www.mlypredcsed.com. Yun Zuo 0001, Xingze Fang, Jiankang Chen, Jiayi Ji, Xiangrong Liu, Xiangxiang Zeng, Zhaohong Deng, Hongwei Yin, Anjing Zhao |
Briefings Bioinform. | 7 |
| 2025 | Evaluating large language models for information extraction from gastroscopy and colonoscopy reports through multi-strategy prompting
Zhengqiu Yu, Lexin Fang, Yueping Ding, Yaozheng Cai, Xiangrong Liu |
J. Biomed. Informatics | 7 |
| 2025 | HyperACP: A cutting-edge hybrid framework for anticancer peptide classification via scalable feature extraction and adaptive neighbor-based synthesisabstractCancer remains a major contributor to global mortality, constituting a significant and escalating threat to human health. Anticancer peptides (ACPs) have emerged as promising therapeutic agents due to their specific mechanisms of action, pronounced tumor-targeting capability, and low toxicity. Nevertheless, traditional approaches for ACP identification are constrained by their reliance on shallow, hand-crafted sequence features, which fail to capture deeper semantic and structural characteristics. Moreover, such models exhibit limited robustness and interpretability when confronted with practical challenges such as severe class imbalance. To address these limitations, this study proposes HyperACP, an innovative framework for ACP recognition that integrates deep representation learning, adaptive sampling, and mechanistic interpretability. The framework leverages the ESMC protein language model to extract comprehensive sequence features and employs a novel adaptive algorithm, ANBS, to mitigate class imbalance at the decision boundary. For enhanced model transparency, SHAP-Res is incorporated to elucidate the contributions of individual residues to the final predictions. Comprehensive evaluations demonstrate that HyperACP consistently outperforms state-of-the-art methods across multiple datasets and validation protocols-including 10-fold cross-validation and independent test sets-according to metrics such as Accuracy (ACC), Sensitivity (SN), Specificity (SP), Matthews Correlation Coefficient (MCC), and Area Under the Curve (AUC). Furthermore, the model yields biologically interpretable results, pinpointing key residues (K, L, F, G) known to play pivotal roles in anticancer activity. These findings provide not only a robust predictive tool (available at www.hyperacp.com) but also novel insights into the structure-function relationships underlying ACPs. Bangyi Zhang, Yun Zuo 0001, Jiayue Liu, Xiangrong Liu, Xiangxiang Zeng, Zhaohong Deng |
PLoS Comput. Biol. | 5 |
| 2024 | FlexMem: Proactive Memory Deduplication for Qcow2-Based VMs with Virtual Persistent MemoryabstractVirtualization-based cloud computing has gained popularity owing to its system-isolation security and near-native performance. However, a critical challenge for cloud computing arises from the presence of numerous redundant memory pages across independent virtual machines (VMs), significantly reducing memory efficiency. For instance, each VM typically loads data from its virtual disk into its memory without considering the existing data in other VMs’ memory. The state-of-the-art solution for addressing memory efficiency is kernel same-page merging (KSM), which performs memory deduplication by merging pages with identical data from different processes on the host. However, KSM operates reactively after memory redundancy has occurred, adding computational overhead due to periodic scanning, rendering it inefficient for cloud computing environments with high-density VMs. Xiangrong Liu, Yiming Zhang 0003 |
APNet | 3 |
| 2024 | LMGAN: A Progressive End-to-End Chinese Landscape Painting Generation ModelabstractAI Generated Content (AIGC) has caught the attention of researchers around the world. Very good results have been achieved in the generation of Western art painting, but little research has been done on the generation of Chinese Landscape Painting (CLP). In addition, the quality of CLP generated by existing studies is still not high enough, they only imitate the color and often lack of details in their scenes. To overcome those weaknesses, we propose a progressive end-to-end Long Memory Adversarial Generation Network (LMGAN) to generate CLP. First, we create a new high-quality CLP dataset. Then we design a memory module and a mapping module to capture more low-dimensional features, as well as a layer extraction module to improve the layers of the generated results. Experiments show that LMGAN can generate richer and more detailed CLP compared to former methods. Yingqian Zhang 0002, Shaomin Xie, Xiangrong Liu |
IJCNN | 3 |
| 2024 | MSlocPRED: deep transfer learning-based identification of multi-label mRNA subcellular localizationabstractSubcellular localization of messenger ribonucleic acid (mRNA) is a universal mechanism for precise and efficient control of the translation process. Although many computational methods have been constructed by researchers for predicting mRNA subcellular localization, very few of these computational methods have been designed to predict subcellular localization with multiple localization annotations, and their generalization performance could be improved. In this study, the prediction model MSlocPRED was constructed to identify multi-label mRNA subcellular localization. First, the preprocessed Dataset 1 and Dataset 2 are transformed into the form of images. The proposed MDNDO-SMDU resampling technique is then used to balance the number of samples in each category in the training dataset. Finally, deep transfer learning was used to construct the predictive model MSlocPRED to identify subcellular localization for 16 classes (Dataset 1) and 18 classes (Dataset 2). The results of comparative tests of different resampling techniques show that the resampling technique proposed in this study is more effective in preprocessing for subcellular localization. The prediction results of the datasets constructed by intercepting different NC end (Both the 5' and 3' untranslated regions that flank the protein-coding sequence and influence mRNA function without encoding proteins themselves.) lengths show that for Dataset 1 and Dataset 2, the prediction performance is best when the NC end is intercepted by 35 nucleotides, respectively. The results of both independent testing and five-fold cross-validation comparisons with established prediction tools show that MSlocPRED is significantly better than established tools for identifying multi-label mRNA subcellular localization. Additionally, to understand how the MSlocPRED model works during the prediction process, SHapley Additive exPlanations was used to explain it. The predictive model and associated datasets are available on the following github: https://github.com/ZBYnb1/MSlocPRED/tree/main. Yun Zuo 0001, Bangyi Zhang, Wenying He, Yue Bi, Xiangrong Liu, Xiangxiang Zeng, Zhaohong Deng |
Briefings Bioinform. | 5 |
| 2024 | Surface-based multimodal protein-ligand binding affinity predictionabstractMOTIVATION: In the field of drug discovery, accurately and effectively predicting the binding affinity between proteins and ligands is crucial for drug screening and optimization. However, current research primarily utilizes representations based on sequence or structure to predict protein-ligand binding affinity, with relatively less study on protein surface information, which is crucial for protein-ligand interactions. Moreover, when dealing with multimodal information of proteins, traditional approaches typically concatenate features from different modalities in a straightforward manner without considering the heterogeneity among them, which results in an inability to effectively exploit the complementary between modalities. RESULTS: We introduce a novel multimodal feature extraction (MFE) framework that, for the first time, incorporates information from protein surfaces, 3D structures, and sequences, and uses cross-attention mechanism for feature alignment between different modalities. Experimental results show that our method achieves state-of-the-art performance in predicting protein-ligand binding affinity. Furthermore, we conduct ablation studies that demonstrate the effectiveness and necessity of protein surface information and multimodal feature alignment within the framework. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at https://github.com/Sultans0fSwing/MFE. Lian Shen, Menglong Zhang, Changzhi Jiang, Yanni Xu, Xiangrong Liu |
Bioinform. | 8 |
| 2024 | MoleMCL: a multi-level contrastive learning framework for molecular pre-trainingabstractMOTIVATION: Molecular representation learning plays an indispensable role in crucial tasks such as property prediction and drug design. Despite the notable achievements of molecular pre-training models, current methods often fail to capture both the structural and feature semantics of molecular graphs. Moreover, while graph contrastive learning has unveiled new prospects, existing augmentation techniques often struggle to retain their core semantics. To overcome these limitations, we propose a gradient-compensated encoder parameter perturbation approach, ensuring efficient and stable feature augmentation. By merging enhancement strategies grounded in attribute masking and parameter perturbation, we introduce MoleMCL, a new MOLEcular pre-training model based on multi-level contrastive learning. RESULTS: Experimental results demonstrate that MoleMCL adeptly dissects the structure and feature semantics of molecular graphs, surpassing current state-of-the-art models in molecular prediction tasks, paving a novel avenue for molecular modeling. AVAILABILITY AND IMPLEMENTATION: The code and data underlying this work are available in GitHub at https://github.com/BioSequenceAnalysis/MoleMCL. Yanni Xu, Changzhi Jiang, Lian Shen, Xiangrong Liu |
Bioinform. | 5 |
| 2024 | EPI-Trans: an effective transformer-based deep learning model for enhancer promoter interaction predictionabstractBACKGROUND: Recognition of enhancer-promoter Interactions (EPIs) is crucial for human development. EPIs in the genome play a key role in regulating transcription. However, experimental approaches for classifying EPIs are too expensive in terms of effort, time, and resources. Therefore, more and more studies are being done on developing computational techniques, particularly using deep learning and other machine learning techniques, to address such problems. Unfortunately, the majority of current computational methods are based on convolutional neural networks, recurrent neural networks, or a combination of them, which don't take into consideration contextual details and the long-range interactions between the enhancer and promoter sequences. A new transformer-based model called EPI-Trans is presented in this study to overcome the aforementioned limitations. The multi-head attention mechanism in the transformer model automatically learns features that represent the long interrelationships between enhancer and promoter sequences. Furthermore, a generic model is created with transferability that can be utilized as a pre-trained model for various cell lines. Moreover, the parameters of the generic model are fine-tuned using a particular cell line dataset to improve performance. RESULTS: Based on the results obtained from six benchmark cell lines, the average AUROC for the specific, generic, and best models is 94.2%, 95%, and 95.7%, while the average AUPR is 80.5%, 66.1%, and 79.6% respectively. CONCLUSIONS: This study proposed a transformer-based deep learning model for EPI prediction. The comparative results on certain cell lines show that EPI-Trans outperforms other cutting-edge techniques and can provide superior performance on the challenge of recognizing EPI. Fatma S. Ahmed, Saleh K. H. Aly, Xiangrong Liu |
BMC Bioinform. | 3 |
| 2024 | A heterogeneous graph neural network with automatic discovery of effective metapaths for drug-target interaction prediction
Menglong Zhang, Lian Shen, Yanni Xu, Xiangrong Liu |
Future Gener. Comput. Syst. | 8 |
| 2024 | PreMLS: The undersampling technique based on ClusterCentroids to predict multiple lysine sitesabstractThe translated protein undergoes a specific modification process, which involves the formation of covalent bonds on lysine residues and the attachment of small chemical moieties. The protein's fundamental physicochemical properties undergo a significant alteration. The change significantly alters the proteins' 3D structure and activity, enabling them to modulate key physiological processes. The modulation encompasses inhibiting cancer cell growth, delaying ovarian aging, regulating metabolic diseases, and ameliorating depression. Consequently, the identification and comprehension of post-translational lysine modifications hold substantial value in the realms of biological research and drug development. Post-translational modifications (PTMs) at lysine (K) sites are among the most common protein modifications. However, research on K-PTMs has been largely centered on identifying individual modification types, with a relative scarcity of balanced data analysis techniques. In this study, a classification system is developed for the prediction of concurrent multiple modifications at a single lysine residue. Initially, a well-established multi-label position-specific triad amino acid propensity algorithm is utilized for feature encoding. Subsequently, PreMLS: a novel ClusterCentroids undersampling algorithm based on MiniBatchKmeans was introduced to eliminate redundant or similar major class samples, thereby mitigating the issue of class imbalance. A convolutional neural network architecture was specifically constructed for the analysis of biological sequences to predict multiple lysine modification sites. The model, evaluated through five-fold cross-validation and independent testing, was found to significantly outperform existing models such as iMul-kSite and predML-Site. The results presented here aid in prioritizing potential lysine modification sites, facilitating subsequent biological assays and advancing pharmaceutical research. To enhance accessibility, an open-access predictive script has been crafted for the multi-label predictive model developed in this study. Yun Zuo 0001, Xingze Fang, Jiayong Wan, Wenying He, Xiangrong Liu, Xiangxiang Zeng, Zhaohong Deng |
PLoS Comput. Biol. | 5 |
| 2024 | Regularity-constrained point cloud reconstruction of building models via global alignment
Juan Cao 0002, Xiangrong Liu, Zhonggui Chen |
Vis. Comput. | 3 |
| 2023 | RareDR: A Drug Repositioning Approach for Rare Diseases Based on Knowledge Graph
Yuehan Huang, Shuting Jin, Changzhi Jiang, Zhengqiu Yu, Xiangrong Liu, Shaohui Huang |
ICIC (3) | 6 |
| 2023 | MGCPI: A Multi-granularity Neural Network for Predicting Compound-Protein Interactions
Peixuan Lin, Likun Jiang, Fatma S. Ahmed, Xinru Ruan, Xiangrong Liu, Juan Liu 0003 |
ICIC (3) | 5 |
| 2023 | Attention-Aware Contrastive Learning for Predicting Peptide-HLA Binding Specificity
Pengyu Luo, Yuehan Huang, Lian Shen, Xiangrong Liu |
ICIC (3) | 6 |
| 2023 | Fast Accurate Fish Recognition with Deep Learning Based on a Domain-Specific Large-Scale Fish Dataset
Zhaoqi Chu, Jari Korhonen, Xiangrong Liu, Juan Liu 0003, Lvping Fang, Weidi Yang, Debasish Ghose, Junyong You |
MMM (1) | 5 |
| 2023 | Uncertainty-Confidence Fused Pseudo-labeling for Graph Neural Networks
Pingjiang Long, Zihao Jian, Xiangrong Liu |
PRCV (9) | 3 |
| 2023 | MSGCL: inferring miRNA-disease associations based on multi-view self-supervised graph structure contrastive learningabstractPotential miRNA-disease associations (MDA) play an important role in the discovery of complex human disease etiology. Therefore, MDA prediction is an attractive research topic in the field of biomedical machine learning. Recently, several models have been proposed for this task, but their performance limited by over-reliance on relevant network information with noisy graph structure connections. However, the application of self-supervised graph structure learning to MDA tasks remains unexplored. Our study is the first to use multi-view self-supervised contrastive learning (MSGCL) for MDA prediction. Specifically, we generated a learner view without association labels of miRNAs and diseases as input, and utilized the known association network to generate an anchor view that provides guiding signals for the learner view. The graph structure was optimized by designing a contrastive loss to maximize the consistency between the anchor and learner views. Our model is similar to a pre-trained model that continuously optimizes upstream tasks for high-quality association graph topology, thereby enhancing the latent representation of association predictions. The experimental results show that our proposed method outperforms state-of-the-art methods by 2.79$\%$ and 3.20$\%$ in area under the receiver operating characteristic curve (AUC) and area under the precision/recall curve (AUPR), respectively. Xinru Ruan, Changzhi Jiang, Peixuan Lin, Juan Liu 0003, Shaohui Huang, Xiangrong Liu |
Briefings Bioinform. | 7 |
| 2023 | Chemical structure-aware molecular image representation learningabstractCurrent methods of molecular image-based drug discovery face two major challenges: (1) work effectively in absence of labels, and (2) capture chemical structure from implicitly encoded images. Given that chemical structures are explicitly encoded by molecular graphs (such as nitrogen, benzene rings and double bonds), we leverage self-supervised contrastive learning to transfer chemical knowledge from graphs to images. Specifically, we propose a novel Contrastive Graph-Image Pre-training (CGIP) framework for molecular representation learning, which learns explicit information in graphs and implicit information in images from large-scale unlabeled molecules via carefully designed intra- and inter-modal contrastive learning. We evaluate the performance of CGIP on multiple experimental settings (molecular property prediction, cross-modal retrieval and distribution similarity), and the results show that CGIP can achieve state-of-the-art performance on all 12 benchmark datasets and demonstrate that CGIP transfers chemical knowledge in graphs to molecular images, enabling image encoder to perceive chemical structures in images. We hope this simple and effective framework will inspire people to think about the value of image for molecular representation learning. Hongxin Xiang, Shuting Jin, Xiangrong Liu, Xiangxiang Zeng |
Briefings Bioinform. | 3 |
| 2023 | A general hypergraph learning algorithm for drug multi-task predictions in micro-to-macro biomedical networksabstractThe powerful combination of large-scale drug-related interaction networks and deep learning provides new opportunities for accelerating the process of drug discovery. However, chemical structures that play an important role in drug properties and high-order relations that involve a greater number of nodes are not tackled in current biomedical networks. In this study, we present a general hypergraph learning framework, which introduces Drug-Substructures relationship into Molecular interaction Networks to construct the micro-to-macro drug centric heterogeneous network (DSMN), and develop a multi-branches HyperGraph learning model, called HGDrug, for Drug multi-task predictions. HGDrug achieves highly accurate and robust predictions on 4 benchmark tasks (drug-drug, drug-target, drug-disease, and drug-side-effect interactions), outperforming 8 state-of-the-art task specific models and 6 general-purpose conventional models. Experiments analysis verifies the effectiveness and rationality of the HGDrug model architecture as well as the multi-branches setup, and demonstrates that HGDrug is able to capture the relations between drugs associated with the same functional groups. In addition, our proposed drug-substructure interaction networks can help improve the performance of existing network models for drug-related prediction tasks. Shuting Jin, Yinghui Jiang, Leyi Wei, Zhuohang Yu, Xiangxiang Zeng, Xiangrong Liu |
PLoS Comput. Biol. | 9 |
| 2023 | Bioactive Peptide Recognition Based on NLP Pre-Train AlgorithmabstractBioactive peptides are defined as peptide sequences within a protein that can regulate important bodily functions through their myriad activities. With the development of machine learning, more computational methods were proposed for bioactive peptides recognition so that this task does not only rely on tedious and time-consuming wet-experiment. But the training and testing process of existing models are limited to small datasets, which affects model performance. Inspired by the success of sequence classification in natural language processing with unlabeled data, we proposed a pre-training method for Bioactive peptides recognition. By pre-trained with large-scale of protein sequences, our method achieved the best performance in multiple functional peptides identification including anti-cancer, anti-diabetic, anti-hypertensive, anti-inflammatory and anti-microbial peptides. Compared with the advanced model, our model's precision, coverage, accuracy and absolute true are improved by 7.2%, 6.9%, 6.1% and 4.2% in the result of 5-fold cross-validation. In addition, the results indicate the model has superior prediction performance in single functional peptides recognition, especially for anti-cancer peptides and anti-microbial peptides which with longer sequences. Likun Jiang, Xiangrong Liu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2023 | KGNMDA: A Knowledge Graph Neural Network Method for Predicting Microbe-Disease AssociationsabstractAccumulated studies discovered that various microbes in human bodies were closely related to complex human diseases and could provide new insight into drug development. Multiple computational methods were constructed to predict microbes that were potentially associated with diseases. However, most previous methods were based on single characteristics of microbes or diseases, that lacked important biological information related to microorganisms or diseases. Therefore, we constructed a knowledge graph centered on microorganisms and diseases from several existed databases to provide knowledgeable information for microbes and diseases. Then, we adopted a graph neural network method to learn representations of microbes and diseases from the constructed knowledge graph. After that, we introduced the Gaussian kernel similarity features of microbes and diseases to generate final representations of microbes and diseases. At last, we proposed a score function on final representations of microbes and diseases to predict scores of microbe-disease associations. Comprehensive experiments on the Human Microbe-Disease Association Database (HMDAD) dataset had demonstrated that our approach outperformed baseline methods. Furthermore, we implemented case studies on two important diseases (asthma and inflammatory bowel disease), the result demonstrated that our proposed model was effective in revealing the relationship between diseases and microbes. The source code of our model and the data were available on https://github.com/ChangzhiJiang/KGNMDA_master. Changzhi Jiang, Minli Tang, Shuting Jin, Xiangrong Liu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2022 | Drug-target interactions prediction via deep collaborative filtering with multiembeddingsabstractDrug-target interactions (DTIs) prediction research presents important significance for promoting the development of modern medicine and pharmacology. Traditional biochemical experiments for DTIs prediction confront the challenges including long time period, high cost and high failure rate, and finally leading to a low-drug productivity. Chemogenomic-based computational methods can realize high-throughput prediction. In this study, we develop a deep collaborative filtering prediction model with multiembeddings, named DCFME (deep collaborative filtering prediction model with multiembeddings), which can jointly utilize multiple feature information from multiembeddings. Two different representation learning algorithms are first employed to extract heterogeneous network features. DCFME uses the generated low-dimensional dense vectors as input, and then simulates the drug-target relationship from the perspective of both couplings and heterogeneity. In addition, the model employs focal loss that concentrates the loss on sparse and hard samples in the training process. Comparative experiments with five baseline methods show that DCFME achieves more significant performance improvement on sparse datasets. Moreover, the model has better robustness and generalization capacity under several harder prediction scenarios. Ruolan Chen, Feng Xia 0007, Shuting Jin, Xiangrong Liu |
Briefings Bioinform. | 5 |
| 2022 | DeepTTA: a transformer-based model for predicting cancer drug responseabstractIdentifying new lead molecules to treat cancer requires more than a decade of dedicated effort. Before selected drug candidates are used in the clinic, their anti-cancer activity is generally validated by in vitro cellular experiments. Therefore, accurate prediction of cancer drug response is a critical and challenging task for anti-cancer drugs design and precision medicine. With the development of pharmacogenomics, the combination of efficient drug feature extraction methods and omics data has made it possible to use computational models to assist in drug response prediction. In this study, we propose DeepTTA, a novel end-to-end deep learning model that utilizes transformer for drug representation learning and a multilayer neural network for transcriptomic data prediction of the anti-cancer drug responses. Specifically, DeepTTA uses transcriptomic gene expression data and chemical substructures of drugs for drug response prediction. Compared to existing methods, DeepTTA achieved higher performance in terms of root mean square error, Pearson correlation coefficient and Spearman's rank correlation coefficient on multiple test sets. Moreover, we discovered that anti-cancer drugs bortezomib and dactinomycin provide a potential therapeutic option with multiple clinical indications. With its excellent performance, DeepTTA is expected to be an effective method in cancer drug design. Likun Jiang, Changzhi Jiang, Shuting Jin, Xiangrong Liu |
Briefings Bioinform. | 6 |
| 2022 | preMLI: a pre-trained method to uncover microRNA-lncRNA potential interactionsabstractThe interaction between microribonucleic acid and long non-coding ribonucleic acid plays a very important role in biological processes, and the prediction of the one is of great significance to the study of its mechanism of action. Due to the limitations of traditional biological experiment methods, more and more computational methods are applied to this field. However, the existing methods often have problems, such as inadequate acquisition of potential features of the sequence due to simple coding and the need to manually extract features as input. We propose a deep learning model, preMLI, based on rna2vec pre-training and deep feature mining mechanism. We use rna2vec to train the ribonucleic acid (RNA) dataset and to obtain the RNA word vector representation and then mine the RNA sequence features separately and finally concatenate the two feature vectors as the input of the prediction task. The preMLI performs better than existing methods on benchmark datasets and has cross-species prediction capabilities. Experiments show that both pre-training and deep feature mining mechanisms have a positive impact on the prediction performance of the model. To be more specific, pre-training can provide more accurate word vector representations. The deep feature mining mechanism also improves the prediction performance of the model. Meanwhile, The preMLI only needs RNA sequence as the input of the model and has better cross-species prediction performance than the most advanced prediction models, which have reference value for related research. Likun Jiang, Shuting Jin, Xiangxiang Zeng, Xiangrong Liu |
Briefings Bioinform. | 5 |
| 2022 | MLysPRED: graph-based multi-view clustering and multi-dimensional normal distribution resampling techniques to predict multiple lysine sitesabstractPosttranslational modification of lysine residues, K-PTM, is one of the most popular PTMs. Some lysine residues in proteins can be continuously or cascaded covalently modified, such as acetylation, crotonylation, methylation and succinylation modification. The covalent modification of lysine residues may have some special functions in basic research and drug development. Although many computational methods have been developed to predict lysine PTMs, up to now, the K-PTM prediction methods have been modeled and learned a single class of K-PTM modification. In view of this, this study aims to fill this gap by building a multi-label computational model that can be directly used to predict multiple K-PTMs in proteins. In this study, a multi-label prediction model, MLysPRED, is proposed to identify multiple lysine sites using features generated from human protein sequences. In MLysPRED, three kinds of multi-label sequence encoding algorithms (MLDBPB, MLPSDAAP, MLPSTAAP) are proposed and combined with three encoding strategies (CHHAA, DR and Kmer) to convert preprocessed lysine sequences into effective numerical features. A multidimensional normal distribution oversampling technique and graph-based multi-view clustering under-sampling algorithm were first proposed and incorporated to reduce the proportion of the original training samples, and multi-label nearest neighbor algorithm is used for classification. It is observed that MLysPRED achieved an Aiming of 92.21%, Coverage of 94.98%, Accuracy of 89.63%, Absolute-True of 81.46% and Absolute-False of 0.0682 on the independent datasets. Additionally, comparison of results with five existing predictors also indicated that MLysPRED is very promising and encouraging to predict multiple K-PTMs in proteins. For the convenience of the experimental scientists, 'MLysPRED' has been deployed as a user-friendly web-server at http://47.100.136.41:8181. Yun Zuo 0001, Xiangxiang Zeng, Qiang Zhang 0026, Xiangrong Liu |
Briefings Bioinform. | 5 |
| 2022 | LaGAT: link-aware graph attention network for drug-drug interaction predictionabstractMOTIVATION: Drug-drug interaction (DDI) prediction is a challenging problem in pharmacology and clinical applications. With the increasing availability of large biomedical databases, large-scale biological knowledge graphs containing drug information have been widely used for DDI prediction. However, large knowledge graphs inevitably suffer from data noise problems, which limit the performance and interpretability of models based on the knowledge graph. Recent studies attempt to improve models by introducing inductive bias through an attention mechanism. However, they all only depend on the topology of entity nodes independently to generate fixed attention pathways, without considering the semantic diversity of entity nodes in different drug pair links. This makes it difficult for models to select more meaningful nodes to overcome data quality limitations and make more interpretable predictions. RESULTS: To address this issue, we propose a Link-aware Graph Attention method for DDI prediction, called LaGAT, which is able to generate different attention pathways for drug entities based on different drug pair links. For a drug pair link, the LaGAT uses the embedding representation of one of the drugs as a query vector to calculate the attention weights, thereby selecting the appropriate topological neighbor nodes to obtain the semantic information of the other drug. We separately conduct experiments on binary and multi-class classification and visualize the attention pathways generated by the model. The results prove that LaGAT can better capture semantic relationships and achieves remarkably superior performance over both the classical and state-of-the-art models on DDI prediction. AVAILABILITYAND IMPLEMENTATION: The source code and data are available at https://github.com/Azra3lzz/LaGAT. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Pengyu Luo, Shuting Jin, Xiangrong Liu |
Bioinform. | 4 |
| 2022 | The challenges of explainable AI in biomedical data scienceabstractWith the surge of biomedical data science, more and more AI techniques are employed to discover knowledge, unveil latent data behavior, generate new insight, and seek optimal strategies in decision making.Different AI methods have been proposed and developed in almost all different biomedical data science fields that range from drug discovery, electronic medical records (EMRs) data automation, single-cell RNA sequencing, early disease diagnosis, COVID research, and healthcare analytics.The AI methods and systems also generate a massive amount of data or big data that not only bring unpreceded progress in biomedical fields but also new challenges for AI.One of the key challenges should be the explainability of AI in biomedical data science problem-solving.It refers to that an AI method or system should not only bring good results but also have good interpretability, i.e., let users know why this way is the optimal one rather than the others.The existing AI methods employed in biomedical data science generally lack good explainability and may not create trustworthiness and transparency in usage well.For example, a deep learning model may bring good accuracy in disease diagnosis by analyzing corresponding bioimages, but it can be hard to explain well about the setting of thousands of parameters in the model.It can be possible that some small perturbations of the parameters may generate totally different learning results and challenge the robustness and stability of the deep learning model.Since AI models cannot explain themselves well, it is likely to encounter a high risk to make an incorrect decision making and decrease its trustworthiness and reliability, even if it has the advantage in accuracy, speed, or complicate data relationship revealing.On the other hand, the AI interpretation issue has been raised almost ten years ago in some subfield of biomedical data science such as bioinformatics.For instance, bioinformaticians found that the gene markers or network markers recommended from an AI disease diagnosis system may not explain themselves, i.e., the identified markers not only cannot apply themselves well in clinical practice, but also those markers that do well in the clinical practice may not be recommended from the AI system [1]. Henry Han, Xiangrong Liu |
BMC Bioinform. | 2 |
| 2022 | A Deep Ensemble Predictor for Identifying Anti-Hypertensive Peptides Using Pretrained Protein EmbeddingabstractHypertension (HT), or high blood pressure is one of the most common and main causes in cardiovascular diseases, which is also related to a series of detrimental diseases in humans. Deficiencies in effective treatment in HT are often associated with a series of diseases including multi-infarct dementia, amputation, and renal failure. Therefore, identifying anti-hypertension peptides has the vital realistic significance. Although many bioactive peptides have been developed to reduce blood pressure, they are time-consuming and laborious. In views of the obstacles of the intrinsic methods in antihypertensive peptide (AHTP) classification, computational methods are suggested as a supplement to identify AHTPs. In this study, we develop a comprehensive feature representation algorithm based on pretrained model and convolutional neural network and apply the deep ensemble model to construct the prediction model. The new predictor is used to identify AHTPs in benchmark and independent datasets. It has been shown in the independent test set that the performance is better than the recent methods. Comparative results indicate that our model can shed some light on hypertension therapy and gains more insights of classifying AHTPs. The implements and codes can be found in https://github.com/yuanying566/AHPred-DE. Yuanying Zhuang, Xiangrong Liu, Longxin Wu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2021 | A Meta-Path based Drug-Target Prediction Model with Collaborative Attention MechanismsabstractThe discovery and confirmation o f d rug-target interactions plays an important role for multiple pharmacology, drug discovery, drug repositioning, side effect prediction and drug resistance. In this paper, we propose a meta-path-based collaborative attention prediction model that effectively learns the explicit representation of drugs, targets, and meta-path contexts to predict the potential relationships between drug and target. Experimental results prove that the model has an excellent performance in predictability and interpretability. Feng Xia 0007, Ruolan Chen, Shuting Jin, Xiangrong Liu |
BIBM | 5 |
| 2021 | Pm6 A: an Integrated Classification Algorithm for 2021 IEEE International Conference on Bioinformatics and Biomedicine (BIBM) Identifying m6 A SitesabstractAs a major RNA methylation modification, $\mathrm{N}^{6}_{-}$ methyladenosine (m6A) affects the occurrence and development of various human cancers through a variety of mechanisms. It has been reported that m6A RNA methylation involves different physiological and pathological processes. Therefore, the detection of m6A is helpful to reveal its biological function. Due to the high cost and time-consuming of high-throughput sequencing and the inaccurate sites identified, computational tools are needed to guide the accurate prediction of m6A modified sites and help reduce the costs associated with high-throughput sequencing. In this study, an integrated classification algorithm, Pm6A, is proposed to identify S. cerevisiae m6A sites using features generated from RNA sequences. In Pm6A, six sequence encoding schemes (pseudo dinucleotide composition, dinucleotide-based auto covariance, dinucleotide-based cross covariance, dinucleotide-based auto-cross covariance, mismatch and subsequence) are used for feature extraction, the VIBES ensemble algorithm based on four base classifiers, namely nearest neighbor, support vector machine, discriminant analysis and artificial neural network is used for classification, the optimized forward search algorithm is used to find the optimal parameters of model. The results of 10-fold cross-validation show that the proposed approach achieved better specificity (at Dataset 1) and better accuracy (at Dataset 2) than other methods. It is expected that Pm6A will be a useful tool for predicting m6A sites. Yun Zuo 0001, Xiangrong Liu, Xiangxiang Zeng |
BIBM | 2 |
| 2021 | Application of deep learning methods in biological networksabstractThe increase in biological data and the formation of various biomolecule interaction databases enable us to obtain diverse biological networks. These biological networks provide a wealth of raw materials for further understanding of biological systems, the discovery of complex diseases and the search for therapeutic drugs. However, the increase in data also increases the difficulty of biological networks analysis. Therefore, algorithms that can handle large, heterogeneous and complex data are needed to better analyze the data of these network structures and mine their useful information. Deep learning is a branch of machine learning that extracts more abstract features from a larger set of training data. Through the establishment of an artificial neural network with a network hierarchy structure, deep learning can extract and screen the input information layer by layer and has representation learning ability. The improved deep learning algorithm can be used to process complex and heterogeneous graph data structures and is increasingly being applied to the mining of network data information. In this paper, we first introduce the used network data deep learning models. After words, we summarize the application of deep learning on biological networks. Finally, we discuss the future development prospects of this field. Shuting Jin, Xiangxiang Zeng, Feng Xia 0007, Xiangrong Liu |
Briefings Bioinform. | 5 |
| 2021 | Predicting enhancer-promoter interactions by deep learning and matching heuristicabstractEnhancer-promoter interactions (EPIs) play an important role in transcriptional regulation. Recently, machine learning-based methods have been widely used in the genome-scale identification of EPIs due to their promising predictive performance. In this paper, we propose a novel method, termed EPI-DLMH, for predicting EPIs with the use of DNA sequences only. EPI-DLMH consists of three major steps. First, a two-layer convolutional neural network is used to learn local features, and an bidirectional gated recurrent unit network is used to capture long-range dependencies on the sequences of promoters and enhancers. Second, an attention mechanism is used for focusing on relatively important features. Finally, a matching heuristic mechanism is introduced for the exploration of the interaction between enhancers and promoters. We use benchmark datasets in evaluating and comparing the proposed method with existing methods. Comparative results show that our model is superior to currently existing models in multiple cell lines. Specifically, we found that the matching heuristic mechanism introduced into the proposed model mainly contributes to the improvement of performance in terms of overall accuracy. Additionally, compared with existing models, our model is more efficient with regard to computational speed. Xiaoping Min, Congmin Ye, Xiangrong Liu, Xiangxiang Zeng |
Briefings Bioinform. | 3 |
| 2021 | A novel antibacterial peptide recognition algorithm based on BERTabstractAs the best substitute for antibiotics, antimicrobial peptides (AMPs) have important research significance. Due to the high cost and difficulty of experimental methods for identifying AMPs, more and more researches are focused on using computational methods to solve this problem. Most of the existing calculation methods can identify AMPs through the sequence itself, but there is still room for improvement in recognition accuracy, and there is a problem that the constructed model cannot be universal in each dataset. The pre-training strategy has been applied to many tasks in natural language processing (NLP) and has achieved gratifying results. It also has great application prospects in the field of AMP recognition and prediction. In this paper, we apply the pre-training strategy to the model training of AMP classifiers and propose a novel recognition algorithm. Our model is constructed based on the BERT model, pre-trained with the protein data from UniProt, and then fine-tuned and evaluated on six AMP datasets with large differences. Our model is superior to the existing methods and achieves the goal of accurate identification of datasets with small sample size. We try different word segmentation methods for peptide chains and prove the influence of pre-training steps and balancing datasets on the recognition effect. We find that pre-training on a large number of diverse AMP data, followed by fine-tuning on new data, is beneficial for capturing both new data's specific features and common features between AMP sequences. Finally, we construct a new AMP dataset, on which we train a general AMP recognition model. Jianyuan Lin, Lianmin Zhao, Xiangxiang Zeng, Xiangrong Liu |
Briefings Bioinform. | 5 |
| 2021 | CarSite-II: an integrated classification algorithm for identifying carbonylated sites based on K-means similarity-based undersampling and synthetic minority oversampling techniquesabstractBACKGROUND: Carbonylation is a non-enzymatic irreversible protein post-translational modification, and refers to the side chain of amino acid residues being attacked by reactive oxygen species and finally converted into carbonyl products. Studies have shown that protein carbonylation caused by reactive oxygen species is involved in the etiology and pathophysiological processes of aging, neurodegenerative diseases, inflammation, diabetes, amyotrophic lateral sclerosis, Huntington's disease, and tumor. Current experimental approaches used to predict carbonylation sites are expensive, time-consuming, and limited in protein processing abilities. Computational prediction of the carbonylation residue location in protein post-translational modifications enhances the functional characterization of proteins. RESULTS: In this study, an integrated classifier algorithm, CarSite-II, was developed to identify K, P, R, and T carbonylated sites. The resampling method K-means similarity-based undersampling and the synthetic minority oversampling technique (SMOTE-KSU) were incorporated to balance the proportions of K, P, R, and T carbonylated training samples. Next, the integrated classifier system Rotation Forest uses "support vector machine" subclassifications to divide three types of feature spaces into several subsets. CarSite-II gained Matthew's correlation coefficient (MCC) values of 0.2287/0.3125/0.2787/0.2814, False Positive rate values of 0.2628/0.1084/0.1383/0.1313, False Negative rate values of 0.2252/0.0205/0.0976/0.0608 for K/P/R/T carbonylation sites by tenfold cross-validation, respectively. On our independent test dataset, CarSite-II yield MCC values of 0.6358/0.2910/0.4629/0.3685, False Positive rate values of 0.0165/0.0203/0.0188/0.0094, False Negative rate values of 0.1026/0.1875/0.2037/0.3333 for K/P/R/T carbonylation sites. The results show that CarSite-II achieves remarkably better performance than all currently available prediction tools. CONCLUSION: The related results revealed that CarSite-II achieved better performance than the currently available five programs, and revealed the usefulness of the SMOTE-KSU resampling approach and integration algorithm. For the convenience of experimental scientists, the web tool of CarSite-II is available in http://47.100.136.41:8081/. Yun Zuo 0001, Jianyuan Lin, Xiangxiang Zeng, Quan Zou 0001, Xiangrong Liu |
BMC Bioinform. | 5 |
| 2021 | Neural-like P systems with plasmids
Francis George Cabarle, Xiangxiang Zeng, Niall Murphy, Tao Song 0001, Alfonso Rodríguez-Patón, Xiangrong Liu |
Inf. Comput. | 6 |
| 2020 | A multi-task learning method for analyzing microbiota as cancer immunotherapy signalabstractResearches have found that tumor immunotherapy can only work for some patients, and the intestinal microbiota is one of the important factors affecting the responses of patients with cancer to immune checkpoint blockade therapy. It is highly desirable to develop computational methods that can predict whether a patient with cancer will have positive effects on cancer immunotherapy by analyzing intestinal microorganisms of the patient. In this study, a multi-task model is introduced to predict the efficacy of cancer immunotherapy on a patient who suffered from non-small cell lung cancer or renal cell carcinoma. The results demonstrate the multi-task model outperforms several single-task methods. Therefore, we believe the multi-task idea can be used to predict the efficacy of cancer immunotherapy based on the gut microbe, which would be important to cancer patients. Changzhi Jiang, Yousi Fu, Shuting Jin, Xiangrong Liu, Baishan Fang, Xiangxiang Zeng |
BIBM | 4 |
| 2020 | Computational methods for identifying the critical nodes in biological networksabstractA biological network is complex. A group of critical nodes determines the quality and state of such a network. Increasing studies have shown that diseases and biological networks are closely and mutually related and that certain diseases are often caused by errors occurring in certain nodes in biological networks. Thus, studying biological networks and identifying critical nodes can help determine the key targets in treating diseases. The problem is how to find the critical nodes in a network efficiently and with low cost. Existing experimental methods in identifying critical nodes generally require much time, manpower and money. Accordingly, many scientists are attempting to solve this problem by researching efficient and low-cost computing methods. To facilitate calculations, biological networks are often modeled as several common networks. In this review, we classify biological networks according to the network types used by several kinds of common computational methods and introduce the computational methods used by each type of network. Xiangrong Liu, Zengyan Hong, Juan Liu 0003, Alfonso Rodríguez-Patón, Quan Zou 0001, Xiangxiang Zeng |
Briefings Bioinform. | 1 |
| 2020 | Sequence clustering in bioinformatics: an empirical studyabstractSequence clustering is a basic bioinformatics task that is attracting renewed attention with the development of metagenomics and microbiomics. The latest sequencing techniques have decreased costs and as a result, massive amounts of DNA/RNA sequences are being produced. The challenge is to cluster the sequence data using stable, quick and accurate methods. For microbiome sequencing data, 16S ribosomal RNA operational taxonomic units are typically used. However, there is often a gap between algorithm developers and bioinformatics users. Different software tools can produce diverse results and users can find them difficult to analyze. Understanding the different clustering mechanisms is crucial to understanding the results that they produce. In this review, we selected several popular clustering tools, briefly explained the key computing principles, analyzed their characters and compared them using two independent benchmark datasets. Our aim is to assist bioinformatics users in employing suitable clustering tools effectively to analyze big sequencing data. Related data, codes and software tools were accessible at the link http://lab.malab.cn/∼lg/clustering/. Quan Zou 0001, Xingpeng Jiang, Xiangrong Liu, Xiangxiang Zeng |
Briefings Bioinform. | 4 |
| 2020 | Identifying enhancer-promoter interactions with neural network based on pre-trained DNA vectors and attention mechanismabstractMOTIVATION: Identification of enhancer-promoter interactions (EPIs) is of great significance to human development. However, experimental methods to identify EPIs cost too much in terms of time, manpower and money. Therefore, more and more research efforts are focused on developing computational methods to solve this problem. Unfortunately, most existing computational methods require a variety of genomic data, which are not always available, especially for a new cell line. Therefore, it limits the large-scale practical application of methods. As an alternative, computational methods using sequences only have great genome-scale application prospects. RESULTS: In this article, we propose a new deep learning method, namely EPIVAN, that enables predicting long-range EPIs using only genomic sequences. To explore the key sequential characteristics, we first use pre-trained DNA vectors to encode enhancers and promoters; afterwards, we use one-dimensional convolution and gated recurrent unit to extract local and global features; lastly, attention mechanism is used to boost the contribution of key features, further improving the performance of EPIVAN. Benchmarking comparisons on six cell lines show that EPIVAN performs better than state-of-the-art predictors. Moreover, we build a general model, which has transfer ability and can be used to predict EPIs in various cell lines. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at: https://github.com/hzy95/EPIVAN. Zengyan Hong, Xiangxiang Zeng, Leyi Wei, Xiangrong Liu |
Bioinform. | 4 |
| 2020 | Multiobjective Particle Swarm Optimization Based on Network Embedding for Complex Network Community DetectionabstractCommunity detection in complex networks is significant to social network analysis. Most of the algorithms take advantage of single-objective optimization methods, which may not be effective for complex networks. Compared with single-objective algorithms, multiobjective evolutionary algorithms can avoid local optimization. However, multiobjective evolutionary algorithms often encounter problems of excessive search space and low efficiency. To solve these issues, this study introduces network embedding into the multiobjective particle swarm algorithm and maps nodes into a low-latitude space, thereby effectively reducing the search space while increasing search efficiency via a consensus propagation strategy. Experimental results demonstrate that a novel effective algorithm based on multiobjective particle swarm optimization (NE-PSO) performs effectively and has competitive performance in comparison with state-of-the-art approaches on synthetic and real-world networks, especially the large-scale ones. Xiangrong Liu, Yanzi Du, Min Jiang 0005, Xiangxiang Zeng |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2019 | A Deep Neural Network for Antimicrobial Peptide RecognitionabstractWith the widespread use of antibiotics, many bacteria have developed resistance. Antimicrobial peptides have broad applications in medicine because of their high antibacterial activity. In this paper, a neural network model is introduced to recognize and detect antimicrobial peptides. Our model consists of an embedded, convolutional, bidirectional LSTM, and full connection layers. The embedded layer is used to code different amino acid residues into different vectors. The convolutional layer and bidirectional LSTM extract peptide amino acid residue sequence information. The full connection layer maps the sequence information linearly to the interval from 0 to 1, as the peptide for the probability of antimicrobial peptides. Training and testing on several different datasets reveal that our model performs better than other proposed models. Jianyuan Lin, Xiangxiang Zeng, Yun Zuo 0001, Ying Ju 0002, Xiangrong Liu |
BIBM | 5 |
| 2019 | Two-Dimensional Imaging With Stationary Nonuniform Frequency Diverse Array TransmitterabstractThis paper proposes a simple two-dimensional imaging via stationary nonuniform frequency diverse array (FDA) transmitter by exploiting FDA both range and angle dependent transmit beampattern. The nonuniform FDA transmitter only generates range-dependent beampattern, so that two-dimensional (range-angle dimension) imaging of targets can be achieved without the employment of moving radar platform, by jointly utilizing the only range-dependent transmitter beampattern and general single-antenna or phased-array antenna. The effectiveness of the proposed method is verified by simulation results. Xiangrong Liu |
IGARSS | 1 |
| 2019 | PolSAR Image Edge Detection via Structure Tensor AnalysisabstractThe weighted structure tensor (ST) can be usually used on edge detection of multi-channel images, including polarimetric synthetic aperture radar (polSAR) images. The traditional weighting method is average weighting, each channel considered to provide the same edge information. But in fact different channels tend to contain different amount of edge information, so the traditional weighting method cannot be adopted to extract full edge information of the multi-channel images. This letter proposes a novel weighting method for ST, in which the weight of each channel is obtained by the eigenvalue-measured way. Experimental results show that the proposed method outperforms the traditional method on edge detection. Xiangrong Liu, Shunsheng Zhang, Wen-Qin Wang |
IGARSS | 1 |
| 2019 | A Novel Deceptive Jamming Method Via Frequency Diverse ArrayabstractDifferent from traditional single-element and phased-array antennas, frequency diverse array (FDA) uses a slight frequency offset across the elements to produce angle-, range- and even time-dependent transmit beampattern, which provides promising potentials to develop new jamming techniques on synthetic aperture radar (SAR) imaging. In this paper, we propose a novel deception jamming method against SAR imaging, which uses FDA to quickly generate false targets. Moreover, the position of the false target is controlled by the jammer, and the number of false targets is determined by the number of FDA elements. The proposed methods are validated by extensive simulation results with range-Doppler imaging algorithm. Shunsheng Zhang, Xiangrong Liu |
IGARSS | 4 |
| 2019 | deepDR: a network-based deep learning approach to in silico drug repositioningabstractMOTIVATION: Traditional drug discovery and development are often time-consuming and high risk. Repurposing/repositioning of approved drugs offers a relatively low-cost and high-efficiency approach toward rapid development of efficacious treatments. The emergence of large-scale, heterogeneous biological networks has offered unprecedented opportunities for developing in silico drug repositioning approaches. However, capturing highly non-linear, heterogeneous network structures by most existing approaches for drug repositioning has been challenging. RESULTS: In this study, we developed a network-based deep-learning approach, termed deepDR, for in silico drug repurposing by integrating 10 networks: one drug-disease, one drug-side-effect, one drug-target and seven drug-drug networks. Specifically, deepDR learns high-level features of drugs from the heterogeneous networks by a multi-modal deep autoencoder. Then the learned low-dimensional representation of drugs together with clinically reported drug-disease pairs are encoded and decoded collectively via a variational autoencoder to infer candidates for approved drugs for which they were not originally approved. We found that deepDR revealed high performance [the area under receiver operating characteristic curve (AUROC) = 0.908], outperforming conventional network-based or machine learning-based approaches. Importantly, deepDR-predicted drug-disease associations were validated by the ClinicalTrials.gov database (AUROC = 0.826) and we showcased several novel deepDR-predicted approved drugs for Alzheimer's disease (e.g. risperidone and aripiprazole) and Parkinson's disease (e.g. methylphenidate and pergolide). AVAILABILITY AND IMPLEMENTATION: Source code and data can be downloaded from https://github.com/ChengF-Lab/deepDR. SUPPLEMENTARY INFORMATION: Supplementary data are available online at Bioinformatics. Xiangxiang Zeng, Siyi Zhu, Xiangrong Liu, Yadi Zhou, Ruth Nussinov, Feixiong Cheng |
Bioinform. | 3 |
| 2019 | An advanced approach to identify antimicrobial peptides and their function types for penaeus through machine learning strategiesabstractBACKGROUND: Antimicrobial peptides (AMPs) are essential components of the innate immune system and can protect the host from various pathogenic bacteria. The marine environment is known to be one of the richest sources for AMPs. Effective usage of AMPs and their derivatives can greatly improve the immunity and breeding survival rate of aquatic products. It is highly desirable to develop computational tools for rapidly and accurately identifying AMPs and their functional types, for the purpose of helping design new and more effective antimicrobial agents. RESULTS: In this study, we made an attempt to develop an advanced machine learning based computational approach, MAMPs-Pred, for identification of AMPs and its function types. Initially, SVM-prot 188-D features were extracted that were subsequently used as input to a two-layer multi-label classifier. In specific, the first layer is to identify whether it is an AMP by applying RF classifier, and the second layer addresses the multi-type problem by identifying the activites or function types of AMPs by applying PS-RF and LC-RF classifiers. To benchmark the methods,the MAMPs-Pred method is also compared with existing best-performing methods in literature and has shown an improved identification accuracy. CONCLUSIONS: The results reported in this study indicate that the MAMP-Pred method achieves high performance for identifying AMPs and its functional types.The proposed approach is believed to supplement the tools and techniques that have been developed in the past for predicting AMPs and their function types. Yinyin Cai, Juan Liu 0003, Chen Lin 0001, Xiangrong Liu |
BMC Bioinform. | 5 |
| 2019 | On solutions and representations of spiking neural P systems with rules on synapses
Francis George Cabarle, Ren Tristan A. de la Cruz, Dionne Peter P. Cailipan, Xiangrong Liu, Xiangxiang Zeng |
Inf. Sci. | 5 |
| 2019 | Computational Prediction of Sigma-54 Promoters in Bacterial Genomes by Integrating Motif Finding and Machine Learning StrategiesabstractSigma factor, as a unit of RNA polymerase holoenzyme, is a critical factor in the process of gene transcriptional regulation. It recognizes the specific DNA sites and brings the core enzyme of RNA polymerase to the upstream regions of target genes. Therefore, the prediction of the promoters for a particular sigma factor is essential for interpreting functional genomic data and observation. This paper develops a new method to predict sigma-54 promoters in bacterial genomes. The new method organically integrates motif finding and machine learning strategies to capture the intrinsic features of sigma-54 promoters. The experiments on E. coli benchmark test set show that our method has good capability to distinguish sigma-54 promoters from surrounding or randomly selected DNA sequences. The applications of the other three bacterial genomes indicate the potential robustness and applicative power of our method on a large number of bacterial genomes. The source code of our method can be freely downloaded at https://github.com/maqin2001/PromotePredictor. Bingqiang Liu, Xiangrong Liu, Jichang Wu, Qin Ma 0003 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2017 | Deep-based fisher vector for mobile visual searchabstractWe tackle the problem of mobile visual search. Moving pictures experts group (MPEG) has completed a standard named compact descriptor for visual search (CDVS) to provide a standardized syntax in the context of image retrieval application. CDVS applies principal components analysis to reduce the dimension of local feature descriptor as the input of global descriptor pipeline, and utilizes traditional fisher vector as the local feature descriptor aggregation algorithm. However, the descriptor components of SIFT and Fisher Vector (FV) have highly non-Gaussian statistics, and applying a single PCA transform can in-fact hurt compression performance at high rates. We develop a net-based architecture combining neural networks with FV layer to obtain fisher vector. There are two advantages in our architecture comparing with CDVS global descriptor pipeline. One is that we employ “autoencoder” networks to reduce the dimensionality of data, the other is that we exploit a trainable system to learn parameters after the FV codebook obtained. The experiments demonstrate an obvious advantage of our proposed architecture in terms of CDVS retrieval task. Shengchuan Zhang, Xianming Lin, Xiangrong Liu, Rongrong Ji |
ICIP | 4 |
| 2016 | Investment behavior prediction in heterogeneous information network
Xiangxiang Zeng, Stephen C. H. Leung, Ziyu Lin, Xiangrong Liu |
Neurocomputing | 5 |
| 2016 | Question microblog identification and answer recommendation
Xiangrong Liu, Runquan Xie, Liujuan Cao |
Multim. Syst. | 1 |
| 2015 | The power of time-free tissue P systems: Attacking NP-complete problems
Xiangrong Liu, Juan Suo, Stephen C. H. Leung, Juan Liu 0003, Xiangxiang Zeng |
Neurocomputing | 1 |
| 2015 | Asynchronous spiking neural P systems with rules on synapses
Tao Song 0001, Quan Zou 0001, Xiangrong Liu, Xiangxiang Zeng |
Neurocomputing | 3 |
| 2015 | A uniform solution to integer factorization using time-free spiking neural P system
Xiangrong Liu, Ziming Li 0001, Juan Suo, Juan Liu 0003, Xiaoping Min |
Neural Comput. Appl. | 1 |
| 2015 | Asynchronous Spiking Neural P Systems with Anti-Spikes
Tao Song 0001, Xiangrong Liu, Xiangxiang Zeng |
Neural Process. Lett. | 2 |
| 2014 | Identifying Reliable Posts and Users in Online Social NetworksabstractIn the age of Web 2.0, generalizing concepts are mainly based on information sharing and user cooperation. The user-centric Internet mode encourages more users to participate in the Internet. However, this mode causes the flooding of social networks with low-quality information. Thus, identifying trusted information in online social networks (OSNs) so as to enable the Internet to serve humans better has gradually become a research hotspot. However, most studies evaluate the quality of user-created content or discriminate reliable users in isolation. These results are specific to particular social networks and lack generality. Intuitively, reliable content and users are closely related. In this study, the authors attempt to combine these two types of reliable data by separately identifying reliable posts and users in a social network. Posts and users are found to improve each other based on the preliminary identification results. To deal with imbalanced data, an algorithm that combines oversampling and undersampling is used to build balanced data. An ensemble classifier is adopted for data classification. Experiments show that the proposed framework is both effective and efficient for most types of OSNs. The contributions of this study are two-fold: (i) a framework combining user-created content and reliable user recognition is proposed, and (ii) an ensemble classifier is built for use in data classification. Sifa Xie, Wei Weng 0002, Xiangrong Liu |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2014 | On languages generated by spiking neural P systems with weights
Xiangxiang Zeng, Lei Xu 0002, Xiangrong Liu, Linqiang Pan |
Inf. Sci. | 3 |
| 2010 | A novel computing model of the maximum clique problem based on circular DNA
Cheng Zhang 0017, Xiangrong Liu, Xiaoli Qiang |
Sci. China Inf. Sci. | 4 |
| 2006 | Predicting Melting Temperature (Tm) of DNA Duplex Based on Neural Network
Xiangrong Liu, Juan Liu 0003, Linqiang Pan, Jin Xu 0015 |
ICIC (3) | 1 |