EDBT 2026 Demo / reviewers in the wild / expert
Liang Yu 0002
dblp:28/1433-2
· DBLP profile ↗
30ranked-venue papers
11as first author
26since 2021 · last 2026
0000-0002-8351-3332ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 26 · 10 first-author · 23 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpaMFG: a spatial multi-omics integration method based on feature groupingabstractMOTIVATION: The rapid development of spatial multi-omics technology enables the simultaneous measurement of gene and protein expression alongside spatial location, providing valuable insights into tissue heterogeneity. However, challenges such as low spatial resolution and high feature dimensionality complicate data integration and biological interpretation. RESULTS: To address these issues, we propose SpaMFG, an innovative feature-group-level framework for interpretable spatial multi-omics integration. SpaMFG leverages spatial location information and introduces a spatial proximity weighting method to improve feature grouping accuracy. Additionally, it employs a new cross-omics feature group matching method that combines spatial location and Jaccard similarity to construct a weighted cost matrix, which is optimized using the Hungarian algorithm. This approach enhances the biological interpretability of cross-omics feature relationships. We evaluated SpaMFG's performance through comparative analysis on the human lymph node dataset, demonstrating its effectiveness. Further applications on human tonsils, mouse spleens, and mouse thymus datasets confirmed the robustness of SpaMFG in various biological contexts. AVAILABILITY AND IMPLEMENTATION: The source code for SpaMFG is available at https://github.com/LiangYu-Xidian/SpaMFG. Zilin Li, Litian Ma, Jingtao Liu, Liang Yu 0002 |
Bioinform. | 7 |
| 2026 | Enhancing cross-context generalization in drug perturbation prediction with a multimodal conditional diffusion frameworkabstractMOTIVATION: Predicting drug-induced transcriptional perturbations is critical for precision medicine, yet existing models fail to capture multimodal biological context, limiting generalization across unseen drugs and cell lines. RESULTS: We present PertDiff, a conditional diffusion framework that integrates control gene expression, LLM-derived cell semantics, and pretrained molecular graph representations to predict transcriptome-wide perturbations. PertDiff outperforms state-of-the-art baselines in prediction accuracy and generalizes robustly across drugs and cell lines. It further demonstrates translational utility through accurate drug sensitivity prediction, therapeutic repurposing for pancreatic cancer, and concordance with real-world clinical treatment outcomes, establishing it as a biologically grounded transcriptomic modeling tool. AVAILABILITY: The source code and data are available at https://github.com/Panda-myj/PertDiff and https://doi.org/10.5281/zenodo.18427848. Yanjie Ma, Pengyong Li, Liang Yu 0002 |
Bioinform. | 5 |
| 2026 | Predicting enhancer-promoter interactions using a stacking-based ensemble strategyabstractMOTIVATION: Enhancer-promoter interactions (EPIs) are essential for gene regulation and disease progression. Recent studies have shown that distal enhancers can regulate target genes through interactions with nearby promoters, providing important insights into transcriptional regulation mechanisms. Although high-throughput experimental techniques have enabled large-scale identification of EPIs, these methods are often costly and time-consuming. In addition, existing computational approaches still face challenges in effectively integrating heterogeneous feature representations from different cell lines. RESULTS: We propose a stacked ensemble framework for EPI prediction that integrates feature representations from diverse cell line datasets using multiple machine learning algorithms. The extracted complementary patterns are further combined by an XGBoost classifier to improve robustness against overfitting. Experiments on six independent datasets show that the proposed method achieves superior accuracy and generalization compared with existing EPI prediction models, with an average AUROC of 0.909 while maintaining computational efficiency. AVAILABILITY: The source code and its archived release are available at GitHub and Zenodo. The Zenodo archive provides a versioned snapshot of the repository: https://zenodo.org/records/19952998. Zhichao Xiao, Haibo Ji, Quan Zou 0001, Yijie Ding, Liang Yu 0002 |
Bioinform. | 5 |
| 2025 | MT-IDR: Disordered Flexible Linkers and Molecular Recognition Features Prediction Based on Multi-Task LearningabstractWith the continuous development of research on intrinsically disordered proteins, experimental functional annotation methods provide researchers with some functional annotations of intrinsically disordered regions(IDRs), shifting research on the functions of IDRs toward data-driven approaches. Disordered Flexible Linkers(DFLs) and Molecular Recognition Features(MoRFs) are two of the most annotated functions. Moreover, the amount of labeled data is too small, and it is difficult to build a sequence-level model. Therefore, for DFLs and MoRFs prediction, we proposed MT-IDR (intrinsically disordered region function prediction based on multi-task learning), a solution based on sequence hierarchy with multi-task learning. First, based on the amino acid sequence of the protein, evolutionary information, physicochemical properties and structural properties were calculated, then the enhanced protein representation module was used to improve the features. Second, the enhanced protein representation is fed into a multi-task learning model for training. In the independent test sets of two functional prediction tasks, MT-IDR has excellent performance. In the DFLs prediction task, the AUC values of both the TE82 and TE64 reached 0.790 and 0.812. In the MoRFs prediction task, the EXP53 achieves the best performance of single models. The higher performance of MT-IDR provides the opportunity for large-scale screening of DFLs and MoRFs. Liang Yu 0002, Haozheng Li, JingTao Liu, Yuchuan Peng, Bin Liu 0014 |
BIBM | 1 |
| 2025 | DeCoST: unveiling cell type heterogeneity in spatial transcriptomics based on inter-domain alignment and Gaussian kernel conditional autoregressiveabstractSpatial transcriptomics (STs) has emerged as a transformative approach to elucidate cellular heterogeneity and spatial organization within complex tissue microenvironments. However, the analysis of ST data is challenged by limited spatial resolution, resulting in mixed expression profiles at each spatial location. Moreover, the precious spatial information is rarely exploited, and noise issues in spatial transcriptomes (STs) are often overlooked by computational deconvolution methods. In this study, a novel computational framework for STs deconvolution (DeCoST), called DeCoST, is presented. DeCoST capitalizes on the valuable spatial context information by integrating a Gaussian kernel-based conditional autoregressive model. Additionally, the method employs domain adaptation techniques to address platform effects between single-cell and ST data, enabling robust cell type identification. Evaluations on simulated datasets under diverse spatial configurations, as well as real-world case studies on human pancreatic ductal adenocarcinoma, mouse olfactory bulb, and mouse brain samples, demonstrate the superior performance of DeCoST compared to existing deconvolution approaches. The method's ability to accurately map region-specific cell types and uncover spatial interactions advances our understanding of complex tissue organization and function, with broad applications in disease research and developmental biology. Xinyang Guo, Zilin Li, Liang Yu 0002 |
Briefings Bioinform. | 6 |
| 2025 | FORAlign: accelerating gap-affine DNA pairwise sequence alignment using FOR-blocks based on Four Russians approach with linear space complexityabstractPairwise sequence alignment (PSA) serves as the cornerstone in computational bioinformatics, facilitating multiple sequence alignment and phylogenetic analysis. This paper introduces the FORAlign algorithm, leveraging the Four Russians algorithm with identical upper-bound time and space complexity as the Hirschberg divide-and-conquer PSA algorithm, aimed at accelerating Hirschberg PSA algorithm in parallel. Particularly notable is its capability to achieve up to 16.79 times speedup when aligning sequences with low sequence similarity, compared to the conventional Needleman-Wunsch PSA method using non-heuristic methods. Empirical evaluations underscore FORAlign's superiority over existing wavefront alignment (WFA) series software, especially in scenarios characterized by low sequence similarity during PSA tasks. Our method is capable of directly aligning monkeypox sequences with other sequences using non-heuristic methods. The algorithm was implemented within the FORAlign library, providing functionality for PSA and foundational support for multiple sequence alignment and phylogenetic trees. The FORAlign library is freely available at https://github.com/malabz/FORAlign. Yanming Wei, Tong Zhou 0016, Yixiao Zhai, Liang Yu 0002, Quan Zou 0001 |
Briefings Bioinform. | 4 |
| 2025 | EPIPDLF: a pretrained deep learning framework for predicting enhancer-promoter interactionsabstractMOTIVATION: Enhancers and promoters, as regulatory DNA elements, play pivotal roles in gene expression, homeostasis, and disease development across various biological processes. With advancing research, it has been uncovered that distal enhancers may engage with nearby promoters to modulate the expression of target genes. This discovery holds significant implications for deepening our comprehension of various biological mechanisms. In recent years, numerous high-throughput wet-lab techniques have been created to detect possible interactions between enhancers and promoters. However, these experimental methods are often time-intensive and costly. RESULTS: To tackle this issue, we have created an innovative deep learning approach, EPIPDLF, which utilizes advanced deep learning techniques to predict EPIs based solely on genomic sequences in an interpretable manner. Comparative evaluations across six benchmark datasets demonstrate that EPIPDLF consistently exhibits superior performance in EPI prediction. Additionally, by incorporating interpretable analysis mechanisms, our model enables the elucidation of learned features, aiding in the identification and biological analysis of important sequences. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at: https://github.com/xzc196/EPIPDLF. Zhichao Xiao, Yijie Ding, Liang Yu 0002 |
Bioinform. | 4 |
| 2025 | Multiple kernel-based fuzzy system for identifying enhancers
Zhichao Xiao, Yijie Ding, Liang Yu 0002 |
Expert Syst. Appl. | 3 |
| 2025 | Computational approaches for predicting drug-disease associations: a comprehensive review
Zhichao Xiao, Chunyan Ao, Lixin Guan, Liang Yu 0002 |
Frontiers Comput. Sci. | 5 |
| 2024 | MVLncStack: A novel method for lncRNA subcellular localization prediction based on multi-view featuresabstractLncRNA performs different life activities in different subcellulars. For the subcellular multi-classification localization task of lncRNA, this paper proposed an ensemble learning prediction model named MVLncStack based on random forest and deep learning. Random forest in MVLncStack was employed to further analyze the sequence-level features of lncRNA, and selected the optimal feature subset with incremental feature selection. The deep learning model in MVLncStack was employed to learn the residue-level features of lncRNA with position encoding. RNA multi-view features provided the sequence information and structure information of lncRNA for MVLncStack, which can enhance the prediction performance of the model. Compared with the existing methods on the independent test set, the accuracy of multi classification prediction was improved by 5.2%, and other indicators were improved by about 1% - 2%. Through the performance evaluation and related experimental analysis, the importance of multi-view features of RNA was illustrated, and the prediction performance of the models can be effectively improved by selecting appropriate feature coding for RNA. Dongdong Jiang, Bin Liu 0014, Liang Yu 0002 |
BIBM | 5 |
| 2024 | MTTSCL: protein subcellular localization prediction using multi-task learningabstractProtein subcellular localization prediction is an important research area in bioinformatics. Traditionally, machine learning techniques have been utilized to model specific cell types for this task. However, the majority of these approaches solely focus on modeling a specific cell type. It is believed that there are potential similarities and shared semantic information among different biological types that can be leveraged for improved predictions. To address this issue, a deep learning approach based on multi-task learning for protein subcellular localization (Multi-Task Transferable Shared Concept Learning, MTTSCL) is proposed. This approach aims to overcome the limitations imposed by the scale of labeled data and effectively transfer knowledge between different cell types to predict protein subcellular localization. The results demonstrated that the multi-task learning approach outperformed single-task modeling. The joint modeling of diverse biological classifications led to substantial enhancements in the accuracy of protein subcellular localization prediction. In summary, the proposed deep learning approach using multi-task learning provides a promising solution for protein subcellular localization prediction. Liang Yu 0002 |
BIBM | 1 |
| 2024 | Prediction of cancer drug combinations based on multidrug learning and cancer expression information injection
Shujie Ren, Hongxia Hao, Liang Yu 0002 |
Future Gener. Comput. Syst. | 4 |
| 2023 | Application of Multi-Dimensional SNV Features in Supervised LearningabstractSingle nucleotide variation (SNV) is closely related to the occurrence of cancer, and effective extraction of SNV sample information will be beneficial for accurate early diagnosis of cancer. The main method of traditional research to extract SNV features is to combine SNV with two adjacent nucleotides to form a trinucleotide, and mutation features are extracted from the pattern of trinucleotides. However, single-dimensional feature extraction may lead to partial information loss and poor model performance. Therefore, we propose a method to extract single-nucleotide variation (SNV) features under multiple feature dimensions. We treat SNV as a one-dimensional feature, change the feature dimension by adding adjacent nucleotides, and achieve resampling. Simultaneously, we store multiple sets of features to ensure the integrity of the information carried by the SNV. We extend the method to the cancer marker identification application scenario to verify the improvement of the prediction performance by the extracted features. Using a dataset obtained from The Cancer Genome Atlas, based on six supervised learning algorithms, including KNN(K-Nearest Neighbors), SVM(Support Vector Machine), and random forest, we verified the feasibility of multidimensional SNV features. Compared with the original SNV feature extraction method, the feature extraction method proposed in this paper has significantly improved the prediction performance of the model. Additionally, this paper compares multi-dimensional features with the K-mer algorithm under the same dataset, and multi-dimensional features show obvious advantages in terms of acquisition time, storage space, and sample distinction Liang Yu 0002, Lin Gao 0006, Hongxia Hao |
BIBM | 1 |
| 2023 | Gene Expression Profile Prediction under Drug Action Based on Generative Adversarial NetworksabstractGene expression profiles play a significant role in drug research. If the gene expression profile under the action of drugs can be obtained quickly, such as through computational methods, the analysis of the relationship between the drug and the disease will become more comprehensive. The efficiency can be improved and costs can be reduced while exploring the effect of the drug. We developed an algorithm (ppc-GAN, predict-profile-conditional Generative Adversarial Networks) for predicting gene expression profiles for drug effects, which can efficiently and accurately obtain the gene expression profiles after drug administration. Compared with traditional algorithms, ppc-GAN does not require more prior knowledge. Therefore, the final prediction result will not be affected by the preference of prior knowledge. Our ppc-GAN mainly includes two parts—an autoencoder and a generative adversarial network (GAN). We trained the autoencoder through all gene expression profile data in the LINCS database and then merged the trained autoencoder into the GAN for data compression and decompression. Besides, we chose bortezomib as the case drug. Our results show that our model is flexible and has high representative power. Furthermore, the state of the gene expression profile after using the drug can be estimated by the deep learning models. Liang Yu 0002, Huan Zhu, Da Dong, Lin Gao 0006 |
BIBM | 1 |
| 2023 | Potent antibiotic design via guided search from antibacterial activity evaluationsabstractMOTIVATION: The emergence of drug-resistant bacteria makes the discovery of new antibiotics an urgent issue, but finding new molecules with the desired antibacterial activity is an extremely difficult task. To address this challenge, we established a framework, MDAGS (Molecular Design via Attribute-Guided Search), to optimize and generate potent antibiotic molecules. RESULTS: By designing the antibacterial activity latent space and guiding the optimization of functional compounds based on this space, the model MDAGS can generate novel compounds with desirable antibacterial activity without the need for extensive expensive and time-consuming evaluations. Compared with existing antibiotics, candidate antibacterial compounds generated by MDAGS always possessed significantly better antibacterial activity and ensured high similarity. Furthermore, although without explicit constraints on similarity to known antibiotics, these candidate antibacterial compounds all exhibited the highest structural similarity to antibiotics of expected function in the DrugBank database query. Overall, our approach provides a viable solution to the problem of bacterial drug resistance. AVAILABILITY AND IMPLEMENTATION: Code of the model and datasets can be downloaded from GitHub (https://github.com/LiangYu-Xidian/MDAGS). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Liang Yu 0002, Lin Gao 0006 |
Bioinform. | 2 |
| 2022 | NmRF: identification of multispecies RNA 2'-O-methylation modification sites from RNA sequencesabstract2'-O-methylation (Nm) is a post-transcriptional modification of RNA that is catalyzed by 2'-O-methyltransferase and involves replacing the H on the 2'-hydroxyl group with a methyl group. The 2'-O-methylation modification site is detected in a variety of RNA types (miRNA, tRNA, mRNA, etc.), plays an important role in biological processes and is associated with different diseases. There are few functional mechanisms developed at present, and traditional high-throughput experiments are time-consuming and expensive to explore functional mechanisms. For a deeper understanding of relevant biological mechanisms, it is necessary to develop efficient and accurate recognition tools based on machine learning. Based on this, we constructed a predictor called NmRF based on optimal mixed features and random forest classifier to identify 2'-O-methylation modification sites. The predictor can identify modification sites of multiple species at the same time. To obtain a better prediction model, a two-step strategy is adopted; that is, the optimal hybrid feature set is obtained by combining the light gradient boosting algorithm and incremental feature selection strategy. In 10-fold cross-validation, the accuracies of Homo sapiens and Saccharomyces cerevisiae were 89.069 and 93.885%, and the AUC were 0.9498 and 0.9832, respectively. The rigorous 10-fold cross-validation and independent tests confirm that the proposed method is significantly better than existing tools. A user-friendly web server is accessible at http://lab.malab.cn/∼acy/NmRF. Chunyan Ao, Quan Zou 0001, Liang Yu 0002 |
Briefings Bioinform. | 3 |
| 2022 | Multiview network embedding for drug-target Interactions prediction by consistent and complementary information preservingabstractAccurate prediction of drug-target interactions (DTIs) can reduce the cost and time of drug repositioning and drug discovery. Many current methods integrate information from multiple data sources of drug and target to improve DTIs prediction accuracy. However, these methods do not consider the complex relationship between different data sources. In this study, we propose a novel computational framework, called MccDTI, to predict the potential DTIs by multiview network embedding, which can integrate the heterogenous information of drug and target. MccDTI learns high-quality low-dimensional representations of drug and target by preserving the consistent and complementary information between multiview networks. Then MccDTI adopts matrix completion scheme for DTIs prediction based on drug and target representations. Experimental results on two datasets show that the prediction accuracy of MccDTI outperforms four state-of-the-art methods for DTIs prediction. Moreover, literature verification for DTIs prediction shows that MccDTI can predict the reliable potential DTIs. These results indicate that MccDTI can provide a powerful tool to predict new DTIs and accelerate drug discovery. The code and data are available at: https://github.com/ShangCS/MccDTI. Yifan Shang, Xiucai Ye, Yasunori Futamura, Liang Yu 0002, Tetsuya Sakurai |
Briefings Bioinform. | 4 |
| 2022 | A network embedding framework based on integrating multiplex network for drug combination predictionabstractDrug combination is a sensible strategy for disease treatment because it improves the treatment efficacy and reduces concomitant side effects. Due to the large number of possible combinations among candidate compounds, exhaustive screening is prohibitive. Currently, a large number of studies have focused on predicting potential drug combinations. However, these methods are not entirely satisfactory in terms of performance and scalability. In this paper, we proposed a Network Embedding frameWork in MultIplex Network (NEWMIN) to predict synthetic drug combinations. Based on a multiplex drug similarity network, we offered alternative methods to integrate useful information from different aspects and to decide the quantitative importance of each network. For drug combination prediction, we found seven novel drug combinations that have been validated by external sources among the top-ranked predictions of our model. To verify the feasibility of NEWMIN, we compared NEWMIN with other five methods, for which it showed better performance than other methods in terms of the area under the precision-recall curve and receiver operating characteristic curve. Liang Yu 0002, Mingfei Xia |
Briefings Bioinform. | 1 |
| 2022 | MiRNA-disease association prediction based on meta-pathsabstractSince miRNAs can participate in the posttranscriptional regulation of gene expression, they may provide ideas for the development of new drugs or become new biomarkers for drug targets or disease diagnosis. In this work, we propose an miRNA-disease association prediction method based on meta-paths (MDPBMP). First, an miRNA-disease-gene heterogeneous information network was constructed, and seven symmetrical meta-paths were defined according to different semantics. After constructing the initial feature vector for the node, the vector information carried by all nodes on the meta-path instance is extracted and aggregated to update the feature vector of the starting node. Then, the vector information obtained by the nodes on different meta-paths is aggregated. Finally, miRNA and disease embedding feature vectors are used to calculate their associated scores. Compared with the other methods, MDPBMP obtained the highest AUC value of 0.9214. Among the top 50 predicted miRNAs for lung neoplasms, esophageal neoplasms, colon neoplasms and breast neoplasms, 49, 48, 49 and 50 have been verified. Furthermore, for breast neoplasms, we deleted all the known associations between breast neoplasms and miRNAs from the training set. These results also show that for new diseases without known related miRNA information, our model can predict their potential miRNAs. Code and data are available at https://github.com/LiangYu-Xidian/MDPBMP. Liang Yu 0002, Lin Gao 0006 |
Briefings Bioinform. | 1 |
| 2022 | Research progress of miRNA-disease association prediction and comparison of related algorithmsabstractWith an in-depth understanding of noncoding ribonucleic acid (RNA), many studies have shown that microRNA (miRNA) plays an important role in human diseases. Because traditional biological experiments are time-consuming and laborious, new calculation methods have recently been developed to predict associations between miRNA and diseases. In this review, we collected various miRNA-disease association prediction models proposed in recent years and used two common data sets to evaluate the performance of the prediction models. First, we systematically summarized the commonly used databases and similarity data for predicting miRNA-disease associations, and then divided the various calculation models into four categories for summary and detailed introduction. In this study, two independent datasets (D5430 and D6088) were compiled to systematically evaluate 11 publicly available prediction tools for miRNA-disease associations. The experimental results indicate that the methods based on information dissemination and the method based on scoring function require shorter running time. The method based on matrix transformation often requires a longer running time, but the overall prediction result is better than the previous two methods. We hope that the summary of work related to miRNA and disease will provide comprehensive knowledge for predicting the relationship between miRNA and disease and contribute to advanced computation tools in the future. Liang Yu 0002, Bingyi Ju, Chunyan Ao, Lin Gao 0006 |
Briefings Bioinform. | 1 |
| 2022 | Multidrug representation learning based on pretraining model and molecular graph for drug interaction and combination predictionabstractMOTIVATION: Approaches for the diagnosis and treatment of diseases often adopt the multidrug therapy method because it can increase the efficacy or reduce the toxic side effects of drugs. Using different drugs simultaneously may trigger unexpected pharmacological effects. Therefore, efficient identification of drug interactions is essential for the treatment of complex diseases. Currently proposed calculation methods are often limited by the collection of redundant drug features, a small amount of labeled data and low model generalization capabilities. Meanwhile, there is also a lack of unique methods for multidrug representation learning, which makes it more difficult to take full advantage of the originally scarce data. RESULTS: Inspired by graph models and pretraining models, we integrated a large amount of unlabeled drug molecular graph information and target information, then designed a pretraining framework, MGP-DR (Molecular Graph Pretraining for Drug Representation), specifically for drug pair representation learning. The model uses self-supervised learning strategies to mine the contextual information within and between drug molecules to predict drug-drug interactions and drug combinations. The results achieved promising performance across multiple metrics compared with other state-of-the-art methods. Our MGP-DR model can be used to provide a reliable candidate set for the combined use of multiple drugs. AVAILABILITY AND IMPLEMENTATION: Code of the model, datasets and results can be downloaded from GitHub (https://github.com/LiangYu-Xidian/MGP-DR). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shujie Ren, Liang Yu 0002, Lin Gao 0006 |
Bioinform. | 2 |
| 2022 | WMSA: a novel method for multiple sequence alignment of DNA sequencesabstractMOTIVATION: Multiple sequence alignment (MSA) is a fundamental problem in bioinformatics. The quality of alignment will affect downstream analysis. MAFFT has adopted the Fast Fourier Transform method for searching the homologous segments and using them as anchors to divide the sequences, then making alignment only on segments, which can save time and memory without overly reducing the sequence alignment quality. MAFFT becomes slow when the dataset is large. RESULTS: We made a software, WMSA, which uses the divide-and-conquer method to split the sequences into clusters, aligns those clusters into profiles with the center star strategy and then makes a progressive profile-profile alignment. The alignment is conducted by the compiled algorithms of MAFFT, K-Band with multithread parallelism. Our method can balance time, space and quality and performs better than MAFFT in test experiments on highly conserved datasets. AVAILABILITY AND IMPLEMENTATION: Source code is freely available at https://github.com/malabz/WMSA/, which is implemented in C/C++ and supported on Linux, and datasets are available at https://github.com/malabz/WMSA-dataset. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yanming Wei, Quan Zou 0001, Furong Tang, Liang Yu 0002 |
Bioinform. | 4 |
| 2021 | A heterogeneous network embedding framework for predicting similarity-based drug-target interactionsabstractAccurate prediction of drug-target interactions (DTIs) through biological data can reduce the time and economic cost of drug development. The prediction method of DTIs based on a similarity network is attracting increasing attention. Currently, many studies have focused on predicting DTIs. However, such approaches do not consider the features of drugs and targets in multiple networks or how to extract and merge them. In this study, we proposed a Network EmbeDding framework in mulTiPlex networks (NEDTP) to predict DTIs. NEDTP builds a similarity network of nodes based on 15 heterogeneous information networks. Next, we applied a random walk to extract the topology information of each node in the network and learn it as a low-dimensional vector. Finally, the Gradient Boosting Decision Tree model was constructed to complete the classification task. NEDTP achieved accurate results in DTI prediction, showing clear advantages over several state-of-the-art algorithms. The prediction of new DTIs was also verified from multiple perspectives. In addition, this study also proposes a reasonable model for the widespread negative sampling problem of DTI prediction, contributing new ideas to future research. Code and data are available at https://github.com/LiangYu-Xidian/NEDTP. Liang Yu 0002 |
Briefings Bioinform. | 2 |
| 2021 | EPSOL: sequence-based protein solubility prediction using multidimensional embeddingabstractMOTIVATION: The heterologous expression of recombinant protein requires host cells, such as Escherichiacoli, and the solubility of protein greatly affects the protein yield. A novel and highly accurate solubility predictor that concurrently improves the production yield and minimizes production cost, and that forecasts protein solubility in an E.coli expression system before the actual experimental work is highly sought. RESULTS: In this article, EPSOL, a novel deep learning architecture for the prediction of protein solubility in an E.coli expression system, which automatically obtains comprehensive protein feature representations using multidimensional embedding, is presented. EPSOL outperformed all existing sequence-based solubility predictors and achieved 0.79 in accuracy and 0.58 in Matthew's correlation coefficient. The higher performance of EPSOL permits large-scale screening for sequence variants with enhanced manufacturability and predicts the solubility of new recombinant proteins in an E.coli expression system with greater reliability. AVAILABILITY AND IMPLEMENTATION: EPSOL's best model and results can be downloaded from GitHub (https://github.com/LiangYu-Xidian/EPSOL). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Liang Yu 0002 |
Bioinform. | 2 |
| 2021 | Prediction of drug-target interactions based on multi-layer network representation learning
Yifan Shang, Lin Gao 0006, Quan Zou 0001, Liang Yu 0002 |
Neurocomputing | 4 |
| 2021 | Predicting therapeutic drugs for hepatocellular carcinoma based on tissue-specific pathwaysabstractHepatocellular carcinoma (HCC) is a significant health problem worldwide with poor prognosis. Drug repositioning represents a profitable strategy to accelerate drug discovery in the treatment of HCC. In this study, we developed a new approach for predicting therapeutic drugs for HCC based on tissue-specific pathways and identified three newly predicted drugs that are likely to be therapeutic drugs for the treatment of HCC. We validated these predicted drugs by analyzing their overlapping drug indications reported in PubMed literature. By using the cancer cell line data in the database, we constructed a Connectivity Map (CMap) profile similarity analysis and KEGG enrichment analysis on their related genes. By experimental validation, we found securinine and ajmaline significantly inhibited cell viability of HCC cells and induced apoptosis. Among them, securinine has lower toxicity to normal liver cell line, which is worthy of further research. Our results suggested that the proposed approach was effective and accurate for discovering novel therapeutic options for HCC. This method also could be used to indicate unmarked drug-disease associations in the Comparative Toxicogenomics Database. Meanwhile, our method could also be applied to predict the potential drugs for other types of tumors by changing the database. Liang Yu 0002, Fengdan Xu, Lin Gao 0006, Xiangzhi Li |
PLoS Comput. Biol. | 1 |
| 2020 | C3: connect separate connected components to form a succinct disease moduleabstractBACKGROUND: Precise disease module is conducive to understanding the molecular mechanism of disease causation and identifying drug targets. However, due to the fragmentization of disease module in incomplete human interactome, how to determine connectivity pattern and detect a complete neighbourhood of disease based on this is still an open question. RESULTS: In this paper, we perform exploratory analysis leading to an important observation that through a few intermediate nodes, most separate connected components formed by disease-associated proteins can be effectively connected and eventually form a complete disease module. And based on the topological properties of these intermediate nodes, we propose a connect separate connected components (C3) method to detect a succinct disease module by introducing a relatively small number of intermediate nodes, which allows us to obtain more pure disease module than other methods. Then we apply C3 across a large corpus of diseases to validate this connectivity pattern of disease module. Furthermore, the connectivity of the perturbed genes in multi-omics data such as The Cancer Genome Atlas also fits this pattern. CONCLUSIONS: C3 tool is not only useful in detecting a clearly-defined connected disease neighbourhood of 299 diseases and cancer with multi-omics data, but also helpful in better understanding the interconnection of phenotypically related genes in different omics data and studying complex pathological processes. Bingbo Wang, Chenxing Zhang, Yuanjun Zhou, Liang Yu 0002, Xingli Guo, Lin Gao 0006, Yunru Chen |
BMC Bioinform. | 6 |
| 2019 | Human Pathway-Based Disease NetworkabstractConstructing disease-disease similarity network is important in elucidating the associations between the origin and molecular mechanism of diseases, and in researching disease function and medical research. In this paper, we use a high-quality protein interaction network and a collection of pathway databases to construct a Human Pathway-based Disease Network (HPDN) to explore the relationship between diseases and their intrinsic interactions. We find that the similarity of two diseases has a strong correlation with the number of their shared functional pathways and the interaction between their related gene sets. Comparing HPDN with disease networks based on genes and symptoms respectively, we find the three networks have high overlap rates. Additionally, HPDN can predict new disease-disease correlations, which are supported by Comparative Toxicogenomics Database (CTD) benchmark and large-scale biomedical literature database. The comprehensive, high-quality relations between diseases based on pathways can further be applied to study important matters in systems medicine, for instance, drug repurposing. Based on a dense subgraph in our network, we find two drugs, prednisone and folic acid, may have new indications, which will provide potential directions for the treatments of complex diseases. Liang Yu 0002, Lin Gao 0006 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2017 | Drug repositioning based on triangularly balanced structure for tissue-specific diseases in incomplete interactome
Liang Yu 0002, Lin Gao 0006 |
Artif. Intell. Medicine | 1 |
| 2017 | Prediction of Novel Drugs for Hepatocellular Carcinoma Based on Multi-Source Random WalkabstractComputational approaches for predicting drug-disease associations by integrating gene expression and biological network provide great insights to the complex relationships among drugs, targets, disease genes, and diseases at a system level. Hepatocellular carcinoma (HCC) is one of the most common malignant tumors with a high rate of morbidity and mortality. We provide an integrative framework to predict novel d rugs for HCC based on multi-source random walk (PD-MRW). Firstly, based on gene expression and protein interaction network, we construct a gene-gene weighted i nteraction network (GWIN). Then, based on multi-source random walk in GWIN, we build a drug-drug similarity network. Finally, based on the known drugs for HCC, we score all drugs in the drug-drug similarity network. The robustness of our predictions, their overlap with those reported in Comparative Toxicogenomics Database (CTD) and literatures, and their enriched KEGG pathway demonstrate our approach can effectively identify new drug indications. Specifically, regorafenib (Rank = 9 in top-20 list) is proven to be effective in Phase I and II clinical trials of HCC, and the Phase III trial is ongoing. And, it has 11 overlapping pathways with HCC with lower p-values. Focusing on a particular disease, we believe our approach is more accurate and possesses better scalability. Liang Yu 0002, Ruidan Su, Bingbo Wang, Yapeng Zou, Lin Gao 0006 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |