Yaohang Li

dblp:62/6371 · DBLP profile ↗
← Back
86ranked-venue papers
8as first author
39since 2021 · last 2025
0000-0003-0178-1876ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 63 · 2 first-author · 30 since 2021Artificial intelligence and machine learning · 13 · 8 since 2021Systems, architecture and hardware · 8 · 6 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Towards an Event-Level Analysis in Hadronic Physics Using Generative AI-Based Surrogates
abstract
The field of hadronic physics is poised for significant advancements, driven by upcoming experimental programs at Jefferson Lab's 12 GeV facility and the future Electron-Ion Collider at Brookhaven National Laboratory. These experiments will produce vast amounts of data, necessitating sophisticated analytical methods to bridge experimental observations and the underlying theoretical frameworks of quantum chromodynamics (QCD). This paper introduces a generative AI-based surrogate modeling framework to address the computational and analytical challenges of traditional theory-driven approaches in event-level analysis. Specifically, we evaluate and compare the effectiveness of generative AI models, including Generative Adversarial Networks (GANs) and Diffusion Models (DMs), as surrogates for event generation. By replacing complex quantum correlation function (QCF) models with generative AI-based surrogates, our proposed framework enables efficient generation of synthetic events directly from QCF parameters. Moreover, this approach facilitates simulation-based parameter inference by learning the mappings between QCF parameters and observables, bypassing the gradient calculation challenges when directly incorporating the QCF model into solving the QCD inverse problems. Through a proof-of-concept QCF study, we demonstrate the capability of generative AI-based surrogates as an important tool to accurately replicate theoretical QCF models, which can be easily incorporated into simulation-based inference to reconstruct QCF parameters from observables.
Tareq Alghamdi, Jitao Xu 0006, Nesar Ramachandra, Nobuo Sato, Yaohang Li
ICTAI5
2025 CATH-ddG: towards robust mutation effect prediction on protein-protein interactions out of CATH homologous superfamily
abstract
MOTIVATION: Protein-protein interactions (PPIs) are fundamental aspects in understanding biological processes. Accurately predicting the effects of mutations on PPIs remains a critical requirement for drug design and disease mechanistic studies. Recently, deep learning models using protein 3D structures have become predominant for predicting mutation effects. However, significant challenges remain in practical applications, in part due to the considerable disparity in generalization capabilities between easy and hard mutations. Specifically, a hard mutation is defined as one with its maximum TM-score <0.6 when compared to the training set. Additionally, compared to physics-based approaches, deep learning models may overestimate performance due to potential data leakage. RESULTS: We propose new training/test splits that mitigate data leakage according to the CATH homologous superfamily. Under the constraints of physical energy, protein 3D structures, and CATH domain objectives, we employ a hybrid noise strategy as data augmentation and present a geometric encoder scenario, named CATH-ddG, to represent the mutational microenvironment differences between wild-type and mutated protein complexes. Additionally, we fine-tune ESM2 representations by incorporating a lightweight nonlinear module to achieve the transferability of sequence co-evolutionary information. Finally, our study demonstrates that CATH-ddG framework provides enhanced generalization by outperforming other baselines on non-superfamily leakage splits, which plays a crucial role in exploring robust mutation effect regression prediction. Independent case studies demonstrate successful enhancement of binding affinity on 419 antibody variants to human epidermal growth factor receptor 2 (HER2) and 285 variants in the receptor-binding domain (RBD) of SARS-CoV-2 to angiotensin-converting enzyme 2 (ACE2) receptor. AVAILABILITY AND IMPLEMENTATION: CATH-ddG is available at https://github.com/ak422/CATH-ddG.
Guanglei Yu, Xuehua Bi, Yaohang Li, Jianxin Wang 0001
Bioinform.4
2025 LRTM: Left-Right Transition Matrices for Molecular Association Prediction
abstract
Molecular associations are central to most biological processes. The discovery and identification of potential associations between molecules can provide insights into biological exploration, diagnostic and therapeutic interventions, and drug development. So far many relevant computational methods have been proposed, but most of them are usually limited to specific domains and rely on complex preprocessing procedures, which restricts the models' ability to be applied to other tasks. Therefore, it remains a challenge to explore a generalized approach to accurately predicting potential associations. In this study, We propose Left-Right Transition Matrices (LRTM) for molecular association prediction. From the perspective on the diffusion model, we construct two transition matrices to model undirected graph information propagation. This allows modeling the transition probabilities of links, which facilitates link prediction in molecular bipartite networks. The extensive experimental results show that the proposed LRTM algorithm performs better than the compared methods. Also, the proposed algorithm has the potential for cross-task prediction. Furthermore, case studies show that LRTM is a powerful tool that can be effectively applied to practical applications.
Kai Zheng 0020, Guihua Duan, Mengyun Yang, Wei Wu 0011, Yaohang Li, Jianxin Wang 0001
IEEE Trans. Comput. Biol. Bioinform.5
2024 BenchLMM: Benchmarking Cross-Style Visual Capability of Large Multimodal Models
Rizhao Cai, Zirui Song, Dayan Guan, Zhenhao Chen, Yaohang Li, Chenyu Yi, Alex Chichung Kot
ECCV (50)5
2024 Unfolding Particle Detector Acceptance in High Energy Physics with Generative AI
abstract
The “acceptance problem” in high energy physics (HEP) refers to the challenge of accurately modeling detector acceptance to ensure the precision of measurements. This study explores the application of Generative AI to address the acceptance problem in HEP. By training a Generative Adversarial Network (GAN) on simulated detector data (pseudo-data), we demonstrate its capability to learn detector responses and generate synthetic data that closely match measured distributions. A key component of our methodology is a custom generator loss function that incorporates physics-informed principles to improve training. This custom loss function penalizes deviations from the true distribution of event components, ensuring that the generated samples adhere to the underlying physics. Additionally, we trained a binary classifier to distinguish between different topological states (measured and unmeasured events) within the generated Monte Carlo pseudodata, further refining the model's accuracy. Our approach preserves correlations between kinematic variables across multiple dimensions, providing an accurate representation of the underlying physics. Validation with Monte Carlo pseudodata demonstrates the method's ability to recover true distributions even in regions with limited detector sensitivity, establishing a solid foundation for applying our framework to real experimental data. Our results highlight the feasibility and advantages of using generative AI in HEP, paving the way for broader applications in the field.
Tareq Alghamdi, Tommaso Vittorini, Marco Spreafico, Marco Battaglieri, Nobuo Sato, Yaohang Li
ICTAI6
2024 TripHLApan: predicting HLA molecules binding peptides based on triple coding matrix and transfer learning
abstract
Human leukocyte antigen (HLA) recognizes foreign threats and triggers immune responses by presenting peptides to T cells. Computationally modeling the binding patterns between peptide and HLA is very important for the development of tumor vaccines. However, it is still a big challenge to accurately predict HLA molecules binding peptides. In this paper, we develop a new model TripHLApan for predicting HLA molecules binding peptides by integrating triple coding matrix, BiGRU + Attention models, and transfer learning strategy. We have found the main interaction site regions between HLA molecules and peptides, as well as the correlation between HLA encoding and binding motifs. Based on the discovery, we make the preprocessing and coding closer to the natural biological process. Besides, due to the input being based on multiple types of features and the attention module focused on the BiGRU hidden layer, TripHLApan has learned more sequence level binding information. The application of transfer learning strategies ensures the accuracy of prediction results under special lengths (peptides in length 8) and model scalability with the data explosion. Compared with the current optimal models, TripHLApan exhibits strong predictive performance in various prediction environments with different positive and negative sample ratios. In addition, we validate the superiority and scalability of TripHLApan's predictive performance using additional latest data sets, ablation experiments and binding reconstitution ability in the samples of a melanoma patient. The results show that TripHLApan is a powerful tool for predicting the binding of HLA-I and HLA-II molecular peptides for the synthesis of tumor vaccines. TripHLApan is publicly available at https://github.com/CSUBioGroup/TripHLApan.git.
Meng Wang 0067, Chuqi Lei, Jianxin Wang 0001, Yaohang Li, Min Li 0007
Briefings Bioinform.4
2024 Identifying new cancer genes based on the integration of annotated gene sets via hypergraph neural networks
abstract
MOTIVATION: Identifying cancer genes remains a significant challenge in cancer genomics research. Annotated gene sets encode functional associations among multiple genes, and cancer genes have been shown to cluster in hallmark signaling pathways and biological processes. The knowledge of annotated gene sets is critical for discovering cancer genes but remains to be fully exploited. RESULTS: Here, we present the DIsease-Specific Hypergraph neural network (DISHyper), a hypergraph-based computational method that integrates the knowledge from multiple types of annotated gene sets to predict cancer genes. First, our benchmark results demonstrate that DISHyper outperforms the existing state-of-the-art methods and highlight the advantages of employing hypergraphs for representing annotated gene sets. Second, we validate the accuracy of DISHyper-predicted cancer genes using functional validation results and multiple independent functional genomics data. Third, our model predicts 44 novel cancer genes, and subsequent analysis shows their significant associations with multiple types of cancers. Overall, our study provides a new perspective for discovering cancer genes and reveals previously undiscovered cancer genes. AVAILABILITY AND IMPLEMENTATION: DISHyper is freely available for download at https://github.com/genemine/DISHyper.
Hong-Dong Li, Li-Shen Zhang, Yaohang Li, Jianxin Wang 0001
Bioinform.5
2023 CRMSS: predicting circRNA-RBP binding sites based on multi-scale characterizing sequence and structure features
abstract
Circular RNAs (circRNAs) are reverse-spliced and covalently closed RNAs. Their interactions with RNA-binding proteins (RBPs) have multiple effects on the progress of many diseases. Some computational methods are proposed to identify RBP binding sites on circRNAs but suffer from insufficient accuracy, robustness and explanation. In this study, we first take the characteristics of both RNA and RBP into consideration. We propose a method for discriminating circRNA-RBP binding sites based on multi-scale characterizing sequence and structure features, called CRMSS. For circRNAs, we use sequence ${k}\hbox{-}{mer}$ embedding and the forming probabilities of local secondary structures as features. For RBPs, we combine sequence and structure frequencies of RNA-binding domain regions to generate features. We capture binding patterns with multi-scale residual blocks. With BiLSTM and attention mechanism, we obtain the contextual information of high-level representation for circRNA-RBP binding. To validate the effectiveness of CRMSS, we compare its predictive performance with other methods on 37 RBPs. Taking the properties of both circRNAs and RBPs into account, CRMSS achieves superior performance over state-of-the-art methods. In the case study, our model provides reliable predictions and correctly identifies experimentally verified circRNA-RBP pairs. The code of CRMSS is freely available at https://github.com/BioinformaticsCSU/CRMSS.
Lishen Zhang, Chengqian Lu, Min Zeng 0004, Yaohang Li, Jianxin Wang 0001
Briefings Bioinform.4
2023 DFHiC: a dilated full convolution model to enhance the resolution of Hi-C data
abstract
MOTIVATION: Hi-C technology has been the most widely used chromosome conformation capture (3C) experiment that measures the frequency of all paired interactions in the entire genome, which is a powerful tool for studying the 3D structure of the genome. The fineness of the constructed genome structure depends on the resolution of Hi-C data. However, due to the fact that high-resolution Hi-C data require deep sequencing and thus high experimental cost, most available Hi-C data are in low-resolution. Hence, it is essential to enhance the quality of Hi-C data by developing the effective computational methods. RESULTS: In this work, we propose a novel method, so-called DFHiC, which generates the high-resolution Hi-C matrix from the low-resolution Hi-C matrix in the framework of the dilated convolutional neural network. The dilated convolution is able to effectively explore the global patterns in the overall Hi-C matrix by taking advantage of the information of the Hi-C matrix in a way of the longer genomic distance. Consequently, DFHiC can improve the resolution of the Hi-C matrix reliably and accurately. More importantly, the super-resolution Hi-C data enhanced by DFHiC is more in line with the real high-resolution Hi-C data than those done by the other existing methods, in terms of both chromatin significant interactions and identifying topologically associating domains. AVAILABILITY AND IMPLEMENTATION: https://github.com/BinWangCSU/DFHiC.
Bin Wang 0045, Kun Liu 0028, Yaohang Li, Jianxin Wang 0001
Bioinform.3
2023 CellBRF: a feature selection method for single-cell clustering using cell balance and random forest
abstract
MOTIVATION: Single-cell RNA sequencing (scRNA-seq) offers a powerful tool to dissect the complexity of biological tissues through cell sub-population identification in combination with clustering approaches. Feature selection is a critical step for improving the accuracy and interpretability of single-cell clustering. Existing feature selection methods underutilize the discriminatory potential of genes across distinct cell types. We hypothesize that incorporating such information could further boost the performance of single cell clustering. RESULTS: We develop CellBRF, a feature selection method that considers genes' relevance to cell types for single-cell clustering. The key idea is to identify genes that are most important for discriminating cell types through random forests guided by predicted cell labels. Moreover, it proposes a class balancing strategy to mitigate the impact of unbalanced cell type distributions on feature importance evaluation. We benchmark CellBRF on 33 scRNA-seq datasets representing diverse biological scenarios and demonstrate that it substantially outperforms state-of-the-art feature selection methods in terms of clustering accuracy and cell neighborhood consistency. Furthermore, we demonstrate the outstanding performance of our selected features through three case studies on cell differentiation stage identification, non-malignant cell subtype identification, and rare cell identification. CellBRF provides a new and effective tool to boost single-cell clustering accuracy. AVAILABILITY AND IMPLEMENTATION: All source codes of CellBRF are freely available at https://github.com/xuyp-csu/CellBRF.
Yunpei Xu, Hong-Dong Li, Cui-Xiang Lin, Ruiqing Zheng, Yaohang Li, Jinhui Xu 0001, Jianxin Wang 0001
Bioinform.5
2023 MSDRP: a deep learning model based on multisource data for predicting drug response
abstract
MOTIVATION: Cancer heterogeneity drastically affects cancer therapeutic outcomes. Predicting drug response in vitro is expected to help formulate personalized therapy regimens. In recent years, several computational models based on machine learning and deep learning have been proposed to predict drug response in vitro. However, most of these methods capture drug features based on a single drug description (e.g. drug structure), without considering the relationships between drugs and biological entities (e.g. target, diseases, and side effects). Moreover, most of these methods collect features separately for drugs and cell lines but fail to consider the pairwise interactions between drugs and cell lines. RESULTS: In this paper, we propose a deep learning framework, named MSDRP for drug response prediction. MSDRP uses an interaction module to capture interactions between drugs and cell lines, and integrates multiple associations/interactions between drugs and biological entities through similarity network fusion algorithms, outperforming some state-of-the-art models in all performance measures for all experiments. The experimental results of de novo test and independent test demonstrate the excellent performance of our model for new drugs. Furthermore, several case studies illustrate the rationality for using feature vectors derived from drug similarity matrices from multisource data to represent drugs and the interpretability of our model. AVAILABILITY AND IMPLEMENTATION: The codes of MSDRP are available at https://github.com/xyzhang-10/MSDRP.
Qichang Zhao, Yaohang Li, Jianxin Wang 0001
Bioinform.4
2023 Retrieve and rerank for automated ICD coding via Contrastive Learning
Kunying Niu, Yifan Wu 0008, Yaohang Li, Min Li 0007
J. Biomed. Informatics3
2023 A Deep Learning Framework for Predicting Protein Functions With Co-Occurrence of GO Terms
abstract
The understanding of protein functions is critical to many biological problems such as the development of new drugs and new crops. To reduce the huge gap between the increase of protein sequences and annotations of protein functions, many methods have been proposed to deal with this problem. These methods use Gene Ontology (GO) to classify the functions of proteins and consider one GO term as a class label. However, they ignore the co-occurrence of GO terms that is helpful for protein function prediction. We propose a new deep learning model, named DeepPFP-CO, which uses Graph Convolutional Network (GCN) to explore and capture the co-occurrence of GO terms to improve the protein function prediction performance. In this way, we can further deduce the protein functions by fusing the predicted propensity of the center function and its co-occurrence functions. We use Fmax and AUPR to evaluate the performance of DeepPFP-CO and compare DeepPFP-CO with state-of-the-art methods such as DeepGOPlus and DeepGOA. The computational results show that DeepPFP-CO outperforms DeepGOPlus and other methods. Moreover, we further analyze our model at the protein level. The results have demonstrated that DeepPFP-CO improves the performance of protein function prediction. DeepPFP-CO is available at https://csuligroup.com/DeepPFP/.
Min Li 0007, Fuhao Zhang, Min Zeng 0004, Yaohang Li
IEEE ACM Trans. Comput. Biol. Bioinform.5
2023 A Comparison of Topologically Associating Domain Callers Based on Hi-C Data
abstract
Topologically associating domains (TADs) are local chromatin interaction domains, which have been shown to play an important role in gene expression regulation. TADs were originally discovered in the investigation of 3D genome organization based on High-throughput Chromosome Conformation Capture (Hi-C) data. Continuous considerable efforts have been dedicated to developing methods for detecting TADs from Hi-C data. Different computational methods for TADs identification vary in their assumptions and criteria in calling TADs. As a consequence, the TADs called by these methods differ in their similarities and biological features they are enriched in. In this work, we performed a systematic comparison of twenty-six TAD callers. We first compared the TADs and gaps between adjacent TADs across different methods, resolutions, and sequencing depths. We then assessed the quality of TADs and TAD boundaries according to three criteria: the decay of contact frequencies over the genomic distance, enrichment and depletion of regulatory elements around TAD boundaries, and reproducibility of TADs and TAD boundaries in replicate samples. Last, due to the lack of a gold standard of TADs, we also evaluated the performance of the methods on synthetic datasets. We discussed the key principles of TAD callers, and pinpointed current situation in the detection of TADs. We provide a concise, comprehensive, and systematic framework for evaluating the performance of TAD callers, and expect our work will provide useful guidance in choosing suitable approaches for the detection and evaluation of TADs.
Kun Liu 0028, Hong-Dong Li, Yaohang Li, Jun Wang 0153, Jianxin Wang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.3
2023 RNPredATC: A Deep Residual Learning-Based Model With Applications to the Prediction of Drug-ATC Code Association
abstract
The Anatomical Therapeutic Chemical (ATC) classification system, designated by the World Health Organization Collaborating Center (WHOCC), has been widely used in drug screening, repositioning, and similarity research. The ATC classification system assigns different codes to drugs according to the organ or system on which they act and/or their therapeutic and chemical characteristics. Correctly identifying the potential ATC codes for drugs can accelerate drug development and reduce the cost of experiments. Several classifiers have been proposed in this regard. However, they lack of ability to learn basic features from sparsely known drug-ATC code associations. Therefore, there is an urgent need for novel computational methods to precisely predict potential drug-ATC code associations in multiple levels of the ATC classification system based on known associations between drugs and ATC codes. In this paper, we provide a novel end-to-end model, so-called RNPredATC, to predict potential drug-ATC code associations in five ATC classification levels. RNPredATC can extract dense feature vectors from sparsely known drug-ATC code associations and reduce the impact from the degradation problem by a novel deep residual learning. We extensively compare our method with some state-of-the-art methods, including NetPredATC, SPACE, and some multi-label-based methods. Our experimental results show that RNPredATC achieves better performances in five-fold and ten-fold cross validations. Furthermore, the visualization analysis of hidden layers and case studies of predicted associations at the fifth ATC classification level confirm that RNPredATC can effectively identify the potential ATC codes of drugs.
Guihua Duan, Yaohang Li, Jianxin Wang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.5
2023 AttentionDTA: Drug-Target Binding Affinity Prediction by Sequence-Based Deep Learning With Attention Mechanism
abstract
The identification of drug-target relations (DTRs) is substantial in drug development. A large number of methods treat DTRs as drug-target interactions (DTIs), a binary classification problem. The main drawback of these methods are the lack of reliable negative samples and the absence of many important aspects of DTR, including their dose dependence and quantitative affinities. With increasing number of publications of drug-protein binding affinity data recently, DTRs prediction can be viewed as a regression problem of drug-target affinities (DTAs) which reflects how tightly the drug binds to the target and can present more detailed and specific information than DTIs. The growth of affinity data enables the use of deep learning architectures, which have been shown to be among the state-of-the-art methods in binding affinity prediction. Although relatively effective, due to the black-box nature of deep learning, these models are less biologically interpretable. In this study, we proposed a deep learning-based model, named AttentionDTA, which uses attention mechanism to predict DTAs. Different from the models using 3D structures of drug-target complexes or graph representation of drugs and proteins, the novelty of our work is to use attention mechanism to focus on key subsequences which are important in drug and protein sequences when predicting its affinity. We use two separate one-dimensional Convolution Neural Networks (1D-CNNs) to extract the semantic information of drug's SMILES string and protein's amino acid sequence. Furthermore, a two-side multi-head attention mechanism is developed and embedded to our model to explore the relationship between drug features and protein features. We evaluate our model on three established DTA benchmark datasets, Davis, Metz, and KIBA. AttentionDTA outperforms the state-of-the-art deep learning methods under different evaluation metrics. The results show that the attention-based model can effectively extract protein features related to drug information and drug features related to protein information to better predict drug target affinities. It is worth mentioning that we test our model on IC50 dataset, which provides the binding sites between drugs and proteins, to evaluate the ability of our model to locate binding sites. Finally, we visualize the attention weight to demonstrate the biological significance of the model. The source code of AttentionDTA can be downloaded from https://github.com/zhaoqichang/AttentionDTA_TCBB.
Qichang Zhao, Guihua Duan, Mengyun Yang, Zhongjian Cheng, Yaohang Li, Jianxin Wang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.5
2023 GIFDTI: Prediction of Drug-Target Interactions Based on Global Molecular and Intermolecular Interaction Representation Learning
abstract
Drug discovery and drug repurposing often rely on the successful prediction of drug-target interactions (DTIs). Recent advances have shown great promise in applying deep learning to drug-target interaction prediction. One challenge in building deep learning-based models is to adequately represent drugs and proteins that encompass the fundamental local chemical environments and long-distance information among amino acids of proteins (or atoms of drugs). Another challenge is to efficiently model the intermolecular interactions between drugs and proteins, which plays vital roles in the DTIs. To this end, we propose a novel model, GIFDTI, which consists of three key components: the sequence feature extractor (CNNFormer), the global molecular feature extractor (GF), and the intermolecular interaction modeling module (IIF). Specifically, CNNFormer incorporates CNN and Transformer to capture the local patterns and encode the long-distance relationship among tokens (atoms or amino acids) in a sequence. Then, GF and IIF extract the global molecular features and the intermolecular interaction features, respectively. We evaluate GIFDTI on six realistic evaluation strategies and the results show it improves DTI prediction performance compared to state-of-the-art methods. Moreover, case studies confirm that our model can be a useful tool to accurately yield low-cost DTIs. The codes of GIFDTI are available at https://github.com/zhaoqichang/GIFDTI.
Qichang Zhao, Guihua Duan, Kai Zheng 0020, Yaohang Li, Jianxin Wang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.5
2022 Point Cloud-based Variational Autoencoder Inverse Mappers (PC-VAIM) - An Application on Quantum Chromodynamics Global Analysis
abstract
Correctly mapping the experimental data to quantum probability distributions is a critical step to characterize nucleon structure and the emergence of hadrons, in terms of quark and gluon degrees of freedom. Since the actual parameters of interest are not directly measurable, but instead are inferred from experimental observables, this is fundamentally an inverse problem of recovering parameters from observables. In addition to the well-known challenges such as ill-posedness in general inverse problems, an application specific issue here is that the experimental data are observed on kinematics bins which are usually irregular and varying. In this paper, to address this ill- defined, varying observable space problem, we represent the observables together with their kinematics bins as an unstructured, high-dimensional point cloud. We incorporate a permutation invariant neural network framework to handle the observables in unstructured and unordered point cloud representations. We incorporate the point cloud representation into the Variational Autoencoder Inverse Mapper (VAIM) framework. The point cloud-based VAIM (PC-VAIM) enables the underlying deep neural networks to learn how the observables are distributed across kinematics. We demonstrate the effectiveness of PC-VAIM on a toy inverse problem, and then on constructing the inverse function mapping Quantum Correlation Functions (QCF) to observables in a Quantum Chromodynamics (QCD) analysis of nucleon structure.
Manal Almaeen, Yasir Alanazi, Nobuo Sato, Wally Melnitchouk, Yaohang Li
ICMLA5
2022 Image Synthesis Using Conditional GANs for Selective Laser Melting Additive Manufacturing
abstract
In-situ process monitoring for metals additive manufacturing is paramount to the successful build of an object for application in extreme or high stress environments. Yet in selective laser melting additive manufacturing, it is extremely difficult to evaluate the build process. The difficulty is that obtaining enough variety of data to quantify the internal microstructures for the evaluation of its physical properties is problematic, as the laser passes at high speeds over powder grains at a micrometer scale. Using generative models, a type of machine learning, has been shown here to provide new artificially generated data with the same properties as the experimental images. The Generative Adversarial Network (GAN) synthesized new computationally derived data through a process that learns the underlying features of images that correspond to the different laser process parameters in a generator network. While this technique was effective at delivering high-quality images that closely matched the training data when tested against holdout samples, modifications to the general form of the network through a conditional generative adversarial network (CGAN) showed improved capabilities at creating these new images. Using multiple evaluation metrics, it has been shown that generative models can be used to create new data for various laser process parameter combinations, thereby allowing a more comprehensive evaluation of ideal laser conditions for any particular build. The new data can supplement the experimental data, thereby growing the overall knowledge framework for build characteristics.
Andy Ramlatchan, Yaohang Li
IJCNN2
2022 IIFDTI: predicting drug-target interactions through interactive and independent features based on attention mechanism
abstract
MOTIVATION: Identifying drug-target interactions is a crucial step for drug discovery and design. Traditional biochemical experiments are credible to accurately validate drug-target interactions. However, they are also extremely laborious, time-consuming and expensive. With the collection of more validated biomedical data and the advancement of computing technology, the computational methods based on chemogenomics gradually attract more attention, which guide the experimental verifications. RESULTS: In this study, we propose an end-to-end deep learning-based method named IIFDTI to predict drug-target interactions (DTIs) based on independent features of drug-target pairs and interactive features of their substructures. First, the interactive features of substructures between drugs and targets are extracted by the bidirectional encoder-decoder architecture. The independent features of drugs and targets are extracted by the graph neural networks and convolutional neural networks, respectively. Then, all extracted features are fused and inputted into fully connected dense layers in downstream tasks for predicting DTIs. IIFDTI takes into account the independent features of drugs/targets and simulates the interactive features of the substructures from the biological perspective. Multiple experiments show that IIFDTI outperforms the state-of-the-art methods in terms of the area under the receiver operating characteristics curve (AUC), the area under the precision-recall curve (AUPR), precision, and recall on benchmark datasets. In addition, the mapped visualizations of attention weights indicate that IIFDTI has learned the biological knowledge insights, and two case studies illustrate the capabilities of IIFDTI in practical applications. AVAILABILITY AND IMPLEMENTATION: The data and codes underlying this article are available in Github at https://github.com/czjczj/IIFDTI. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhongjian Cheng, Qichang Zhao, Yaohang Li, Jianxin Wang 0001
Bioinform.3
2022 BACPI: a bi-directional attention neural network for compound-protein interaction and binding affinity prediction
abstract
MOTIVATION: The identification of compound-protein interactions (CPIs) is an essential step in the process of drug discovery. The experimental determination of CPIs is known for a large amount of funds and time it consumes. Computational model has therefore become a promising and efficient alternative for predicting novel interactions between compounds and proteins on a large scale. Most supervised machine learning prediction models are approached as a binary classification problem, which aim to predict whether there is an interaction between the compound and the protein or not. However, CPI is not a simple binary on-off relationship, but a continuous value reflects how tightly the compound binds to a particular target protein, also called binding affinity. RESULTS: In this study, we propose an end-to-end neural network model, called BACPI, to predict CPI and binding affinity. We employ graph attention network and convolutional neural network (CNN) to learn the representations of compounds and proteins and develop a bi-directional attention neural network model to integrate the representations. To evaluate the performance of BACPI, we use three CPI datasets and four binding affinity datasets in our experiments. The results show that, when predicting CPIs, BACPI significantly outperforms other available machine learning methods on both balanced and unbalanced datasets. This suggests that the end-to-end neural network model that predicts CPIs directly from low-level representations is more robust than traditional machine learning-based methods. And when predicting binding affinities, BACPI achieves higher performance on large datasets compared to other state-of-the-art deep learning methods. This comparison result suggests that the proposed method with bi-directional attention neural network can capture the important regions of compounds and proteins for binding affinity prediction. AVAILABILITY AND IMPLEMENTATION: Data and source codes are available at https://github.com/CSUBioGroup/BACPI.
Min Li 0007, Zhangli Lu, Yifan Wu 0008, Yaohang Li
Bioinform.4
2022 Accurate Prediction of Human Essential Proteins Using Ensemble Deep Learning
abstract
Essential proteins are considered the foundation of life as they are indispensable for the survival of living organisms. Computational methods for essential protein discovery provide a fast way to identify essential proteins. But most of them heavily rely on various biological information, especially protein-protein interaction networks, which limits their practical applications. With the rapid development of high-throughput sequencing technology, sequencing data has become the most accessible biological data. However, using only protein sequence information to predict essential proteins has limited accuracy. In this paper, we propose EP-EDL, an ensemble deep learning model using only protein sequence information to predict human essential proteins. EP-EDL integrates multiple classifiers to alleviate the class imbalance problem and to improve prediction accuracy and robustness. In each base classifier, we employ multi-scale text convolutional neural networks to extract useful features from protein sequence feature matrices with evolutionary information. Our computational results show that EP-EDL outperforms the state-of-the-art sequence-based methods. Furthermore, EP-EDL provides a more practical and flexible way for biologists to accurately predict essential proteins. The source code and datasets can be downloaded from https://github.com/CSUBioGroup/EP-EDL.
Min Zeng 0004, Yifan Wu 0008, Yaohang Li, Min Li 0007
IEEE ACM Trans. Comput. Biol. Bioinform.4
2022 Biomedical Data and Deep Learning Computational Models for Predicting Compound-Protein Relations
abstract
The identification of compound-protein relations (CPRs), which includes compound-protein interactions (CPIs) and compound-protein affinities (CPAs), is critical to drug development. A common method for compound-protein relation identification is the use of in vitro screening experiments. However, the number of compounds and proteins is massive, and in vitro screening experiments are labor-intensive, expensive, and time-consuming with high failure rates. Researchers have developed a computational field called virtual screening (VS) to aid experimental drug development. These methods utilize experimentally validated biological interaction information to generate datasets and use the physicochemical and structural properties of compounds and target proteins as input information to train computational prediction models. At present, deep learning has been widely used in computer vision and natural language processing and has experienced epoch-making progress. At the same time, deep learning has also been used in the field of biomedicine widely, and the prediction of CPRs based on deep learning has developed rapidly and has achieved good results. The purpose of this study is to investigate and discuss the latest applications of deep learning techniques in CPR prediction. First, we describe the datasets and feature engineering (i.e., compound and protein representations and descriptors) commonly used in CPR prediction methods. Then, we review and classify recent deep learning approaches in CPR prediction. Next, a comprehensive comparison is performed to demonstrate the prediction performance of representative methods on classical datasets. Finally, we discuss the current state of the field, including the existing challenges and our proposed future directions. We believe that this investigation will provide sufficient references and insight for researchers to understand and develop new deep learning methods to enhance CPR predictions.
Qichang Zhao, Mengyun Yang, Zhongjian Cheng, Yaohang Li, Jianxin Wang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.4
2022 A Pseudo Label-Wise Attention Network for Automatic ICD Coding
abstract
Automatic International Classification of Diseases (ICD) coding is defined as a kind of text multi-label classification problem, which is difficult because the number of labels is very large and the distribution of labels is unbalanced. The label-wise attention mechanism is widely used in automatic ICD coding because it can assign weights to every word in full Electronic Medical Records (EMR) for different ICD codes. However, the label-wise attention mechanism is redundant and costly in computing. In this paper, we propose a pseudo label-wise attention mechanism to tackle the problem. Instead of computing different attention modes for different ICD codes, the pseudo label-wise attention mechanism automatically merges similar ICD codes and computes only one attention mode for the similar ICD codes, which greatly compresses the number of attention modes and improves the predicted accuracy. In addition, we apply a more convenient and effective way to obtain the ICD vectors, and thus our model can predict new ICD codes by calculating the similarities between EMR vectors and ICD vectors. Our model demonstrates effectiveness in extensive computational experiments. On the public MIMIC-III dataset and private Xiangya dataset, our model achieves the best performance on micro F1 (0.583 and 0.806), micro AUC (0.986 and 0.994), P@8 (0.756 and 0.413), and costs much smaller GPU memory (about 26.1% of the models with label-wise attention). Furthermore, we verify the ability of our model in predicting new ICD codes. The interpretablility analysis and case study show the effectiveness and reliability of the patterns obtained by the pseudo label-wise attention mechanism.
Yifan Wu 0008, Min Zeng 0004, Yaohang Li, Min Li 0007
IEEE J. Biomed. Health Informatics4
2022 Feature and Nuclear Norm Minimization for Matrix Completion
abstract
Matrix completion, whose goal is to recover a matrix from a few entries observed, is a fundamental model behind many applications. Our study shows that, in many applications, the to-be-complete matrix can be represented as the sum of a low-rank matrix and a sparse matrix associating with side information matrices. The low-rank matrix depicts the global patterns while the sparse matrix characterizes the local patterns, which are often described by the side information. Accordingly, to achieve high-quality matrix completion, we propose a Feature and Nuclear Norm Minimization (FNNM) model. The rationale of FNNM is to employ transductive completion to generalize the global pattern and inductive completion to recover the local pattern. Alternative minimization algorithm based on fixed-point iteration is developed to numerically solve the FNNM model. FNNM has demonstrated promising results on a variety of applications, including movie recommendation, drug-target interaction prediction, and multi-label learning, consistently outperforming the state-of-the-art matrix completion algorithms.
Mengyun Yang, Yaohang Li, Jianxin Wang 0001
IEEE Trans. Knowl. Data Eng.2
2021 Simulation of Electron-Proton Scattering Events by a Feature-Augmented and Transformed Generative Adversarial Network (FAT-GAN)
abstract
We apply generative adversarial network (GAN) technology to build an event generator that simulates particle production in electron-proton scattering that is free of theoretical assumptions about underlying particle dynamics. The difficulty of efficiently training a GAN event simulator lies in learning the complicated patterns of the distributions of the particles physical properties. We develop a GAN that selects a set of transformed features from particle momenta that can be generated easily by the generator, and uses these to produce a set of augmented features that improve the sensitivity of the discriminator. The new Feature-Augmented and Transformed GAN (FAT-GAN) is able to faithfully reproduce the distribution of final state electron momenta in inclusive electron scattering, without the need for input derived from domain-based theoretical assumptions. The developed technology can play a significant role in boosting the science of existing and future accelerator facilities, such as the Electron-Ion Collider.
Yasir Alanazi, Nobuo Sato, Tianbo Liu 0004, Wally Melnitchouk, Pawel Ambrozewicz, Florian Hauenstein, Michelle P. Kuchera, Evan Pritchard, Michael Robertson, Ryan R. Strauss, Luisa Velasco, Yaohang Li
IJCAI12
2021 A Survey of Machine Learning-Based Physics Event Generation
abstract
Event generators in high-energy nuclear and particle physics play an important role in facilitating studies of particle reactions. We survey the state of the art of machine learning (ML) efforts at building physics event generators. We review ML generative models used in ML-based event generators and their specific challenges, and discuss various approaches of incorporating physics into the ML model designs to overcome these challenges. Finally, we explore some open questions related to super-resolution, fidelity, and extrapolation for physics event generation based on ML technology.
Yasir Alanazi, Nobuo Sato, Pawel Ambrozewicz, Astrid N. Hiller Blin, Wally Melnitchouk, Marco Battaglieri, Tianbo Liu 0004, Yaohang Li
IJCAI8
2021 Variational Autoencoder Inverse Mapper: An End-to-End Deep Learning Framework for Inverse Problems
abstract
Inverse problems - using measured observations to determine unknown parameters - are well motivated but challenging in many science and engineering problems. In this paper, we propose an end-to-end deep learning framework, the Variational Autoencoder Inverse Mapper (VAIM), as an autoencoder-based neural network architecture for inverse problems. The encoder and decoder neural networks approximate the forward and backward mapping, respectively, and a variational latent layer is incorporated into VAIM to learn the posterior parameter distributions with respect to given observables. We demonstrate the effectiveness of VAIM for several toy inverse problems, with both finite and infinite solutions, and for constructing the inverse function mapping quantum correlation functions to observables in a Quantum Chromodynamics analysis of nucleon structure.
Manal Almaeen, Yasir Alanazi, Nobuo Sato, Wally Melnitchouk, Michelle P. Kuchera, Yaohang Li
IJCNN6
2021 Biomedical data and computational models for drug repositioning: a comprehensive review
abstract
Drug repositioning can drastically decrease the cost and duration taken by traditional drug research and development while avoiding the occurrence of unforeseen adverse events. With the rapid advancement of high-throughput technologies and the explosion of various biological data and medical data, computational drug repositioning methods have been appealing and powerful techniques to systematically identify potential drug-target interactions and drug-disease interactions. In this review, we first summarize the available biomedical data and public databases related to drugs, diseases and targets. Then, we discuss existing drug repositioning approaches and group them based on their underlying computational models consisting of classical machine learning, network propagation, matrix factorization and completion, and deep learning based models. We also comprehensively analyze common standard data sets and evaluation metrics used in drug repositioning, and give a brief comparison of various prediction methods on the gold standard data sets. Finally, we conclude our review with a brief discussion on challenges in computational drug repositioning, which includes the problem of reducing the noise and incompleteness of biomedical data, the ensemble of various computation drug repositioning methods, the importance of designing reliable negative samples selection methods, new techniques dealing with the data sparseness problem, the construction of large-scale and comprehensive benchmark data sets and the analysis and explanation of the underlying mechanisms of predicted interactions.
Huimin Luo, Min Li 0007, Mengyun Yang, Fang-Xiang Wu, Yaohang Li, Jianxin Wang 0001
Briefings Bioinform.5
2021 DeepDTAF: a deep learning method to predict protein-ligand binding affinity
abstract
Biomolecular recognition between ligand and protein plays an essential role in drug discovery and development. However, it is extremely time and resource consuming to determine the protein-ligand binding affinity by experiments. At present, many computational methods have been proposed to predict binding affinity, most of which usually require protein 3D structures that are not often available. Therefore, new methods that can fully take advantage of sequence-level features are greatly needed to predict protein-ligand binding affinity and accelerate the drug discovery process. We developed a novel deep learning approach, named DeepDTAF, to predict the protein-ligand binding affinity. DeepDTAF was constructed by integrating local and global contextual features. More specifically, the protein-binding pocket, which possesses some special properties for directly binding the ligand, was firstly used as the local input feature for protein-ligand binding affinity prediction. Furthermore, dilated convolution was used to capture multiscale long-range interactions. We compared DeepDTAF with the recent state-of-art methods and analyzed the effectiveness of different parts of our model, the significant accuracy improvement showed that DeepDTAF was a reliable tool for affinity prediction. The resource codes and data are available at https: //github.com/KailiWang1/DeepDTAF.
Renyi Zhou, Yaohang Li, Min Li 0007
Briefings Bioinform.3
2021 Computational drug repositioning based on multi-similarities bilinear matrix factorization
abstract
With the development of high-throughput technology and the accumulation of biomedical data, the prior information of biological entity can be calculated from different aspects. Specifically, drug-drug similarities can be measured from target profiles, drug-drug interaction and side effects. Similarly, different methods and data sources to calculate disease ontology can result in multiple measures of pairwise disease similarities. Therefore, in computational drug repositioning, developing a dynamic method to optimize the fusion process of multiple similarities is a crucial and challenging task. In this study, we propose a multi-similarities bilinear matrix factorization (MSBMF) method to predict promising drug-associated indications for existing and novel drugs. Instead of fusing multiple similarities into a single similarity matrix, we concatenate these similarity matrices of drug and disease, respectively. Applying matrix factorization methods, we decompose the drug-disease association matrix into a drug-feature matrix and a disease-feature matrix. At the same time, using these feature matrices as basis, we extract effective latent features representing the drug and disease similarity matrices to infer missing drug-disease associations. Moreover, these two factored matrices are constrained by non-negative factorization to ensure that the completed drug-disease association matrix is biologically interpretable. In addition, we numerically solve the MSBMF model by an efficient alternating direction method of multipliers algorithm. The computational experiment results show that MSBMF obtains higher prediction accuracy than the state-of-the-art drug repositioning methods in cross-validation experiments. Case studies also demonstrate the effectiveness of our proposed method in practical applications. Availability: The data and code of MSBMF are freely available at https://github.com/BioinformaticsCSU/MSBMF. Corresponding author: Jianxin Wang, School of Computer Science and Engineering, Central South University, Changsha, Hunan 410083, P. R. China. E-mail: [email protected] Supplementary Data: Supplementary data are available online at https://academic.oup.com/bib.
Mengyun Yang, Gaoyan Wu, Qichang Zhao, Yaohang Li, Jianxin Wang 0001
Briefings Bioinform.4
2021 A novel graph attention model for predicting frequencies of drug-side effects from multi-view data
abstract
Identifying the frequencies of the drug-side effects is a very important issue in pharmacological studies and drug risk-benefit. However, designing clinical trials to determine the frequencies is usually time consuming and expensive, and most existing methods can only predict the drug-side effect existence or associations, not their frequencies. Inspired by the recent progress of graph neural networks in the recommended system, we develop a novel prediction model for drug-side effect frequencies, using a graph attention network to integrate three different types of features, including the similarity information, known drug-side effect frequency information and word embeddings. In comparison, the few available studies focusing on frequency prediction use only the known drug-side effect frequency scores. One novel approach used in this work first decomposes the feature types in drug-side effect graph to extract different view representation vectors based on three different type features, and then recombines these latent view vectors automatically to obtain unified embeddings for prediction. The proposed method demonstrates high effectiveness in 10-fold cross-validation. The computational results show that the proposed method achieves the best performance in the benchmark dataset, outperforming the state-of-the-art matrix decomposition model. In addition, some ablation experiments and visual analyses are also supplied to illustrate the usefulness of our method for the prediction of the drug-side effect frequencies. The codes of MGPred are available at https://github.com/zhc940702/MGPred and https://zenodo.org/record/4449613.
Kai Zheng 0020, Yaohang Li, Jianxin Wang 0001
Briefings Bioinform.3
2021 Parallel computing for genome sequence processing
abstract
The rapid increase of genome data brought by gene sequencing technologies poses a massive challenge to data processing. To solve the problems caused by enormous data and complex computing requirements, researchers have proposed many methods and tools which can be divided into three types: big data storage, efficient algorithm design and parallel computing. The purpose of this review is to investigate popular parallel programming technologies for genome sequence processing. Three common parallel computing models are introduced according to their hardware architectures, and each of which is classified into two or three types and is further analyzed with their features. Then, the parallel computing for genome sequence processing is discussed with four common applications: genome sequence alignment, single nucleotide polymorphism calling, genome sequence preprocessing, and pattern detection and searching. For each kind of application, its background is firstly introduced, and then a list of tools or algorithms are summarized in the aspects of principle, hardware platform and computing efficiency. The programming model of each hardware and application provides a reference for researchers to choose high-performance computing tools. Finally, we discuss the limitations and future trends of parallel computing technologies.
You Zou, Yuejie Zhu, Yaohang Li, Fang-Xiang Wu, Jianxin Wang 0001
Briefings Bioinform.3
2021 A convolutional neural network and graph convolutional network-based method for predicting the classification of anatomical therapeutic chemicals
abstract
MOTIVATION: The Anatomical Therapeutic Chemical (ATC) system is an official classification system established by the World Health Organization for medicines. Correctly assigning ATC classes to given compounds is an important research problem in drug discovery, which can not only discover the possible active ingredients of the compounds, but also infer theirs therapeutic, pharmacological and chemical properties. RESULTS: In this article, we develop an end-to-end multi-label classifier called CGATCPred to predict 14 main ATC classes for given compounds. In order to extract rich features of each compound, we use the deep Convolutional Neural Network and shortcut connections to represent and learn the seven association scores between the given compound and others. Moreover, we construct the correlation graph of ATC classes and then apply graph convolutional network on the graph for label embedding abstraction. We use all label embedding to guide the learning process of compound representation. As a result, by using the Jackknife test, CGATCPred obtain reliable Aiming of 81.94%, Coverage of 82.88%, Accuracy 80.81%, Absolute True 76.58% and Absolute False 2.75%, yielding significantly improvements compared to exiting multi-label classifiers. AVAILABILITY AND IMPLEMENTATION: The codes of CGATCPred are available at https://github.com/zhc940702/CGATCPred and https://zenodo.org/record/4552917.
Yaohang Li, Jianxin Wang 0001
Bioinform.2
2021 Protein interaction networks: centrality, modularity, dynamics, and applications
Xiangmao Meng, Xiaoqing Peng, Yaohang Li, Min Li 0007
Frontiers Comput. Sci.4
2021 DeepDSC: A Deep Learning Method to Predict Drug Sensitivity of Cancer Cell Lines
abstract
High-throughput screening technologies have provided a large amount of drug sensitivity data for a panel of cancer cell lines and hundreds of compounds. Computational approaches to analyzing these data can benefit anticancer therapeutics by identifying molecular genomic determinants of drug sensitivity and developing new anticancer drugs. In this study, we have developed a deep learning architecture to improve the performance of drug sensitivity prediction based on these data. We integrated both genomic features of cell lines and chemical information of compounds to predict the half maximal inhibitory concentrations [Formula: see text] on the Cancer Cell Line Encyclopedia (CCLE) and the Genomics of Drug Sensitivity in Cancer (GDSC) datasets using a deep neural network, which we called DeepDSC. Specifically, we first applied a stacked deep autoencoder to extract genomic features of cell lines from gene expression data, and then combined the compounds' chemical features to these genomic features to produce final response data. We conducted 10-fold cross-validation to demonstrate the performance of our deep model in terms of root-mean-square error (RMSE) and coefficient of determination [Formula: see text]. We show that our model outperforms the previous approaches with RMSE of 0.23 and [Formula: see text] of 0.78 on CCLE dataset, and RMSE of 0.52 and [Formula: see text] of 0.78 on GDSC dataset, respectively. Moreover, to demonstrate the prediction ability of our models on novel cell lines or novel compounds, we left cell lines originating from the same tissue and each compound out as the test sets, respectively, and the rest as training sets. The performance was comparable to other methods.
Min Li 0007, Yake Wang, Ruiqing Zheng, Xinghua Shi, Yaohang Li, Fang-Xiang Wu, Jianxin Wang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.5
2021 A Deep Learning Framework for Identifying Essential Proteins by Integrating Multiple Types of Biological Information
abstract
Computational methods including centrality and machine learning-based methods have been proposed to identify essential proteins for understanding the minimum requirements of the survival and evolution of a cell. In centrality methods, researchers are required to design a score function which is based on prior knowledge, yet is usually not sufficient to capture the complexity of biological information. In machine learning-based methods, some selected biological features cannot represent the complete properties of biological information as they lack a computational framework to automatically select features. To tackle these problems, we propose a deep learning framework to automatically learn biological features without prior knowledge. We use node2vec technique to automatically learn a richer representation of protein-protein interaction (PPI) network topologies than a score function. Bidirectional long short term memory cells are applied to capture non-local relationships in gene expression data. For subcellular localization information, we exploit a high dimensional indicator vector to characterize their feature. To evaluate the performance of our method, we tested it on PPI network of S. cerevisiae. Our experimental results demonstrate that the performance of our method is better than traditional centrality methods and is superior to existing machine learning-based methods. To explore which of the three types of biological information is the most vital element, we conduct an ablation study by removing each component in turn. Our results show that the PPI network embedding contributes most to the improvement. In addition, gene expression profiles and subcellular localization information are also helpful to improve the performance in identification of essential proteins.
Min Zeng 0004, Min Li 0007, Zhihui Fei, Fang-Xiang Wu, Yaohang Li, Yi Pan 0001, Jianxin Wang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.5
2021 DMFLDA: A Deep Learning Framework for Predicting lncRNA-Disease Associations
abstract
A growing amount of evidence suggests that long non-coding RNAs (lncRNAs) play important roles in the regulation of biological processes in many human diseases. However, the number of experimentally verified lncRNA-disease associations is very limited. Thus, various computational approaches are proposed to predict lncRNA-disease associations. Current matrix factorization-based methods cannot capture the complex non-linear relationship between lncRNAs and diseases, and traditional machine learning-based methods are not sufficiently powerful to learn the representation of lncRNAs and diseases. Considering these limitations in existing computational methods, we propose a deep matrix factorization model to predict lncRNA-disease associations (DMFLDA in short). DMFLDA uses a cascade of non-linear hidden layers to learn latent representation to represent lncRNAs and diseases. By using non-linear hidden layers, DMFLDA captures the more complex non-linear relationship between lncRNAs and diseases than traditional matrix factorization-based methods. In addition, DMFLDA learns features directly from the lncRNA-disease interaction matrix and thus can obtain more accurate representation learning for lncRNAs and diseases than traditional machine learning methods. The low dimensional representations of the lncRNAs and diseases are fused to estimate the new interaction value. To evaluate the performance of DMFLDA, we perform leave-one-out cross-validation and 5-fold cross-validation on known experimentally verified lncRNA-disease associations. The experimental results show that DMFLDA performs better than the existing methods. The case studies show that many predicted interactions of colorectal cancer, prostate cancer, and renal cancer have been verified by recent biomedical literature. The source code and datasets can be obtained from https://github.com/CSUBioGroup/DMFLDA.
Min Zeng 0004, Chengqian Lu, Zhihui Fei, Fang-Xiang Wu, Yaohang Li, Jianxin Wang 0001, Min Li 0007
IEEE ACM Trans. Comput. Biol. Bioinform.5
2021 A Deep Learning Framework for Gene Ontology Annotations With Sequence- and Network-Based Information
abstract
Knowledge of protein functions plays an important role in biology and medicine. With the rapid development of high-throughput technologies, a huge number of proteins have been discovered. However, there are a great number of proteins without functional annotations. A protein usually has multiple functions and some functions or biological processes require interactions of a plurality of proteins. Additionally, Gene Ontology provides a useful classification for protein functions and contains more than 40,000 terms. We propose a deep learning framework called DeepGOA to predict protein functions with protein sequences and protein-protein interaction (PPI) networks. For protein sequences, we extract two types of information: sequence semantic information and subsequence-based features. We use the word2vec technique to numerically represent protein sequences, and utilize a Bi-directional Long and Short Time Memory (Bi-LSTM) and multi-scale convolutional neural network (multi-scale CNN) to obtain the global and local semantic features of protein sequences, respectively. Additionally, we use the InterPro tool to scan protein sequences for extracting subsequence-based information, such as domains and motifs. Then, the information is plugged into a neural network to generate high-quality features. For the PPI network, the Deepwalk algorithm is applied to generate its embedding information of PPI. Then the two types of features are concatenated together to predict protein functions. To evaluate the performance of DeepGOA, several different evaluation methods and metrics are utilized. The experimental results show that DeepGOA outperforms DeepGO and BLAST.
Fuhao Zhang, Hong Song 0004, Min Zeng 0004, Fang-Xiang Wu, Yaohang Li, Yi Pan 0001, Min Li 0007
IEEE ACM Trans. Comput. Biol. Bioinform.5
2020 A novel approach based on deep residual learning to predict drug's anatomical therapeutic chemical code
abstract
Correctly identifying the potential Anatomical Therapeutic Chemical (ATC) codes for drugs can accelerate drug development and reduce the cost of experiments. However, most of the existing methods only analyze the first-level ATC code of drugs and lack of the ability to learn basic features from sparsely known drug-ATC code associations. In this paper, we propose a novel method based on deep residual network framework, named RNPredATC, to predict potential drug-ATC code associations by integrating drug structure similarity, ATC sematic similarity, and known drug-ATC code associations. RNPredATC can extract dense feature vectors from sparsely known drug-ATC code associations and reduce the impact from degradation problem, such as gradient vanishing or gradient explosion of deep network. The experimental results show that RNPredATC achieves better performances.
Yaohang Li, Jianxin Wang 0001
BIBM4
2020 cFAT-GAN: Conditional Simulation of Electron-Proton Scattering Events with Variate Beam Energies by a Feature Augmented and Transformed Generative Adversarial Network
abstract
The recently proposed generative adversarial network (GAN) based event generator, the Feature Augmented and Transformed GAN (FAT-GAN), has shown impressive capability of reproducing inclusive electron-proton scattering events at a given collision energy. In contrast, many practical applications require the event generator to have the flexibility of allowing users to specify the reaction energy as an input to produce the corresponding synthetic events. In this work, we extend the FAT-GAN framework by conditioning the component neural networks according to the given reaction energy. We demonstrate that this model, referred to as cFAT-GAN, can reliably produce inclusive event feature distributions and correlations for a continuous range of reaction energies by automatically interpolating and extrapolating from a set of trained energies. We employ a continuous energy feature representation to enable the networks to organically learn the distribution relationships between different reaction energies, laying the groundwork for accessing events at untrained energies.This continuous conditional energy provides a degree of versatility to the cFAT-GAN for its further development as a significant research tool in high-energy and nuclear physics.
Luisa Velasco, Evan McClellan, Nobuo Sato, Pawel Ambrozewicz, Tianbo Liu 0004, Wally Melnitchouk, Michelle P. Kuchera, Yasir Alanazi, Yaohang Li
ICMLA9
2020 De novo Prediction of Drug-Target Interaction via Laplacian Regularized Schatten-p Norm Minimization
Gaoyan Wu, Mengyun Yang, Yaohang Li, Jianxin Wang 0001
ISBRA3
2020 CLPred: a sequence-based protein crystallization predictor using BLSTM neural network
abstract
MOTIVATION: Determining the structures of proteins is a critical step to understand their biological functions. Crystallography-based X-ray diffraction technique is the main method for experimental protein structure determination. However, the underlying crystallization process, which needs multiple time-consuming and costly experimental steps, has a high attrition rate. To overcome this issue, a series of in silico methods have been developed with the primary aim of selecting the protein sequences that are promising to be crystallized. However, the predictive performance of the current methods is modest. RESULTS: We propose a deep learning model, so-called CLPred, which uses a bidirectional recurrent neural network with long short-term memory (BLSTM) to capture the long-range interaction patterns between k-mers amino acids to predict protein crystallizability. Using sequence only information, CLPred outperforms the existing deep-learning predictors and a vast majority of sequence-based diffraction-quality crystals predictors on three independent test sets. The results highlight the effectiveness of BLSTM in capturing non-local, long-range inter-peptide interaction patterns to distinguish proteins that can result in diffraction-quality crystals from those that cannot. CLPred has been steadily improved over the previous window-based neural networks, which is able to predict crystallization propensity with high accuracy. CLPred can also be improved significantly if it incorporates additional features from pre-extracted evolutional, structural and physicochemical characteristics. The correctness of CLPred predictions is further validated by the case studies of Sox transcription factor family member proteins and Zika virus non-structural proteins. AVAILABILITY AND IMPLEMENTATION: https://github.com/xuanwenjing/CLPred.
Wenjing Xuan, Yaohang Li, Jianxin Wang 0001
Bioinform.4
2020 Protein-protein interaction site prediction through combining local and global features with deep neural networks
abstract
MOTIVATION: Protein-protein interactions (PPIs) play important roles in many biological processes. Conventional biological experiments for identifying PPI sites are costly and time-consuming. Thus, many computational approaches have been proposed to predict PPI sites. Existing computational methods usually use local contextual features to predict PPI sites. Actually, global features of protein sequences are critical for PPI site prediction. RESULTS: A new end-to-end deep learning framework, named DeepPPISP, through combining local contextual and global sequence features, is proposed for PPI site prediction. For local contextual features, we use a sliding window to capture features of neighbors of a target amino acid as in previous studies. For global sequence features, a text convolutional neural network is applied to extract features from the whole protein sequence. Then the local contextual and global sequence features are combined to predict PPI sites. By integrating local contextual and global sequence features, DeepPPISP achieves the state-of-the-art performance, which is better than the other competing methods. In order to investigate if global sequence features are helpful in our deep learning model, we remove or change some components in DeepPPISP. Detailed analyses show that global sequence features play important roles in DeepPPISP. AVAILABILITY AND IMPLEMENTATION: The DeepPPISP web server is available at http://bioinformatics.csu.edu.cn/PPISP/. The source code can be obtained from https://github.com/CSUBioGroup/DeepPPISP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Min Zeng 0004, Fuhao Zhang, Fang-Xiang Wu, Yaohang Li, Jianxin Wang 0001, Min Li 0007
Bioinform.4
2020 DeepFrag-k: a fragment-based deep learning approach for protein fold recognition
abstract
BACKGROUND: One of the most essential problems in structural bioinformatics is protein fold recognition. In this paper, we design a novel deep learning architecture, so-called DeepFrag-k, which identifies fold discriminative features at fragment level to improve the accuracy of protein fold recognition. DeepFrag-k is composed of two stages: the first stage employs a multi-modal Deep Belief Network (DBN) to predict the potential structural fragments given a sequence, represented as a fragment vector, and then the second stage uses a deep convolutional neural network (CNN) to classify the fragment vector into the corresponding fold. RESULTS: Our results show that DeepFrag-k yields 92.98% accuracy in predicting the top-100 most popular fragments, which can be used to generate discriminative fragment feature vectors to improve protein fold recognition. CONCLUSIONS: There is a set of fragments that can serve as structural "keywords" distinguishing between major protein folds. The deep learning architecture in DeepFrag-k is able to accurately identify these fragments as structure features to improve protein fold recognition.
Wessam Elhefnawy, Min Li 0007, Jianxin Wang 0001, Yaohang Li
BMC Bioinform.4
2020 United Neighborhood Closeness Centrality and Orthology for Predicting Essential Proteins
abstract
Identifying essential proteins plays an important role in disease study, drug design, and understanding the minimal requirement for cellular life. Computational methods for essential proteins discovery overcome the disadvantages of biological experimental methods that are often time-consuming, expensive, and inefficient. The topological features of protein-protein interaction (PPI) networks are often used to design computational prediction methods, such as Degree Centrality (DC), Betweenness Centrality (BC), Closeness Centrality (CC), Subgraph Centrality (SC), Eigenvector Centrality (EC), Information Centrality (IC), and Neighborhood Centrality (NC). However, the prediction accuracies of these individual methods still have space to be improved. Studies show that additional information, such as orthologous relations, helps discover essential proteins. Many researchers have proposed different methods by combining multiple information sources to gain improvement of prediction accuracy. In this study, we find that essential proteins appear in triangular structure in PPI network significantly more often than nonessential ones. Based on this phenomenon, we propose a novel pure centrality measure, so-called Neighborhood Closeness Centrality (NCC). Accordingly, we develop a new combination model, Extended Pareto Optimality Consensus model, named EPOC, to fuse NCC and Orthology information and a novel essential proteins identification method, NCCO, is fully proposed. Compared with seven existing classic centrality methods (DC, BC, IC, CC, SC, EC, and NC) and three consensus methods (PeC, ION, and CSC), our results on S.cerevisiae and E.coli datasets show that NCCO has clear advantages. As a consensus method, EPOC also yields better performance than the random walk model.
Gaoshi Li, Min Li 0007, Jianxin Wang 0001, Yaohang Li, Yi Pan 0001
IEEE ACM Trans. Comput. Biol. Bioinform.4
2020 Identification of Protein Complexes by Using a Spatial and Temporal Active Protein Interaction Network
abstract
The rapid development of proteomics and high-throughput technologies has produced a large amount of Protein-Protein Interaction (PPI) data, which makes it possible for considering dynamic properties of protein interaction networks (PINs) instead of static properties. Identification of protein complexes from dynamic PINs becomes a vital scientific problem for understanding cellular life in the post genome era. Up to now, plenty of models or methods have been proposed for the construction of dynamic PINs to identify protein complexes. However, most of the constructed dynamic PINs just focus on the temporal dynamic information and thus overlook the spatial dynamic information of the complex biological systems. To address the limitation of the existing dynamic PIN analysis approaches, in this paper, we propose a new model-based scheme for the construction of the Spatial and Temporal Active Protein Interaction Network (ST-APIN) by integrating time-course gene expression data and subcellular location information. To evaluate the efficiency of ST-APIN, the commonly used classical clustering algorithm MCL is adopted to identify protein complexes from ST-APIN and the other three dynamic PINs, NF-APIN, DPIN, and TC-PIN. The experimental results show that, the performance of MCL on ST-APIN outperforms those on the other three dynamic PINs in terms of matching with known complexes, sensitivity, specificity, and f-measure. Furthermore, we evaluate the identified protein complexes by Gene Ontology (GO) function enrichment analysis. The validation shows that the identified protein complexes from ST-APIN are more biologically significant. This study provides a general paradigm for constructing the ST-APINs, which is essential for further understanding of molecular systems and the biomedical mechanism of complex diseases.
Min Li 0007, Xiangmao Meng, Ruiqing Zheng, Fang-Xiang Wu, Yaohang Li, Yi Pan 0001, Jianxin Wang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.5
2020 Constructing Disease Similarity Networks Based on Disease Module Theory
abstract
Quantifying the associations between diseases is now playing an important role in modern biology and medicine. Actually discovering associations between diseases could help us gain deeper insights into pathogenic mechanisms of complex diseases, thus could lead to improvements in disease diagnosis, drug repositioning, and drug development. Due to the growing body of high-throughput biological data, a number of methods have been developed for computing similarity between diseases during the past decade. However, these methods rarely consider the interconnections of genes related to each disease in protein-protein interaction network (PPIN). Recently, the disease module theory has been proposed, which states that disease-related genes or proteins tend to interact with each other in the same neighborhood of a PPIN. In this study, we propose a new method called ModuleSim to measure associations between diseases by using disease-gene association data and PPIN data based on disease module theory. The experimental results show that by considering the interactions between disease modules and their modularity, the disease similarity calculated by ModuleSim has a significant correlation with disease classification of Disease Ontology (DO). Furthermore, ModuleSim outperforms other four popular methods which are all using disease-gene association data and PPIN data to measure disease-disease associations. In addition, the disease similarity network constructed by MoudleSim suggests that ModuleSim is capable of finding potential associations between diseases.
Jianxin Wang 0001, Ping Zhong 0002, Yaohang Li, Fang-Xiang Wu, Yi Pan 0001
IEEE ACM Trans. Comput. Biol. Bioinform.4
2020 miRTMC: A miRNA Target Prediction Method Based on Matrix Completion Algorithm
abstract
microRNAs (miRNAs) are small non-coding RNAs which modulate the stability of gene targets and their rates of translation into proteins at transcriptional level and post-transcriptional level. miRNA dysfunctions can lead to human diseases because of dysregulation of their targets. Correct miRNA target prediction will lead to better understanding of the mechanisms of human diseases and provide hints on curing them. In recent years, computational miRNA target prediction methods have been proposed according to the interaction rules between miRNAs and targets. However, these methods suffer from high false positive rates due to the complicated relationship between miRNAs and their targets. The rapidly growing number of experimentally validated miRNA targets enables predicting miRNA targets with high precision via accurate data analysis. Taking advantage of these known miRNA targets, a novel recommendation system model (miRTMC) for miRNA target prediction is established using a new matrix completion algorithm. In miRTMC, a heterogeneous network is constructed by integrating the miRNA similarity network, the gene similarity network, and the miRNA-gene interaction network. Our assumption is that the latent factors determining whether a gene is the target of miRNA or not are highly correlated, i.e., the adjacency matrix of the heterogeneous network is low-rank, which is then completed by using a nuclear norm regularized linear least squares model under non-negative constraints. Alternating direction method of multipliers (ADMM) is adopted to numerically solve the matrix completion problem. Our results show that miRTMC outperforms the competing methods in terms of various evaluation metrics. Our software package is available at https://github.com/hjiangcsu/miRTMC.
Hui Jiang 0008, Mengyun Yang, Xiang Chen 0029, Min Li 0007, Yaohang Li, Jianxin Wang 0001
IEEE J. Biomed. Health Informatics5
2020 Predicting Human lncRNA-Disease Associations Based on Geometric Matrix Completion
abstract
Recently, increasing evidences reveal that dysregulations of long non-coding RNAs (lncRNAs) are relevant to diverse diseases. However, the number of experimentally verified lncRNA-disease associations is limited. Prioritizing potential associations is beneficial not only for disease diagnosis, but also disease treatment, more important apprehending disease mechanisms at lncRNA level. Various computational methods have been proposed, but precise prediction and full use of data's intrinsic structure are still challenging. In this work, we design a new method, denominated GMCLDA (Geometric Matrix Completion lncRNA-Disease Association), to infer underlying associations based on geometric matrix completion. Utilizing association patterns among functionally similar lncRNAs and phenotypically similar diseases, GMCLCA makes use of the intrinsic structure embedded in the association matrix. Besides, limiting the scope of the predicted values gives rise to a certain sparsity in computation and enhances the robustness of GMCLDA. GMCLDA computes disease semantic similarity according to the Disease Ontology (DO) hierarchy and lncRNA Gaussian interaction profile kernel similarity according to known interaction profiles. Then, GMCLDA measures lncRNA sequence similarity using Needleman-Wunsch algorithm. For a new lncRNA, GMCLDA prefills interaction profile on account of its K-nearest neighbors defined by sequence similarity. Finally, GMCLDA estimates the missing entries of the association matrix based on geometric matrix completion model. Compared with state-of-the-art methods, GMCLDA can provide more accurate lncRNA-disease prediction. Further case studies prove that GMCLDA is able to correctly infer possible lncRNAs for renal cancer.
Chengqian Lu, Mengyun Yang, Min Li 0007, Yaohang Li, Fang-Xiang Wu, Jianxin Wang 0001
IEEE J. Biomed. Health Informatics4
2019 LncRNA-disease association prediction through combining linear and non-linear features with matrix factorization and deep learning techniques
abstract
Long non-coding RNAs (lncRNAs) are the foundation for understanding mechanisms of many human diseases. Considering the limited number of known experimentally verified associations between lncRNAs and diseases, it is appealing to develop accurate and effective computational methods to identify lncRNA-disease associations. Conventional matrix factorization-based methods cannot model complicated associations between lncRNAs and diseases. In this study, we propose a novel computational framework, through combining linear and non-linear features, which is used for lncRNA-disease association prediction. In our model, a conventional matrix factorization method is applied to extract linear features between lncRNAs and diseases. Deep learning techniques (fully connected layers) are applied to extract nonlinear features between lncRNAs and diseases. Finally, linear and non-linear features are fused to improve predictive performance. Compared to previous studies, our model can take advantages of the combination of linear and non-linear features between lncRNAs and diseases, and thus can effectively identify potential lncRNA-disease associations. The results show that our method achieves state-of-the-art performance in the leave-one-out cross-validation. The source codes of our method can be found at https://github.com/CSUBioGroup/DMFLDA2.
Min Zeng 0004, Chengqian Lu, Fuhao Zhang, Zhangli Lu, Fang-Xiang Wu, Yaohang Li, Min Li 0007
BIBM6
2019 AttentionDTA: prediction of drug-target binding affinity using attention model
abstract
In bioinformatics, machine learning-based prediction of drug-target interaction (DTI) plays an important role in virtual screening of drug discovery. DTI prediction, which have been treated as a binary classification problem, depends on the concentration of two molecules, the interaction between two molecules, and other factors. The degree of affinity between a drug molecule (such as a drug compound) and a target molecule (such as a receptor or protein kinase) reflects how tightly the drug binds to a particular target and is quantified by the measurement which can reflect more detailed and specific information than binary relationship. In this study, we proposed an end-to-end model, named AttentionDTA, based on deep learning, which associates attention mechanism to predict the binding affinity of DTI. The novelty in this work is to use attentional mechanisms to consider which subsequences in a protein are more important for a drug and which subsequences in a drug are more important for a protein when predicting its affinity. So that the representational ability of the model is stronger. The model uses one-dimensional Convolution Neural Networks (1D-CNNs) to extract the abstract information of drug and protein, and makes the drug and protein representations mutually adapt through the attention mechanisms. We evaluate our model on two established drug-target affinity benchmark datasets, Davis and KIBA. The model outperforms DeepDTA, a state-of-the-art deep learning method for drug-target binding affinity prediction, with better Mean Squared Error (MSE), Concordance Index (CI), rm2, and Area Under Precision Recall Curve (AUPR). Our results show that the attention-based model can effectively extract effective representations by calculating the weight of the representation between the drug and the protein. Finally, we visualize the attention weight. It proves our model can obtain the information of binding sites.
Qichang Zhao, Fen Xiao, Mengyun Yang, Yaohang Li, Jianxin Wang 0001
BIBM4
2019 Drug repositioning based on bounded nuclear norm regularization
abstract
MOTIVATION: Computational drug repositioning is a cost-effective strategy to identify novel indications for existing drugs. Drug repositioning is often modeled as a recommendation system problem. Taking advantage of the known drug-disease associations, the objective of the recommendation system is to identify new treatments by filling out the unknown entries in the drug-disease association matrix, which is known as matrix completion. Underpinned by the fact that common molecular pathways contribute to many different diseases, the recommendation system assumes that the underlying latent factors determining drug-disease associations are highly correlated. In other words, the drug-disease matrix to be completed is low-rank. Accordingly, matrix completion algorithms efficiently constructing low-rank drug-disease matrix approximations consistent with known associations can be of immense help in discovering the novel drug-disease associations. RESULTS: In this article, we propose to use a bounded nuclear norm regularization (BNNR) method to complete the drug-disease matrix under the low-rank assumption. Instead of strictly fitting the known elements, BNNR is designed to tolerate the noisy drug-drug and disease-disease similarities by incorporating a regularization term to balance the approximation error and the rank properties. Moreover, additional constraints are incorporated into BNNR to ensure that all predicted matrix entry values are within the specific interval. BNNR is carried out on an adjacency matrix of a heterogeneous drug-disease network, which integrates the drug-drug, drug-disease and disease-disease networks. It not only makes full use of available drugs, diseases and their association information, but also is capable of dealing with cold start naturally. Our computational results show that BNNR yields higher drug-disease association prediction accuracy than the current state-of-the-art methods. The most significant gain is in prediction precision measured as the fraction of the positive predictions that are truly positive, which is particularly useful in drug design practice. Cases studies also confirm the accuracy and reliability of BNNR. AVAILABILITY AND IMPLEMENTATION: The code of BNNR is freely available at https://github.com/BioinformaticsCSU/BNNR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mengyun Yang, Huimin Luo, Yaohang Li, Jianxin Wang 0001
Bioinform.3
2019 DeepEP: a deep learning framework for identifying essential proteins
abstract
BACKGROUND: Essential proteins are crucial for cellular life and thus, identification of essential proteins is an important topic and a challenging problem for researchers. Recently lots of computational approaches have been proposed to handle this problem. However, traditional centrality methods cannot fully represent the topological features of biological networks. In addition, identifying essential proteins is an imbalanced learning problem; but few current shallow machine learning-based methods are designed to handle the imbalanced characteristics. RESULTS: We develop DeepEP based on a deep learning framework that uses the node2vec technique, multi-scale convolutional neural networks and a sampling technique to identify essential proteins. In DeepEP, the node2vec technique is applied to automatically learn topological and semantic features for each protein in protein-protein interaction (PPI) network. Gene expression profiles are treated as images and multi-scale convolutional neural networks are applied to extract their patterns. In addition, DeepEP uses a sampling method to alleviate the imbalanced characteristics. The sampling method samples the same number of the majority and minority samples in a training epoch, which is not biased to any class in training process. The experimental results show that DeepEP outperforms traditional centrality methods. Moreover, DeepEP is better than shallow machine learning-based methods. Detailed analyses show that the dense vectors which are generated by node2vec technique contribute a lot to the improved performance. It is clear that the node2vec technique effectively captures the topological and semantic properties of PPI network. The sampling method also improves the performance of identifying essential proteins. CONCLUSION: We demonstrate that DeepEP improves the prediction performance by integrating multiple deep learning techniques and a sampling method. DeepEP is more effective than existing methods.
Min Zeng 0004, Min Li 0007, Fang-Xiang Wu, Yaohang Li, Yi Pan 0001
BMC Bioinform.4
2019 Decoding the Structural Keywords in Protein Structure Universe
Wessam Elhefnawy, Min Li 0007, Jianxin Wang 0001, Yaohang Li
J. Comput. Sci. Technol.4
2019 Overlap matrix completion for predicting drug-associated indications
abstract
Identification of potential drug-associated indications is critical for either approved or novel drugs in drug repositioning. Current computational methods based on drug similarity and disease similarity have been developed to predict drug-disease associations. When more reliable drug- or disease-related information becomes available and is integrated, the prediction precision can be continuously improved. However, it is a challenging problem to effectively incorporate multiple types of prior information, representing different characteristics of drugs and diseases, to identify promising drug-disease associations. In this study, we propose an overlap matrix completion (OMC) for bilayer networks (OMC2) and tri-layer networks (OMC3) to predict potential drug-associated indications, respectively. OMC is able to efficiently exploit the underlying low-rank structures of the drug-disease association matrices. In OMC2, first of all, we construct one bilayer network from drug-side aspect and one from disease-side aspect, and then obtain their corresponding block adjacency matrices. We then propose the OMC2 algorithm to fill out the values of the missing entries in these two adjacency matrices, and predict the scores of unknown drug-disease pairs. Moreover, we further extend OMC2 to OMC3 to handle tri-layer networks. Computational experiments on various datasets indicate that our OMC methods can effectively predict the potential drug-disease associations. Compared with the other state-of-the-art approaches, our methods yield higher prediction accuracy in 10-fold cross-validation and de novo experiments. In addition, case studies also confirm the effectiveness of our methods in identifying promising indications for existing drugs in practical applications.
Mengyun Yang, Huimin Luo, Yaohang Li, Fang-Xiang Wu, Jianxin Wang 0001
PLoS Comput. Biol.3
2019 Automated ICD-9 Coding via A Deep Learning Approach
abstract
ICD-9 (the Ninth Revision of International Classification of Diseases) is widely used to describe a patient's diagnosis. Accurate automated ICD-9 coding is important because manual coding is expensive, time-consuming, and inefficient. Inspired by the recent successes of deep learning, in this study, we present a deep learning framework called DeepLabeler to automatically assign ICD-9 codes. DeepLabeler combines the convolutional neural network with the 'Document to Vector' technique to extract and encode local and global features. Our proposed DeepLabeler demonstrates its effectiveness by achieving state-of-the-art performance, i.e., 0.335 micro F-measure on MIMIC-II dataset and 0.408 micro F-measure on MIMIC-III dataset. It outperforms classical hierarchy-based SVM and flat-SVM both on these two datasets by at least 14 percent. Furthermore, we analyze the deep neural network structure to discover the vital elements in the success of DeepLabeler. We find that the convolutional neural network is the most effective component in our network and the 'Document to Vector' technique is also necessary for enhancing classification performance since it extracts well-recognized global features. Extensive experimental results demonstrate that the great promise of deep learning techniques in the field of text multi-label classification and automated medical coding.
Min Li 0007, Zhihui Fei, Min Zeng 0004, Fang-Xiang Wu, Yaohang Li, Yi Pan 0001, Jianxin Wang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.5
2019 MGT-SM: A Method for Constructing Cellular Signal Transduction Networks
abstract
A cellular signal transduction network is an important means to describe biological responses to environmental stimuli and exchange of biological signals. Constructing the cellular signal transduction network provides an important basis for the study of the biological activities, the mechanism of the diseases, drug targets and so on. The statistical approaches to network inference are popular in literature. Granger test has been used as an effective method for causality inference. Compared with bivariate granger tests, multivariate granger tests reduce the indirect causality and were used widely for the construction of cellular signal transduction networks. A multivariate Granger test requires that the number of time points in the time-series data is more than the number of nodes involved in the network. However, there are many real datasets with a few time points which are much less than the number of nodes in the network. In this study, we propose a new multivariate Granger test-based framework to construct cellular signal transduction network, called MGT-SM. Our MGT-SM uses SVD to compute the coefficient matrix from gene expression data and adopts Monte Carlo simulation to estimate the significance of directed edges in the constructed networks. We apply the proposed MGT-SM to Yeast Synthetic Network and MDA-MB-468, and evaluate its performance in terms of the recall and the AUC. The results show that MGT-SM achieves better results, compared with other popular methods (CGC2SPR, PGC, and DBN).
Min Li 0007, Ruiqing Zheng, Yaohang Li, Fang-Xiang Wu, Jianxin Wang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.3
2018 Disease Inference with Symptom Extraction and Bidirectional Recurrent Neural Network
Donglin Guo, Min Li 0007, Yaohang Li, Guihua Duan, Fang-Xiang Wu, Jianxin Wang 0001
BIBM4
2018 A Deep Learning Framework for Identifying Essential Proteins Based on Protein-Protein Interaction Network and Gene Expression Data
Min Zeng 0004, Min Li 0007, Zhihui Fei, Fang-Xiang Wu, Yaohang Li, Yi Pan 0001
BIBM5
2018 Using Deep Neural Network to Predict Drug Sensitivity of Cancer Cell Lines
Yake Wang, Min Li 0007, Ruiqing Zheng, Xinghua Shi, Yaohang Li, Fang-Xiang Wu, Jianxin Wang 0001
ICIC (2)5
2018 Faster Matrix Completion Using Randomized SVD
abstract
Matrix completion is a widely used technique for image inpainting and personalized recommender system, etc. In this work, we focus on accelerating the matrix completion using faster randomized singular value decomposition (rSVD). Firstly, two fast randomized algorithms (rSVD-PI and rSVDBKI) are proposed for handling sparse matrix. They make use of an eigSVD procedure and several accelerating skills. Then, with the rSVD-BKI algorithm and a new subspace recycling technique, we accelerate the singular value thresholding (SVT) method in [1] to realize faster matrix completion. Experiments show that the proposed rSVD algorithms can be 6× faster than the basic rSVD algorithm [2] while keeping same accuracy. For image inpainting and movie-rating estimation problems (including up to 2 × 107ratings), the proposed accelerated SVT algorithm consumes 15× and 8× less CPU time than the methods using svds and lansvd respectively, without loss of accuracy.
Wenjian Yu, Yaohang Li
ICTAI3
2018 PBMarsNet: A Multivariate Adaptive Regression Splines Based Method to Reconstruct Gene Regulatory Networks
Ruiqing Zheng, Xiang Chen 0029, Yaohang Li, Fang-Xiang Wu, Min Li 0007
ISBRA4
2018 Prediction of lncRNA-disease associations based on inductive matrix completion
abstract
Motivation: Accumulating evidences indicate that long non-coding RNAs (lncRNAs) play pivotal roles in various biological processes. Mutations and dysregulations of lncRNAs are implicated in miscellaneous human diseases. Predicting lncRNA-disease associations is beneficial to disease diagnosis as well as treatment. Although many computational methods have been developed, precisely identifying lncRNA-disease associations, especially for novel lncRNAs, remains challenging. Results: In this study, we propose a method (named SIMCLDA) for predicting potential lncRNA-disease associations based on inductive matrix completion. We compute Gaussian interaction profile kernel of lncRNAs from known lncRNA-disease interactions and functional similarity of diseases based on disease-gene and gene-gene onotology associations. Then, we extract primary feature vectors from Gaussian interaction profile kernel of lncRNAs and functional similarity of diseases by principal component analysis, respectively. For a new lncRNA, we calculate the interaction profile according to the interaction profiles of its neighbors. At last, we complete the association matrix based on the inductive matrix completion framework using the primary feature vectors from the constructed feature matrices. Computational results show that SIMCLDA can effectively predict lncRNA-disease associations with higher accuracy compared with previous methods. Furthermore, case studies show that SIMCLDA can effectively predict candidate lncRNAs for renal cancer, gastric cancer and prostate cancer. Availability and implementation: https://github.com//bioinfomaticsCSU/SIMCLDA. Supplementary information: Supplementary data are available at Bioinformatics online.
Chengqian Lu, Mengyun Yang, Feng Luo 0001, Fang-Xiang Wu, Min Li 0007, Yi Pan 0001, Yaohang Li, Jianxin Wang 0001
Bioinform.7
2018 Computational drug repositioning using low-rank matrix approximation and randomized algorithms
abstract
Motivation: Computational drug repositioning is an important and efficient approach towards identifying novel treatments for diseases in drug discovery. The emergence of large-scale, heterogeneous biological and biomedical datasets has provided an unprecedented opportunity for developing computational drug repositioning methods. The drug repositioning problem can be modeled as a recommendation system that recommends novel treatments based on known drug-disease associations. The formulation under this recommendation system is matrix completion, assuming that the hidden factors contributing to drug-disease associations are highly correlated and thus the corresponding data matrix is low-rank. Under this assumption, the matrix completion algorithm fills out the unknown entries in the drug-disease matrix by constructing a low-rank matrix approximation, where new drug-disease associations having not been validated can be screened. Results: In this work, we propose a drug repositioning recommendation system (DRRS) to predict novel drug indications by integrating related data sources and validated information of drugs and diseases. Firstly, we construct a heterogeneous drug-disease interaction network by integrating drug-drug, disease-disease and drug-disease networks. The heterogeneous network is represented by a large drug-disease adjacency matrix, whose entries include drug pairs, disease pairs, known drug-disease interaction pairs and unknown drug-disease pairs. Then, we adopt a fast Singular Value Thresholding (SVT) algorithm to complete the drug-disease adjacency matrix with predicted scores for unknown drug-disease pairs. The comprehensive experimental results show that DRRS improves the prediction accuracy compared with the other state-of-the-art approaches. In addition, case studies for several selected drugs further demonstrate the practical usefulness of the proposed method. Availability and implementation: http://bioinformatics.csu.edu.cn/resources/softs/DrugRepositioning/DRRS/index.html. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Huimin Luo, Min Li 0007, Shaokai Wang, Yaohang Li, Jianxin Wang 0001
Bioinform.5
2017 Single-Pass PCA of Large High-Dimensional Data
abstract
Principal component analysis (PCA) is a fundamental dimension reduction tool in statistics and machine learning. For large and high-dimensional data, computing the PCA (i.e., the top singular vectors of the data matrix) becomes a challenging task. In this work, a single-pass randomized algorithm is proposed to compute PCA with only one pass over the data. It is suitable for processing extremely large and high-dimensional data stored in slow memory (hard disk) or the data generated in a streaming fashion. Experiments with synthetic and real data validate the algorithm's accuracy, which has orders of magnitude smaller error than an existing single-pass algorithm. For a set of high-dimensional data stored as a 150 GB file, the algorithm is able to compute the first 50 principal components in just 24 minutes on a typical 24-core computer, with less than 1 GB memory cost.
Wenjian Yu, Shenghua Liu, Yaohang Li
IJCAI5
2017 Construction of Protein Backbone Fragments Libraries on Large Protein Sets Using a Randomized Spectral Clustering Algorithm
Wessam Elhefnawy, Min Li 0007, Jianxin Wang 0001, Yaohang Li
ISBRA4
2017 Relating Diseases Based on Disease Module Theory
Min Li 0007, Ping Zhong 0002, Guihua Duan, Jianxin Wang 0001, Yaohang Li, Fang-Xiang Wu
ISBRA6
2017 Calcium Ion Fluctuations Alter Channel Gating in a Stochastic Luminal Calcium Release Site Model
abstract
Stochasticity and small system size effects in complex biochemical reaction networks can greatly alter transient and steady-state system properties. A common approach to modeling reaction networks, which accounts for system size, is the chemical master equation that governs the dynamics of the joint probability distribution for molecular copy number. However, calculation of the stationary distribution is often prohibitive, due to the large state-space associated with most biochemical reaction networks. Here, we analyze a network representing a luminal calcium release site model and investigate to what extent small system size effects and calcium fluctuations, driven by ion channel gating, influx and diffusion, alter steady-state ion channel properties including open probability. For a physiological ion channel gating model and number of channels, the state-space may be between approximately$10^6-10^8$elements, and a novel modified block power method is used to solve the associated dominant eigenvector problem required to calculate the stationary distribution. We demonstrate that both small local cytosolic domain volume and a small number of ion channels drive calcium fluctuations that result in deviation from the corresponding model that neglects small system size effects.
Hao Ji 0005, Yaohang Li, Seth H. Weinberg
IEEE ACM Trans. Comput. Biol. Bioinform.2
2016 FLEXc: protein flexibility prediction using context-based statistics, predicted structural features, and sequence information
abstract
BACKGROUND: The fluctuation of atoms around their average positions in protein structures provides important information regarding protein dynamics. This flexibility of protein structures is associated with various biological processes. Predicting flexibility of residues from protein sequences is significant for analyzing the dynamic properties of proteins which will be helpful in predicting their functions. RESULTS: In this paper, an approach of improving the accuracy of protein flexibility prediction is introduced. A neural network method for predicting flexibility in 3 states is implemented. The method incorporates sequence and evolutionary information, context-based scores, predicted secondary structures and solvent accessibility, and amino acid properties. Context-based statistical scores are derived, using the mean-field potentials approach, for describing the different preferences of protein residues in flexibility states taking into consideration their amino acid context. The 7-fold cross validated accuracy reached 61 % when context-based scores and predicted structural states are incorporated in the training process of the flexibility predictor. CONCLUSIONS: Incorporating context-based statistical scores with predicted structural states are important features to improve the performance of predicting protein flexibility, as shown by our computational results. Our prediction method is implemented as web service called "FLEXc" and available online at: http://hpcr.cs.odu.edu/flexc .
Ashraf Yaseen, Mais Nijim, Brandon Williams, Min Li 0007, Jianxin Wang 0001, Yaohang Li
BMC Bioinform.7
2016 Predicting drug-target interaction using positive-unlabeled learning
Wei Lan 0001, Jianxin Wang 0001, Min Li 0007, Jin Liu 0012, Yaohang Li, Fang-Xiang Wu, Yi Pan 0001
Neurocomputing5
2016 A load-balancing workload distribution scheme for three-body interaction computation on Graphics Processing Units (GPU)
Ashraf Yaseen, Hao Ji 0005, Yaohang Li
J. Parallel Distributed Comput.3
2015 Calcium Ion Fluctuations Alter Channel Gating in a Stochastic Luminal Calcium Release Site Model
Hao Ji 0005, Yaohang Li, Seth H. Weinberg
ISBRA2
2015 Using Workflow Technology to Create Scenario-based Workflows for Information Security Education: Scenario-based Workflows (Abstract Only)
abstract
Teaching information security courses is technically challenging. In an information security course, students and instructors often end up struggling in low-level and complicated software installation, system setup, service configuration, command operations, and data manipulation while losing concentration in learning the important information security principles. To help students in information security courses learn information security principles more effectively and efficiently, we used the workflow technology to create scenario-based workflows in order to improve the effectiveness of teaching and learning of several key information security principles and techniques. Two case studies simulating real-life scenarios, including one for an online banking system and one for an online grading system, are recreated within a laboratory setting using workflow technology and are then presented in information security classes. Our educational practice shows that the benefits of using workflow technology in information security education have been well received by students.
Wu He, Ashish Kshirsagar, Alexander C. Nwala, Yaohang Li
SIGCSE4
2014 Template-based C8-SCORPION: a protein 8-state secondary structure prediction method using structural information and context-based features
abstract
BACKGROUND: Secondary structures prediction of proteins is important to many protein structure modeling applications. Correct prediction of secondary structures can significantly reduce the degrees of freedom in protein tertiary structure modeling and therefore reduces the difficulty of obtaining high resolution 3D models. METHODS: In this work, we investigate a template-based approach to enhance 8-state secondary structure prediction accuracy. We construct structural templates from known protein structures with certain sequence similarity. The structural templates are then incorporated as features with sequence and evolutionary information to train two-stage neural networks. In case of structural templates absence, heuristic structural information is incorporated instead. RESULTS: After applying the template-based 8-state secondary structure prediction method, the 7-fold cross-validated Q8 accuracy is 78.85%. Even templates from structures with only 20%~30% sequence similarity can help improve the 8-state prediction accuracy. More importantly, when good templates are available, the prediction accuracy of less frequent secondary structures, such as 3-10 helices, turns, and bends, are highly improved, which are useful for practical applications. CONCLUSIONS: Our computational results show that the templates containing structural information are effective features to enhance 8-state secondary structure predictions. Our prediction algorithm is implemented on a web server named "C8-SCORPION" available at: http://hpcr.cs.odu.edu/c8scorpion.
Ashraf Yaseen, Yaohang Li
BMC Bioinform.2
2013 Dinosolve: a protein disulfide bonding prediction server using context-based features to enhance prediction accuracy
abstract
BACKGROUND: Disulfide bonds play an important role in protein folding and structure stability. Accurately predicting disulfide bonds from protein sequences is important for modeling the structural and functional characteristics of many proteins. METHODS: In this work, we introduce an approach of enhancing disulfide bonding prediction accuracy by taking advantage of context-based features. We firstly derive the first-order and second-order mean-force potentials according to the amino acid environment around the cysteine residues from large number of cysteine samples. The mean-force potentials are integrated as context-based scores to estimate the favorability of a cysteine residue in disulfide bonding state as well as a cysteine pair in disulfide bond connectivity. These context-based scores are then incorporated as features together with other sequence and evolutionary information to train neural networks for disulfide bonding state prediction and connectivity prediction. RESULTS: The 10-fold cross validated accuracy is 90.8% at residue-level and 85.6% at protein-level in classifying an individual cysteine residue as bonded or free, which is around 2% accuracy improvement. The average accuracy for disulfide bonding connectivity prediction is also improved, which yields overall sensitivity of 73.42% and specificity of 91.61%. CONCLUSIONS: Our computational results have shown that the context-based scores are effective features to enhance the prediction accuracies of both disulfide bonding state prediction and connectivity prediction. Our disulfide prediction algorithm is implemented on a web server named "Dinosolve" available at: http://hpcr.cs.odu.edu/dinosolve.
Ashraf Yaseen, Yaohang Li
BMC Bioinform.2
2013 High-dimensional MRI data analysis using a large-scale manifold learning approach
Loc Tran, Debrup Banerjee, Ashok J. Kumar, Frederic D. McKenzie, Yaohang Li, Jiang Li 0001
Mach. Vis. Appl.6
2012 Accelerating knowledge-based energy evaluation in protein structure modeling with Graphics Processing Units
Ashraf Yaseen, Yaohang Li
J. Parallel Distributed Comput.2
2010 Integrating multiple scoring functions to improve protein loop structure conformation space sampling
abstract
In this article, we present a new protein structure modeling approach based on multi-scoring functions sampling. The rationale is to integrate multiple carefully-selected physics-or knowledge-based scoring functions to tolerate insensitivity and inaccuracy existing in an individual scoring function so as to improve protein structure modeling accuracy. We apply the multi-scoring function sampling approach to protein loop backbone structure modeling. Our computational results show that sampling the scoring function space of a physics-based soft-sphere potential function and a knowledge-based scoring function based on pairwise atoms distance has led to resolution improvement in the predicted decoy populations in a set of 12-residue benchmark loop targets.
Yaohang Li, Ionel Rata, Eric Jakobsson
CIBCB1
2009 Resource-efficient computing paradigm for computational protein modeling applications
abstract
Many computational protein modeling applications using numerical methods such as Molecular Dynamics (MD), Monte Carlo (MC), or Genetic Algorithms (GA) require a large number of energy estimations of the protein molecular system. A typical energy function describing the protein energy is a combination of a number of terms characterizing various interactions within the protein molecule as well as the protein-solvent interactions. Evaluating the energy function of a relatively large protein molecule is rather computationally costly and usually occupies the major computation time in the protein simulation process. In this paper, we present a resource-efficient computing paradigm based on ldquoconsolidationrdquo to reduce the computational time of evaluating the energy function of large protein molecule. The fundamental idea of consolidation is to increase computational density to a computer in order to increase the CPU utilizations. Consolidation will be particularly efficient when the consolidated computations have heterogeneous resource demands. In computational protein modeling applications with costly energy function evaluation, we advocate the use of ldquothread consolidation,rdquo which is to spawn concurrent threads to carry out parallel energy function terms computations. Our computational results show that 7%~11% speedup in a protein loop structure prediction program on various hardware architectures where memory-intensive and computation-intensive terms coexist in the energy function. For an MD protein simulation program where computation-intensive energy function evaluations are divided and carried out by concurrent threads, we also find slight performance improvement when the thread consolidation technique is applied.
Yaohang Li, Douglas Wardell, Vincent W. Freeh
IPDPS1
2009 A decentralized parallel implementation for parallel tempering algorithm
Yaohang Li, Michael Mascagni, Andrey Gorin
Parallel Comput.1
2008 An adaptive and trustworthy software testing framework on the grid
Yaohang Li, Yongduan Song 0001
J. Supercomput.1
2007 Decentralized Replica Exchange Parallel Tempering: An Efficient Implementation of Parallel Tempering Using MPI and SPRNG
Yaohang Li, Michael Mascagni, Andrey Gorin
ICCSA (3)1
2007 Trustworthy remote compiling services for grid-based scientific applications
Yaohang Li, Xiaohong Yuan
J. Supercomput.1
2003 Improving Performance via Computational Replication on a Large-Scale Computational Grid
abstract
High performance computing on a large-scale computational grid is complicated by the heterogeneous computational capabilities of each node, node unavailability, and unreliable network connectivity. Replicating computation on multiple nodes can significantly improve performance by reducing task completion time on a grid's dynamic environment. We develop an analytical model to determine the number of task replicas to meet the performance goals in different computational grid configurations. Furthermore, taking advantage of the statistical nature of grid-based Monte Carlo applications, we extend the computational replication technique to an N-out-of-M scheduling strategy for grid-based Monte Carlo applications, which can potentially form a large category of grid-computing applications. In addition, we establish a corresponding model for the N-out-of-M scheduling mechanism. Simulations are used to validate the computational replication models. Our preliminary results show that the models we use are effective in predicting the required number of replicas to achieve short task completion time with a given high probability.
Yaohang Li, Michael Mascagni
CCGRID1
2003 Grid-based Nonequilibrium Multiple-Time Scale Molecular Dynamics/Brownian Dynamics Simulations of Ligand-Receptor Interactions in Structured Protein Systems
abstract
In the hybrid Molecular Dynamics (MD)/Brownian Dynamics (BD) algorithm for simulating the long-time, nonequilibrium dynamics of receptor-ligand interactions, the evaluation of the force autocorrelation function can be computationally costly but fortunately is highly amenable to multimode processing methods. In this paper, taking advantage of the computational grid's large-scale computational resources and the nice characteristics of grid-based Monte Carlo applications, we developed a grid-based receptor-ligand interactions simulation application using the MD/BD algorithm. We expect to provide high-performance and trustworthy computing for analyzing long-time dynamics of proteins and protein-protein interaction to predict and understand cell signaling processes and small molecule drug efficacies. Our preliminary results showed that our grid-based application could provide a faster and more accurate computation for the force autocorrelation function in our MD/BD simulation than previous parallel implementations.
Yaohang Li, Michael Mascagni, Michael H. Peter
CCGRID1