VLDB 2026 Research / reviewers in the wild / expert
Hui Liu 0026
dblp:93/4010-26
· DBLP profile ↗
25ranked-venue papers
9as first author
11since 2021 · last 2025
0000-0001-7158-913XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning Cross-Domain Representations for Transferable Drug Perturbations on Single-Cell Transcriptional ResponsesabstractPhenotypic drug discovery has attracted widespread attention because of its potential to identify bioactive molecules. Transcriptomic profiling provides a comprehensive reflection of phenotypic changes in cellular responses to external perturbations. In this paper, we propose XTransferCDR, a novel generative framework designed for feature decoupling and transferable representation learning across domains. Given a pair of perturbed expression profiles, our approach decouples the perturbation representations from basal states through domain separation encoders and then cross-transfers them in the latent space. The transferred representations are then used to reconstruct the corresponding perturbed expression profiles via a shared decoder. This cross-transfer constraint effectively promotes the learning of transferable drug perturbation representations. We conducted extensive evaluations of our model on multiple datasets, including single-cell transcriptional responses to drugs and single- and combinatorial genetic perturbations. The experimental results show that XTransferCDR achieved better performance than current state-of-the-art methods, showcasing its potential to advance phenotypic drug discovery. Hui Liu 0026, Shikai Jin |
AAAI | 1 |
| 2025 | Predicting Single-Cell Drug Sensitivity Utilizing Adaptive Weighted Features for Multi-Source Domain AdaptationabstractThe advancement of single-cell sequencing technology has promoted the generation of a large amount of single-cell transcriptional profiles, providing unprecedented opportunities to identify drug-resistant cell subpopulations within a tumor. However, few studies have focused on drug response prediction at single-cell level, and their performance remains suboptimal. This paper proposed scAdaDrug, a novel multi-source domain adaptation model powered by adaptive importance-aware representation learning to predict drug response of individual cells. We used a shared encoder to extract domain-invariant features related to drug response from multiple source domains by utilizing adversarial domain adaptation. Particularly, we introduced a plug-and-play module to generate importance-aware and mutually independent weights, which could adaptively modulate the latent representation of each sample in element-wise manner between source and target domains. Extensive experimental results showed that our model achieved state-of-the-art performance in predicting drug response on multiple independent datasets, including single-cell datasets derived from both cell lines and patient-derived xenografts (PDX) models, as well as clinical tumor patient cohorts. Moreover, the ablation experiments demonstrated our model effectively captured the underlying patterns determining drug response from multiple source domains. Hui Liu 0026, Judong Luo |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | TVPR: Text-to-Video Person Retrieval and a New BenchmarkabstractMost existing methods for text-based person retrieval focus on text-to-image person retrieval. Nevertheless, due to the lack of dynamic information provided by isolated frames, the performance is hampered when the person is obscured or variable motion details are missed in isolated frames. To overcome this, we propose a novel Text-to-Video Person Retrieval (TVPR) task. Since there is no dataset or benchmark that describes person videos with natural language, we construct a large-scale cross-modal person video dataset containing detailed natural language annotations, termed as Text-to-Video Person Reidentification (TVPReid) dataset. In this paper, we introduce a Multielement Feature Guided Fragments Learning (MFGF) strategy, which leverages the cross-modal text-video representations to provide strong text-visual and text-motion matching information to tackle uncertain occlusion conflicting and variable motion details. Specifically, we establish two potential cross-modal spaces for text and video feature collaborative learning to progressively reduce the semantic difference between text and video. To evaluate the effectiveness of the proposed MFGF, extensive experiments have been conducted on TVPReid dataset. To the best of our knowledge, MFGF is the first successful attempt to use video for text-based person retrieval task and has achieved state-of-the-art performance on TVPReid dataset. The TVPReid dataset will be publicly available to benefit future research. Xu Zhang 0075, Fan Ni, Guannan Dong, Aichun Zhu, Mingcheng Ni, Hui Liu 0026 |
ACM Multimedia | 7 |
| 2024 | Predicting single-cell cellular responses to perturbations using cycle consistency learningabstractSUMMARY: Phenotype-based drug screening emerges as a powerful approach for identifying compounds that actively interact with cells. Transcriptional and proteomic profiling of cell lines and individual cells provide insights into the cellular state alterations that occur at the molecular level in response to external perturbations, such as drugs or genetic manipulations. In this paper, we propose cycleCDR, a novel deep learning framework to predict cellular response to external perturbations. We leverage the autoencoder to map the unperturbed cellular states to a latent space, in which we postulate the effects of drug perturbations on cellular states follow a linear additive model. Next, we introduce the cycle consistency constraints to ensure that unperturbed cellular state subjected to drug perturbation in the latent space would produces the perturbed cellular state through the decoder. Conversely, removal of perturbations from the perturbed cellular states can restore the unperturbed cellular state. The cycle consistency constraints and linear modeling in the latent space enable to learn transferable representations of external perturbations, so that our model can generalize well to unseen drugs during training stage. We validate our model on four different types of datasets, including bulk transcriptional responses, bulk proteomic responses, and single-cell transcriptional responses to drug/gene perturbations. The experimental results demonstrate that our model consistently outperforms existing state-of-the-art methods, indicating our method is highly versatile and applicable to a wide range of scenarios. AVAILABILITY AND IMPLEMENTATION: The source code is available at: https://github.com/hliulab/cycleCDR. Hui Liu 0026 |
Bioinform. | 2 |
| 2024 | Cross-Domain Feature Disentanglement for Interpretable Modeling of Tumor Microenvironment Impact on Drug ResponseabstractHigh-throughput screening technology has enabled the generation of large-scale drug responses across hundreds of cancer cell lines. There remains a significant gap between in vitro cell lines and actual tumors in vivo in terms of their response to drug treatments yet. This is because tumors consist of a complex cellular composition and histopathology structure, known as the tumor microenvironment (TME), which greatly impacts the drug cytotoxicity against tumor cells. To date, no study has focused on modeling the impact of the TME on clinical drug response. In this study, we postulated that the intricate complexity of an actual tumor can be conceptually simplified into two separable components: cancerous cells and the tumor microenvironment. This assumption allowed us to model the influence of these two constituent parts on drug response through feature disentanglement. We employed a domain adaptation network to decouple and extract features from tumor transcriptional profiles. Specifically, two denoising autoencoders were separately used to extract features from cell lines (source domain) and tumors (target domain) for partial domain alignment and feature decoupling. The private encoder was enforced to extract information only about the TME. Moreover, to ensure generalizability to novel drugs, we employed a graph attention network to learn the latent representation of drugs, enabling us to linearly model the drug perturbation on cellular state in latent space. We validated our model on a benchmark dataset and demonstrated its superior performance in predicting clinical drug response and dissecting the influence of the TME on drug efficacy. Hui Liu 0026 |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | Contrastive learning-based computational histopathology predict differential expression of cancer driver genesabstractMOTIVATION: Digital pathological analysis is run as the main examination used for cancer diagnosis. Recently, deep learning-driven feature extraction from pathology images is able to detect genetic variations and tumor environment, but few studies focus on differential gene expression in tumor cells. RESULTS: In this paper, we propose a self-supervised contrastive learning framework, HistCode, to infer differential gene expression from whole slide images (WSIs). We leveraged contrastive learning on large-scale unannotated WSIs to derive slide-level histopathological features in latent space, and then transfer it to tumor diagnosis and prediction of differentially expressed cancer driver genes. Our experiments showed that our method outperformed other state-of-the-art models in tumor diagnosis tasks, and also effectively predicted differential gene expression. Interestingly, we found the genes with higher fold change can be more precisely predicted. To intuitively illustrate the ability to extract informative features from pathological images, we spatially visualized the WSIs colored by the attention scores of image tiles. We found that the tumor and necrosis areas were highly consistent with the annotations of experienced pathologists. Moreover, the spatial heatmap generated by lymphocyte-specific gene expression patterns was also consistent with the manually labeled WSIs. Gongming Zhou, Lei Deng 0002, Dachuan Zhang, Hui Liu 0026 |
Briefings Bioinform. | 7 |
| 2022 | Attention-wise masked graph contrastive learning for predicting molecular propertyabstractMOTIVATION: Accurate and efficient prediction of the molecular property is one of the fundamental problems in drug research and development. Recent advancements in representation learning have been shown to greatly improve the performance of molecular property prediction. However, due to limited labeled data, supervised learning-based molecular representation algorithms can only search limited chemical space and suffer from poor generalizability. RESULTS: In this work, we proposed a self-supervised learning method, ATMOL, for molecular representation learning and properties prediction. We developed a novel molecular graph augmentation strategy, referred to as attention-wise graph masking, to generate challenging positive samples for contrastive learning. We adopted the graph attention network as the molecular graph encoder, and leveraged the learned attention weights as masking guidance to generate molecular augmentation graphs. By minimization of the contrastive loss between original graph and augmented graph, our model can capture important molecular structure and higher order semantic information. Extensive experiments showed that our attention-wise graph mask contrastive learning exhibited state-of-the-art performance in a couple of downstream molecular property prediction tasks. We also verified that our model pretrained on larger scale of unlabeled data improved the generalization of learned molecular representation. Moreover, visualization of the attention heatmaps showed meaningful patterns indicative of atoms and atomic groups important to specific molecular property. Hui Liu 0026, Yibiao Huang, Lei Deng 0002 |
Briefings Bioinform. | 1 |
| 2022 | DeepDDS: deep graph neural network with attention mechanism to predict synergistic drug combinationsabstractMOTIVATION: Drug combination therapy has become an increasingly promising method in the treatment of cancer. However, the number of possible drug combinations is so huge that it is hard to screen synergistic drug combinations through wet-lab experiments. Therefore, computational screening has become an important way to prioritize drug combinations. Graph neural network has recently shown remarkable performance in the prediction of compound-protein interactions, but it has not been applied to the screening of drug combinations. RESULTS: In this paper, we proposed a deep learning model based on graph neural network and attention mechanism to identify drug combinations that can effectively inhibit the viability of specific cancer cells. The feature embeddings of drug molecule structure and gene expression profiles were taken as input to multilayer feedforward neural network to identify the synergistic drug combinations. We compared DeepDDS (Deep Learning for Drug-Drug Synergy prediction) with classical machine learning methods and other deep learning-based methods on benchmark data set, and the leave-one-out experimental results showed that DeepDDS achieved better performance than competitive methods. Also, on an independent test set released by well-known pharmaceutical enterprise AstraZeneca, DeepDDS was superior to competitive methods by more than 16% predictive precision. Furthermore, we explored the interpretability of the graph attention network and found the correlation matrix of atomic features revealed important chemical substructures of drugs. We believed that DeepDDS is an effective tool that prioritized synergistic drug combinations for further wet-lab experiment validation. AVAILABILITY AND IMPLEMENTATION: Source code and data are available at https://github.com/Sinwang404/DeepDDS/tree/master. Jinxian Wang, Lei Deng 0002, Hui Liu 0026 |
Briefings Bioinform. | 5 |
| 2022 | Graph2MDA: a multi-modal variational graph embedding model for predicting microbe-drug associationsabstractMOTIVATION: Accumulated clinical studies show that microbes living in humans interact closely with human hosts, and get involved in modulating drug efficacy and drug toxicity. Microbes have become novel targets for the development of antibacterial agents. Therefore, screening of microbe-drug associations can benefit greatly drug research and development. With the increase of microbial genomic and pharmacological datasets, we are greatly motivated to develop an effective computational method to identify new microbe-drug associations. RESULTS: In this article, we proposed a novel method, Graph2MDA, to predict microbe-drug associations by using variational graph autoencoder (VGAE). We constructed multi-modal attributed graphs based on multiple features of microbes and drugs, such as molecular structures, microbe genetic sequences and function annotations. Taking as input the multi-modal attribute graphs, VGAE was trained to learn the informative and interpretable latent representations of each node and the whole graph, and then a deep neural network classifier was used to predict microbe-drug associations. The hyperparameter analysis and model ablation studies showed the sensitivity and robustness of our model. We evaluated our method on three independent datasets and the experimental results showed that our proposed method outperformed six existing state-of-the-art methods. We also explored the meaning of the learned latent representations of drugs and found that the drugs show obvious clustering patterns that are significantly consistent with drug ATC classification. Moreover, we conducted case studies on two microbes and two drugs and found 75-95% predicted associations have been reported in PubMed literature. Our extensive performance evaluations validated the effectiveness of our proposed method. AVAILABILITY AND IMPLEMENTATION: Source codes and preprocessed data are available at https://github.com/moen-hyb/Graph2MDA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lei Deng 0002, Yibiao Huang, Hui Liu 0026 |
Bioinform. | 4 |
| 2022 | MSPCD: predicting circRNA-disease associations via integrating multi-source data and hierarchical neural networkabstractBACKGROUND: Increasing evidence shows that circRNA plays an essential regulatory role in diseases through interactions with disease-related miRNAs. Identifying circRNA-disease associations is of great significance to precise diagnosis and treatment of diseases. However, the traditional biological experiment is usually time-consuming and expensive. Hence, it is necessary to develop a computational framework to infer unknown associations between circRNA and disease. RESULTS: In this work, we propose an efficient framework called MSPCD to infer unknown circRNA-disease associations. To obtain circRNA similarity and disease similarity accurately, MSPCD first integrates more biological information such as circRNA-miRNA associations, circRNA-gene ontology associations, then extracts circRNA and disease high-order features by the neural network. Finally, MSPCD employs DNN to predict unknown circRNA-disease associations. CONCLUSIONS: Experiment results show that MSPCD achieves a significantly more accurate performance compared with previous state-of-the-art methods on the circFunBase dataset. The case study also demonstrates that MSPCD is a promising tool that can effectively infer unknown circRNA-disease associations. Lei Deng 0002, Dayun Liu, Yizhan Li, Runqi Wang, Hui Liu 0026 |
BMC Bioinform. | 7 |
| 2021 | A multi-task graph convolutional network modeling of drug-drug interactions and synergistic efficacyabstractIdentification of drug-drug interaction(DDI) is critical for safer and more effective drug co-prescription. As wetlab screening assays are time-consuming, labor-intensive and expensive, it is highly desired to develop an effective computational method to predict drug-drug interactions. In this work, we aim to predict of drug-drug interactions and synergistic drug combinations by proposing an end-to-end multi-task learning framework based on graph convolutional network (GCN). Precisely, we first convert the drug into a molecular graph, in which vertices represent atoms and edges represent chemical bonds. Next, the r-radius subgraph method is applied to molecular graph so that a series of subgraphs are produced for each drug. Next, the subgraphs are used as the input of the graph convolutional network to learn the embedding vector. Finally, the pairwise drug embeddings learned by GCN are concatenated as input into a fully-connected layer for predicting drug-drug interaction and synergistic effects. We conducted extensive performance evaluations on different data sets, including benchmark DDI data sets and manually collected drug combination data sets, and the results show that our proposed method is significantly better than the newly proposed methods (DeepCCI) and four typical machine learning methods (FFNN, SVM, RF, AdaBoost). In addition, our case study showed that 11 our of top 20 predicted DDIs have been reported by PubMed literature. Yuanyuan Deng, Lei Deng 0002, Hui Liu 0026 |
BIBM | 4 |
| 2020 | Predicting circRNA-disease associations using meta path-based representation learning on heterogenous networkabstractCircular RNA (circRNA) is a new class of regulatory non-coding RNAs modulating gene expression by acting as a microRNA (miRNA) sponge, RNA binding protein sponge and translational regulator. A increasing number of experimental studies have shown that circRNA plays an important role in the development of diseases, and circRNA biomarkers are helpful for the diagnosis and treatment of various human diseases. There is a pressing demand to establish an effective computational method to identify the associations between circRNAs and diseases. In this paper, we propose a new computational framework for the prediction of the circRNA-disease associations. In particular, we calculated meta path-based feature vectors for each circRNAdisease pair on a heterogeneous information network (HIN) that integrated multiple subnetworks, including circRNA similarity network, disease similarity network, protein similarity network, circRNA-disease associations, circRNA-protein associations and protein-disease associations. A positive-unlabeled learning algorithm was adopted to generate negative samples, and a random forest classifier was trained to predicted circRNA-disease associations. We conducted performance comparison with three popular methods on the CircR2Disease dataset. The experimental results show that our method outperform other existing methods by achieving AUC 0.983 on 5-fold cross-validation. Lei Deng 0002, Hui Liu 0026 |
BIBM | 3 |
| 2019 | A deep neural network approach using distributed representations of RNA sequence and structure for identifying binding site of RNA-binding proteinsabstractRNA-binding proteins (RBPs) play a crucial role in the post-transcriptional regulation of RNAs. Identification of RBP binding sites is a key step to understand the biological mechanism of post-transcriptional regulation. Although many computational methods have been developed for predicting RNA-protein binding sites, few study considers the k-mer embedding representation of RNA primary sequence and secondary structure specificities. In this paper, we develop a general deep learning framework, named deepRKE, to predict RNA-protein binding sites. deepRKE takes an unsupervised shallow two-layer neural network to automatically learn the distributed representation of k-mers by taking their neighbor context into account. Compared to conventional k-mers approach, distributed representations effectively detect the latent relationship and similarity between k-mers. The distributed representations of the sequences and secondary structures are fed into CNN convolutional neural network (CNN) and a bidirectional long short term memory network (BLSTM) to discriminate the RBP binding sites from unbound sites. We comprehensively evaluate deepRKE on two large-scale RBP binding sites datasets, and the experimental results show that deepRKE achieves better performance than five competitive methods. Lei Deng 0002, Youzhi Liu, Yechuan Shi, Hui Liu 0026 |
BIBM | 4 |
| 2019 | D2VCB: A Hybrid Deep Neural Network for the Prediction of in-vivo Protein-DNA Binding from Combined DNA SequenceabstractPrediction of in-vivo protein-DNA binding is an important, but challenging task in the broad field of computational biology. Although some methods based on deep learning have succeed in modeling in-vivo protein-DNA binding, they often simply extract the sequence features from the original DNA sequence without consideration of other sequence features, such as their reverse, complementary and reverse complementary sequences. Also, one-hot encoding of DNA sequence is vulnerable to the curse of dimensionality, which leads to unwanted equidistance of pairwise sequences. To address these problems, we propose D2VCB (dna2vec, convolution, bi-LSTM), a novel hybrid deep neural network framework using dna2vec to predict in-vivo protein-DNA binding events. We extract input features from DNA original sequences, reverse sequences, complementary and complementary reverse sequences, and then use dna2vec to compute a distributed representation of k-mer. In our D2VCB model, the convolution layer captures motif features, while the recurrent layer captures long-term dependencies among motif features so as to improve prediction accuracy. Our performance comparison experiments show that D2VCB outperforms significantly other existing methods in terms of multiple performance metrics. Lei Deng 0002, Hui Liu 0026 |
BIBM | 3 |
| 2019 | MADOKA: an ultra-fast approach for large-scale protein structure similarity searchingabstractBACKGROUND: Protein comparative analysis and similarity searches play essential roles in structural bioinformatics. A couple of algorithms for protein structure alignments have been developed in recent years. However, facing the rapid growth of protein structure data, improving overall comparison performance and running efficiency with massive sequences is still challenging. RESULTS: Here, we propose MADOKA, an ultra-fast approach for massive structural neighbor searching using a novel two-phase algorithm. Initially, we apply a fast alignment between pairwise structures. Then, we employ a score to select pairs with more similarity to carry out a more accurate fragment-based residue-level alignment. MADOKA performs about 6-100 times faster than existing methods, including TM-align and SAL, in massive alignments. Moreover, the quality of structural alignment of MADOKA is better than the existing algorithms in terms of TM-score and number of aligned residues. We also develop a web server to search structural neighbors in PDB database (About 360,000 protein chains in total), as well as additional features such as 3D structure alignment visualization. The MADOKA web server is freely available at: http://madoka.denglab.org/ CONCLUSIONS: MADOKA is an efficient approach to search for protein structure similarity. In addition, we provide a parallel implementation of MADOKA which exploits massive power of multi-core CPUs. Lei Deng 0002, Guolun Zhong, Chenzhe Liu, Judong Luo, Hui Liu 0026 |
BMC Bioinform. | 5 |
| 2019 | Predicting effective drug combinations using gradient tree boosting based on features extracted from drug-protein heterogeneous networkabstractBACKGROUND: Although targeted drugs have contributed to impressive advances in the treatment of cancer patients, their clinical benefits on tumor therapies are greatly limited due to intrinsic and acquired resistance of cancer cells against such drugs. Drug combinations synergistically interfere with protein networks to inhibit the activity level of carcinogenic genes more effectively, and therefore play an increasingly important role in the treatment of complex disease. RESULTS: In this paper, we combined the drug similarity network, protein similarity network and known drug-protein associations into a drug-protein heterogenous network. Next, we ran random walk with restart (RWR) on the heterogenous network using the combinatorial drug targets as the initial probability, and obtained the converged probability distribution as the feature vector of each drug combination. Taking these feature vectors as input, we trained a gradient tree boosting (GTB) classifier to predict new drug combinations. We conducted performance evaluation on the widely used drug combination data set derived from the DCDB database. The experimental results show that our method outperforms seven typical classifiers and traditional boosting algorithms. CONCLUSIONS: The heterogeneous network-derived features introduced in our method are more informative and enriching compared to the primary ontology features, which results in better performance. In addition, from the perspective of network pharmacology, our method effectively exploits the topological attributes and interactions of drug targets in the overall biological network, which proves to be a systematic and reliable approach for drug discovery. Hui Liu 0026, Lixia Nie, Xiancheng Ding, Judong Luo, Ling Zou 0002 |
BMC Bioinform. | 1 |
| 2018 | XPredRBR: Accurate and Fast Prediction of RNA-Binding Residues in Proteins Using eXtreme Gradient Boosting
Lei Deng 0002, Zuojin Dong, Hui Liu 0026 |
ISBRA | 3 |
| 2018 | PDRLGB: precise DNA-binding residue prediction using a light gradient boosting machineabstractBACKGROUND: Identifying specific residues for protein-DNA interactions are of considerable importance to better recognize the binding mechanism of protein-DNA complexes. Despite the fact that many computational DNA-binding residue prediction approaches have been developed, there is still significant room for improvement concerning overall performance and availability. RESULTS: Here, we present an efficient approach termed PDRLGB that uses a light gradient boosting machine (LightGBM) to predict binding residues in protein-DNA complexes. Initially, we extract a wide variety of 913 sequence and structure features with a sliding window of 11. Then, we apply the random forest algorithm to sort the features in descending order of importance and obtain the optimal subset of features using incremental feature selection. Based on the selected feature set, we use a light gradient boosting machine to build the prediction model for DNA-binding residues. Our PDRLGB method shows better overall predictive accuracy and relatively less training time than other widely used machine learning (ML) methods such as random forest (RF), Adaboost and support vector machine (SVM). We further compare PDRLGB with various existing approaches on the independent test datasets and show improvement in results over the existing state-of-the-art approaches. CONCLUSIONS: PDRLGB is an efficient approach to predict specific residues for protein-DNA interactions. Lei Deng 0002, Juan Pan, Wenyi Yang, Chuyao Liu, Hui Liu 0026 |
BMC Bioinform. | 6 |
| 2018 | Effectively Identifying Compound-Protein Interactions by Learning from Positive and Unlabeled ExamplesabstractPrediction of compound-protein interactions (CPIs) is to find new compound-protein pairs where a protein is targeted by at least a compound, which is a crucial step in new drug design. Currently, a number of machine learning based methods have been developed to predict new CPIs in the literature. However, as there is not yet any publicly available set of validated negative CPIs, most existing machine learning based approaches use the unknown interactions (not validated CPIs) selected randomly as the negative examples to train classifiers for predicting new CPIs. Obviously, this is not quite reasonable and unavoidably impacts the CPI prediction performance. In this paper, we simply take the unknown CPIs as unlabeled examples, and propose a new method called PUCPI (the abbreviation of PU learning for Compound-Protein Interaction identification) that employs biased-SVM (Support Vector Machine) to predict CPIs using only positive and unlabeled examples. PU learning is a class of learning methods that leans from positive and unlabeled (PU) samples. To the best of our knowledge, this is the first work that identifies CPIs using only positive and unlabeled examples. We first collect known CPIs as positive examples and then randomly select compound-protein pairs not in the positive set as unlabeled examples. For each CPI/compound-protein pair, we extract protein domains as protein features and compound substructures as chemical features, then take the tensor product of the corresponding compound features and protein features as the feature vector of the CPI/compound-protein pair. After that, biased-SVM is employed to train classifiers on different datasets of CPIs and compound-protein pairs. Experiments over various datasets show that our method outperforms six typical classifiers, including random forest, L1- and L2-regularized logistic regression, naive Bayes, SVM and k-nearest neighbor (kNN), and three types of existing CPI prediction models. More information can be found at http://admis.fudan.edu.cn/projects/pucpi.html. Zhanzhan Cheng, Shuigeng Zhou, Yang Wang 0100, Hui Liu 0026, Jihong Guan, Yi-Ping Phoebe Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2017 | A new hybrid method for learning bayesian networks: Separation and reunion
Hui Liu 0026, Shuigeng Zhou, Wai Lam, Jihong Guan |
Knowl. Based Syst. | 1 |
| 2016 | Inferring new indications for approved drugs via random walk on drug-disease heterogenous networksabstractBACKGROUND: Since traditional drug research and development is often time-consuming and high-risk, there is an increasing interest in establishing new medical indications for approved drugs, referred to as drug repositioning, which provides a relatively low-cost and high-efficiency approach for drug discovery. With the explosive growth of large-scale biochemical and phenotypic data, drug repositioning holds great potential for precision medicine in the post-genomic era. It is urgent to develop rational and systematic approaches to predict new indications for approved drugs on a large scale. RESULTS: In this paper, we propose the two-pass random walks with restart on a heterogenous network, TP-NRWRH for short, to predict new indications for approved drugs. Rather than random walk on bipartite network, we integrated the drug-drug similarity network, disease-disease similarity network and known drug-disease association network into one heterogenous network, on which the two-pass random walks with restart is implemented. We have conducted performance evaluation on two datasets of drug-disease associations, and the results show that our method has higher performance than six existing methods. A case study on the Alzheimer's disease showed that nine of top 10 predicted drugs have been approved or investigational for neurodegenerative diseases. The experimental results show that our method achieves state-of-the-art performance in predicting new indications for approved drugs. CONCLUSIONS: We proposed a two-pass random walk with restart on the drug-disease heterogeneous network, referred to as TP-NRWRH, to predict new indications for approved drugs. Performance evaluation on two independent datasets showed that TP-NRWRH achieved higher performance than six existing methods on 10-fold cross validations. The case study on the Alzheimer's disease showed that nine of top 10 predicted drugs have been approved or are investigational for neurodegenerative diseases. The results show that our method achieves state-of-the-art performance in predicting new indications for approved drugs. Hui Liu 0026, Yinglong Song, Jihong Guan, Libo Luo, Ziheng Zhuang |
BMC Bioinform. | 1 |
| 2015 | Improving compound-protein interaction prediction by building up highly credible negative samplesabstractMOTIVATION: Computational prediction of compound-protein interactions (CPIs) is of great importance for drug design and development, as genome-scale experimental validation of CPIs is not only time-consuming but also prohibitively expensive. With the availability of an increasing number of validated interactions, the performance of computational prediction approaches is severely impended by the lack of reliable negative CPI samples. A systematic method of screening reliable negative sample becomes critical to improving the performance of in silico prediction methods. RESULTS: This article aims at building up a set of highly credible negative samples of CPIs via an in silico screening method. As most existing computational models assume that similar compounds are likely to interact with similar target proteins and achieve remarkable performance, it is rational to identify potential negative samples based on the converse negative proposition that the proteins dissimilar to every known/predicted target of a compound are not much likely to be targeted by the compound and vice versa. We integrated various resources, including chemical structures, chemical expression profiles and side effects of compounds, amino acid sequences, protein-protein interaction network and functional annotations of proteins, into a systematic screening framework. We first tested the screened negative samples on six classical classifiers, and all these classifiers achieved remarkably higher performance on our negative samples than on randomly generated negative samples for both human and Caenorhabditis elegans. We then verified the negative samples on three existing prediction models, including bipartite local model, Gaussian kernel profile and Bayesian matrix factorization, and found that the performances of these models are also significantly improved on the screened negative samples. Moreover, we validated the screened negative samples on a drug bioactivity dataset. Finally, we derived two sets of new interactions by training an support vector machine classifier on the positive interactions annotated in DrugBank and our screened negative interactions. The screened negative samples and the predicted interactions provide the research community with a useful resource for identifying new drug targets and a helpful supplement to the current curated compound-protein databases. AVAILABILITY: Supplementary files are available at: http://admis.fudan.edu.cn/negative-cpi/. Hui Liu 0026, Jianjiang Sun, Jihong Guan, Jie Zheng 0002, Shuigeng Zhou |
Bioinform. | 1 |
| 2014 | A comparative evaluation on prediction methods of nucleosome positioningabstractNucleosome positioning plays an essential role in cellular processes by modulating accessibility of DNA to proteins. Many computational models have been developed to predict genome-wide nucleosome positions from DNA sequences. Comparative analysis of predicted and experimental nucleosome positioning maps facilitates understanding the regulatory mechanisms of transcription and DNA replication. Therefore, a comprehensive evaluation of existing computational methods is important and useful for biologists to choose appropriate ones in their research. In this article, we carried out a performance comparison among eight widely used computational methods on four species including yeast, fruitfly, mouse and human. In particular, we compared these methods on different regions of each species such as gene sequences, promoters and 5'UTR exons. The experimental results show that the performances of the two latest versions of the thermodynamic model are relatively steadier than the other four methods. Moreover, these methods are workable on four species, but their performances decrease gradually from yeast to human, indicating that the fundamental mechanism of nucleosome positioning is conserved through the evolution process, but more and more factors participate in the determination of nucleosome positions, which leads to sophisticated regulation mechanisms. Hui Liu 0026, Ruichang Zhang, Jihong Guan, Ziheng Zhuang, Shuigeng Zhou |
Briefings Bioinform. | 1 |
| 2013 | Protein function prediction by collective classification with explicit and implicit edges in protein-protein interaction networksabstractBACKGROUND: Protein function prediction is an important problem in the post-genomic era. Recent advances in experimental biology have enabled the production of vast amounts of protein-protein interaction (PPI) data. Thus, using PPI data to functionally annotate proteins has been extensively studied. However, most existing network-based approaches do not work well when annotation and interaction information is inadequate in the networks. RESULTS: In this paper, we proposed a new method that combines PPI information and protein sequence information to boost the prediction performance based on collective classification. Our method divides function prediction into two phases: First, the original PPI network is enriched by adding a number of edges that are inferred from protein sequence information. We call the added edges implicit edges, and the existing ones explicit edges correspondingly. Second, a collective classification algorithm is employed on the new network to predict protein function. CONCLUSIONS: We conducted extensive experiments on two real, publicly available PPI datasets. Compared to four existing protein function prediction approaches, our method performs better in many situations, which shows that adding implicit edges can indeed improve the prediction performance. Furthermore, the experimental results also indicate that our method is significantly better than the compared approaches in sparsely-labeled networks, and it is robust to the change of the proportion of annotated proteins. Hui Liu 0026, Jihong Guan, Shuigeng Zhou |
BMC Bioinform. | 2 |
| 2013 | Identifying Mammalian MicroRNA Targets Based on Supervised Distance Metric LearningabstractMicroRNAs (miRNAs) have been emerged as a novel class of endogenous post-transcriptional regulators in a variety of animal and plant species. One challenge facing miRNA research is to accurately identify the target mRNAs, because of the very limited sequence complementarity between miRNAs and their target sites, and the scarcity of experimentally validated targets to guide accurate prediction. In this paper, we propose a new method called SuperMirTar that exploits supervised distance learning to predict miRNA targets. Specifically, we use the experimentally supported miRNA-mRNA pairs as training set to learn a distance metric function that minimizes the distances between miRNAs and mRNAs with validated interactions, then use the learned function to calculate the distances of test miRNAmRNA interactions, and those with smaller distances than a predefined threshold are regarded as true interactions. We carry out performance comparison between the proposed approach and seven existing methods on independent datasets, the results show that our method achieves superior performance and can effectively narrow the gap between the number of predicted miRNA targets and the number of experimentally validated ones. Hui Liu 0026, Shuigeng Zhou, Jihong Guan |
IEEE J. Biomed. Health Informatics | 1 |