Hui Lu 0004

dblp:65/4062-4 · DBLP profile ↗
← Back
24ranked-venue papers
0as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 20 · 6 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021
YearPublicationVenuePosition
2026 A causal bidirectional selective state space model for imaging genetics in neurodegenerative diseases
Hongrui Liu 0001, Yuanyuan Gui, Binglei Zhao 0001, Hui Lu 0004, Manhua Liu
Neural Networks4
2025 A novel prognostic framework for HBV-infected hepatocellular carcinoma: insights from ferroptosis and iron metabolism proteomics
abstract
Effective classification methods and prognostic models enable more accurate classification and treatment of hepatocellular carcinoma (HCC) patients. However, the weak correlation between RNA and protein data has limited the clinical utility of previous RNA-based prognostic models for HCC. In this work, we constructed a novel prognostic framework for HCC patients using seven differentially expressed proteins associated with ferroptosis and iron metabolism. Furthermore, this prognostic model robustly classifies HCC patients into three clinically relevant risk groups. Significant differences in overall survival, age, tumor differentiation, microvascular invasion, distant metastasis, and alpha-fetoprotein levels were observed among the risk groups. Based on the prognostic model and known biological pathways, we explored the potential mechanisms underlying the inconsistent differential expression patterns of FTH1 (Ferritin heavy chain 1) mRNA and protein. Our findings demonstrated that tumor tissues in HCC patients promote liver cancer progression by downregulating FTH1 protein expression, rather than upregulating FTH1 mRNA expression, ultimately leading to poor prognosis. Subsequently, based on risk score and tumor size, we developed a nomogram for predicting the prognosis of HCC patients, which demonstrated superior predictive performance in both the training and validation cohorts (C-index: 0.774; AUC for 1-5 years: 0.783-0.964). Additionally, our findings demonstrated that the adverse prognosis of high-risk HCC patients was closely correlated with ferroptosis in liver cancer tissues, alterations in iron metabolism, and changes in the tumor immune microenvironment. In conclusion, our prognostic model and predictive nomogram offer novel insights and tools for the effective classification of HCC patients, potentially enhancing clinical decision-making and outcomes.
Yongyong Ren, Xinbo Wang, Yuening Zhang, Yingqi Hua, Hongyu Zhao 0003, Hui Lu 0004
Briefings Bioinform.7
2025 scGO: interpretable deep neural network for cell status annotation and disease diagnosis
abstract
Machine learning has emerged as a transformative tool for elucidating cellular heterogeneity in single-cell RNA sequencing. However, a significant challenge lies in the "black box" nature of deep learning models, which obscures the decision-making process and limits interpretability in cell status annotation. In this study, we introduced scGO, a Gene Ontology (GO)-inspired deep learning framework designed to provide interpretable cell status annotation for scRNA-seq data. scGO employs sparse neural networks to leverage the intrinsic biological relationships among genes, transcription factors, and GO terms, significantly augmenting interpretability and reducing computational cost. scGO outperforms state-of-the-art methods in the precise characterization of cell subtypes across diverse datasets. Our extensive experimentation across a spectrum of scRNA-seq datasets underscored the remarkable efficacy of scGO in disease diagnosis, prediction of developmental stages, and evaluation of disease severity and cellular senescence status. Furthermore, we incorporated in silico individual gene manipulations into the scGO model, introducing an additional layer for discovering therapeutic targets. Our results provide an interpretable model for accurately annotating cell status, capturing latent biological knowledge, and informing clinical practice.
Yingnan Hou, Hui Lu 0004
Briefings Bioinform.6
2024 Transformer with convolution and graph-node co-embedding: An accurate and interpretable vision backbone for predicting gene expressions from local histopathological image
abstract
Inferring gene expressions from histopathological images has long been a fascinating yet challenging task, primarily due to the substantial disparities between the two modality. Existing strategies using local or global features of histological images are suffering model complexity, GPU consumption, low interpretability, insufficient encoding of local features, and over-smooth prediction of gene expressions among neighboring sites. In this paper, we develop TCGN (Transformer with Convolution and Graph-Node co-embedding method) for gene expression estimation from H&E-stained pathological slide images. TCGN comprises a combination of convolutional layers, transformer encoders, and graph neural networks, and is the first to integrate these blocks in a general and interpretable computer vision backbone. Notably, TCGN uniquely operates with just a single spot image as input for histopathological image analysis, simplifying the process while maintaining interpretability. We validate TCGN on three publicly available spatial transcriptomic datasets. TCGN consistently exhibited the best performance (with median PCC 0.232). TCGN offers superior accuracy while keeping parameters to a minimum (just 86.241 million), and it consumes minimal memory, allowing it to run smoothly even on personal computers. Moreover, TCGN can be extended to handle bulk RNA-seq data while providing the interpretability. Enhancing the accuracy of omics information prediction from pathological images not only establishes a connection between genotype and phenotype, enabling the prediction of costly-to-measure biomarkers from affordable histopathological images, but also lays the groundwork for future multi-modal data modeling. Our results confirm that TCGN is a powerful tool for inferring gene expressions from histopathological images in precision health applications.
Yan Kong, Ronghan Li, Zuoheng Wang, Hui Lu 0004
Medical Image Anal.5
2024 A sparse transformer generation network for brain imaging genetic association
Hongrui Liu 0001, Yuanyuan Gui, Hui Lu 0004, Manhua Liu
Pattern Recognit.3
2023 Superresolved spatial transcriptomics transferred from a histological context
Xiaocheng Zhou, Yan Kong, Hui Lu 0004
Appl. Intell.4
2023 HBV-infected hepatocellular carcinoma can be robustly classified into three clinically relevant subgroups by a novel analytical protocol
abstract
Liver cancer is the third leading cause of cancer-related death worldwide, and hepatocellular carcinoma (HCC) accounts for a relatively large proportion of all primary liver malignancies. Among the several known risk factors, hepatitis B virus (HBV) infection is one of the important causes of HCC. In this study, we demonstrated that the HBV-infected HCC patients could be robustly classified into three clinically relevant subgroups, i.e. Cluster1, Cluster2 and Cluster3, based on consistent differentially expressed mRNAs and proteins, which showed better generalization. The proposed three subgroups showed different molecular characteristics, immune microenvironment and prognostic survival characteristics. The Cluster1 subgroup had near-normal levels of metabolism-related proteins, low proliferation activity and good immune infiltration, which were associated with its good liver function, smaller tumor size, good prognosis, low alpha-fetoprotein (AFP) levels and lower clinical stage. In contrast, the Cluster3 subgroup had the lowest levels of metabolism-related proteins, which corresponded with its severe liver dysfunction. Also, high proliferation activity and poor immune microenvironment in Cluster3 subgroup were associated with its poor prognosis, larger tumor size, high AFP levels, high incidence of tumor thrombus and higher clinical stage. The characteristics of the Cluster2 subgroup were between the Cluster1 and Cluster3 groups. In addition, MCM2-7, RFC2-5, MSH2, MSH6, SMC2, SMC4, NCPAG and TOP2A proteins were significantly upregulated in the Cluster3 subgroup. Meanwhile, abnormally high phosphorylation levels of these proteins were associated with high levels of DNA repair, telomere maintenance and proliferative features. Therefore, these proteins could be identified as potential diagnostic and prognostic markers. In general, our research provided a novel analytical protocol and insights for the robust classification, treatment and prevention of HBV-infected HCC.
Leijie Li, Yuening Zhang, Yongyong Ren, Jianlei Gu, Xinbo Wang, Hongyu Zhao 0003, Hui Lu 0004
Briefings Bioinform.8
2023 diseaseGPS: auxiliary diagnostic system for genetic disorders based on genotype and phenotype
abstract
SUMMARY: The next-generation sequencing brought opportunities for the diagnosis of genetic disorders due to its high-throughput capabilities. However, the majority of existing methods were limited to only sequencing candidate variants, and the process of linking these variants to a diagnosis of genetic disorders still required medical professionals to consult databases. Therefore, we introduce diseaseGPS, an integrated platform for the diagnosis of genetic disorders that combines both phenotype and genotype data for analysis. It offers not only a user-friendly GUI web application for those without a programming background but also scripts that can be executed in batch mode for bioinformatics professionals. The genetic and phenotypic data are integrated using the ACMG-Bayes method and a novel phenotypic similarity method, to prioritize the results of genetic disorders. diseaseGPS was evaluated on 6085 cases from Deciphering Developmental Disorders project and 187 cases from Shanghai Children's hospital. The results demonstrated that diseaseGPS performed better than other commonly used methods. AVAILABILITY AND IMPLEMENTATION: diseaseGPS is available to freely accessed at https://diseasegps.sjtu.edu.cn with source code at https://github.com/BioHuangDY/diseaseGPS.
Daoyi Huang, Pin Li, Yongfen Lyu, Jincai Feng, Mingyue Wei, Zhixing Zhu, Jianlei Gu, Yongyong Ren, Guangjun Yu, Hui Lu 0004
Bioinform.13
2022 MZINBVA: variational approximation for multilevel zero-inflated negative-binomial models for association analysis in microbiome surveys
abstract
As our understanding of the microbiome has expanded, so has the recognition of its critical role in human health and disease, thereby emphasizing the importance of testing whether microbes are associated with environmental factors or clinical outcomes. However, many of the fundamental challenges that concern microbiome surveys arise from statistical and experimental design issues, such as the sparse and overdispersed nature of microbiome count data and the complex correlation structure among samples. For example, in the human microbiome project (HMP) dataset, the repeated observations across time points (level 1) are nested within body sites (level 2), which are further nested within subjects (level 3). Therefore, there is a great need for the development of specialized and sophisticated statistical tests. In this paper, we propose multilevel zero-inflated negative-binomial models for association analysis in microbiome surveys. We develop a variational approximation method for maximum likelihood estimation and inference. It uses optimization, rather than sampling, to approximate the log-likelihood and compute parameter estimates, provides a robust estimate of the covariance of parameter estimates and constructs a Wald-type test statistic for association testing. We evaluate and demonstrate the performance of our method using extensive simulation studies and an application to the HMP dataset. We have developed an R package MZINBVA to implement the proposed method, which is available from the GitHub repository https://github.com/liudoubletian/MZINBVA.
Peirong Xu, Yueyao Du, Hui Lu 0004, Hongyu Zhao 0003, Tao Wang 0067
Briefings Bioinform.4
2020 Predicting viral exposure response from modeling the changes of co-expression networks using time series gene expression data
abstract
BACKGROUND: Deciphering the relationship between clinical responses and gene expression profiles may shed light on the mechanisms underlying diseases. Most existing literature has focused on exploring such relationship from cross-sectional gene expression data. It is likely that the dynamic nature of time-series gene expression data is more informative in predicting clinical response and revealing the physiological process of disease development. However, it remains challenging to extract useful dynamic information from time-series gene expression data. RESULTS: We propose a statistical framework built on considering co-expression network changes across time from time series gene expression data. It first detects change point for co-expression networks and then employs a Bayesian multiple kernel learning method to predict exposure response. There are two main novelties in our method: the use of change point detection to characterize the co-expression network dynamics, and the use of kernel function to measure the similarity between subjects. Our algorithm allows exposure response prediction using dynamic network information across a collection of informative gene sets. Through parameter estimations, our model has clear biological interpretations. The performance of our method on the simulated data under different scenarios demonstrates that the proposed algorithm has better explanatory power and classification accuracy than commonly used machine learning algorithms. The application of our method to time series gene expression profiles measured in peripheral blood from a group of subjects with respiratory viral exposure shows that our method can predict exposure response at early stage (within 24 h) and the informative gene sets are enriched for pathways related to respiratory and influenza virus infection. CONCLUSIONS: The biological hypothesis in this paper is that the dynamic changes of the biological system are related to the clinical response. Our results suggest that when the relationship between the clinical response and a single gene or a gene set is not significant, we may benefit from studying the relationships among genes in gene sets that may lead to novel biological insights.
Fangli Dong, Tao Wang 0067, Hui Lu 0004, Hongyu Zhao 0003
BMC Bioinform.5
2019 CellSim: a novel software to calculate cell similarity and identify their co-regulation networks
abstract
BACKGROUND: Cell direct reprogramming technology has been rapidly developed with its low risk of tumor risk and avoidance of ethical issues caused by stem cells, but it is still limited to specific cell types. Direct reprogramming from an original cell to target cell type needs the cell similarity and cell specific regulatory network. The position and function of cells in vivo, can provide some hints about the cell similarity. However, it still needs further clarification based on molecular level studies. RESULT: CellSim is therefore developed to offer a solution for cell similarity calculation and a tool of bioinformatics for researchers. CellSim is a novel tool for the similarity calculation of different cells based on cell ontology and molecular networks in over 2000 different human cell types and presents sharing regulation networks of part cells. CellSim can also calculate cell types by entering a list of genes, including more than 250 human normal tissue specific cell types and 130 cancer cell types. The results are shown in both tables and spider charts which can be preserved easily and freely. CONCLUSION: CellSim aims to provide a computational strategy for cell similarity and the identification of distinct cell types. Stable CellSim releases (Windows, Linux, and Mac OS/X) are available at: www.cellsim.nwsuaflmz.com , and source code is available at: https://github.com/lileijie1992/CellSim/ .
Leijie Li, Dongxue Che, Siddiq Ur Rahman, Jianbang Zhao, Jiantao Yu, Shiheng Tao, Hui Lu 0004, Mingzhi Liao
BMC Bioinform.9
2018 A new method to measure the semantic similarity from query phenotypic abnormalities to diseases based on the human phenotype ontology
abstract
BACKGROUND: Although rapid developed sequencing technologies make it possible for genotype data to be used in clinical diagnosis, it is still challenging for clinicians to understand the results of sequencing and make correct judgement based on them. Before this, diagnosis based on clinical features held a leading position. With the establishment of the Human Phenotype Ontology (HPO) and the enrichment of phenotype-disease annotations, there throws much more attention to the improvement of phenotype-based diagnosis. RESULTS: In this study, we presented a novel method called RelativeBestPair to measure similarity from the query terms to hereditary diseases based on HPO and then rank the candidate diseases. To evaluate the performance, we simulated a set of patients based on 44 complex diseases. Besides, by adding noise or imprecision or both, cases closer to real clinical conditions were generated. Thus, four simulated datasets were used to make comparison among RelativeBestPair and seven existing semantic similarity measures. RelativeBestPair ranked the underlying disease as top 1 on 93.73% of the simulated dataset without noise and imprecision, 93.64% of the simulated dataset with noise and without imprecision, 39.82% of the simulated dataset without noise and with imprecision, and 33.64% of the simulated dataset with both noise and imprecision. CONCLUSION: Compared with the seven existing semantic similarity measures, RelativeBestPair showed similar performance in two datasets without imprecision. While RelativeBestPair appeared to be equal to Resnik and better than other six methods in the simulated dataset without noise and with imprecision, it significantly outperformed all other seven methods in the simulated dataset with both noise and imprecision. It can be indicated that RelativeBestPair might be of great help in clinical setting.
Xiaofeng Gong, Zhongqu Duan, Hui Lu 0004
BMC Bioinform.4
2017 Prediction of human QT prolongation liability based on pre-clinical RNA expression profiles
abstract
Marked drug-induced prolongation of the QT interval on the electrocardiogram is associated with Torsades de Pointes (TdP), a potentially life-threatening cardiac arrhythmia. Assessment of QT prolongation liability in the drug development process is required but is time and resource intensive. Current pre-clinical safety assessments use patch clamp analysis of the Human Ether-a-Go-Go (hERG) channel, but analyses have broadened to include patch clamp analysis of other ion channels and the use of in silico models. This investigation describes a method for predicting drug-induced QT prolongation liability in humans based on an association with RNA microarray expression profiles from rat liver data, and machine learning implemented in open-source software. Recently reported hERG patch clamp sensitivities and specificities range from between 64-82% and 75-88% respectively. Classification in this study was done using drugs known to prolong the QT interval vs. those that do not, regardless of the drugs' respective indication(s), and then further sub-classified by indication which resulted in 76 sub-groups. Classifier results in this project using 10-fold cross validation had average sensitivities of 85% and specificities of 90% using all available datasets as input, and a mean sensitivity and specificity of 92% and 94%, respectively across 76 drug sub-classifications. While an association between rat liver RNA expression profiles and QT prolongation in human heart tissue does not imply that a specific genetic expression profile is responsible for the QT prolongation, these results suggest that machine learning of gene expression profiles to predict QT liability may be used as a surrogate biomarker as part of the pre-clinical cardiac safety assessment of drugs.
Dennis M. Bergau, Cong Liu 0020, Hui Lu 0004
BIBM3
2016 A novel scoring estimator to screening for oncogenic chimeric transcripts in cancer transcriptome sequencing
abstract
Based on various genomic information of chimeric transcript, recent studies used machine-learning methods to predict the oncogenic potentials for chimeric transcripts, however these works ignored transcriptional signature of those chimeric transcripts. Based on clonal evolution theory, we hypothesized that a chimeric transcript is more likely to be an oncogenic `driver' mutation, if the neoplastic cells harboring this chimeric mutation has larger clonal size than other neoplastic cells in a particular tumor. Here we proposed a novel method, called iFCR (internal Fusion Clone Ratio), to estimate the ratio of subclone carrying chimeric transcripts to the rest of neoplastic cells in transcriptome sequencing data. To evaluate our hypothesis, we applied iFCR method on two public cancer transcriptome sequencing datasets, one for breast cancer cell line and the other for prostate tumors with adjacent normal tissues. Our results demonstrated that the chimeric transcripts in tumor samples appear to have higher iFCR value than normal tissues, the most frequent prostate cancer fusion mutation, TMPRSS2- ERG, has remarkably higher iFCR value in all three independent patients. Our work providing a novel point of view for screening oncogenesis chimeric transcripts in cancer research.
Jianlei Gu, Shi-Yi Liu, Cong Liu 0020, Hui Lu 0004
BIBM5
2016 Implementation of a city-wide Health Information Exchange solution in the largest metropolitan region in China
abstract
Objective: Health Information Exchange (HIE) enables providers to share healthcare information electronically across different organizations to promote safer, more efficient, and less costly patient-centered care. This paper describes the development and implementation of a city-wide HIE system in Shanghai, China. Methods: In 2006, as a product of the Chinese healthcare reform, the Health Information Exchange and Sharing Platform was proposed as a means to facilitate HIE within Shanghai. In collaboration with the Shanghai Hospital Development Center, a state-run nonprofit corporate, the HIE project was implemented across multiple levels within the city. The HIE system is based on the Service-oriented Architecture and complies with industry standards. Results: On September 2010, the first and largest Chinese HIE system was established. As of 2016 the system includes all of Shanghai's 38 tertiary hospitals (highest level hospitals in China), plus 6 district hospitals, and 40 community health centers, with coverage for 39 million patients. The system currently provides a rich source of patient information including encounter history, medication history, laboratory results, radiology images and reports, and clinical notes. Initial outcomes indicate a significant reduction in medication errors and duplication of tests, saving at least 48 million RMB a year, with overall improvement in the quality of care following implementation. Conclusion: The adoption of HIE in Shanghai resulted in improved access to accurate, complete, and relevant clinical information in real time, thus facilitating delivery of high-quality, cost-effective, and efficient care.
Guang-Jun Yu, Wenbin Cui, Li Zhou 0007, David W. Bates, Jianlei Gu, Hui Lu 0004
BIBM6
2016 Gut microbiota community adaption during young children fecal microbiota transplantation by 16s rDNA sequencing
Jianlei Gu, Yi-Zhong Wang, Shi-Yi Liu, Guang-Jun Yu, Hui Lu 0004
Neurocomputing6
2015 A novel essential domain perspective for exploring gene essentiality
abstract
MOTIVATION: Genes with indispensable functions are identified as essential; however, the traditional gene-level studies of essentiality have several limitations. In this study, we characterized gene essentiality from a new perspective of protein domains, the independent structural or functional units of a polypeptide chain. RESULTS: To identify such essential domains, we have developed an Expectation-Maximization (EM) algorithm-based Essential Domain Prediction (EDP) Model. With simulated datasets, the model provided convergent results given different initial values and offered accurate predictions even with noise. We then applied the EDP model to six microbial species and predicted 1879 domains to be essential in at least one species, ranging 10-23% in each species. The predicted essential domains were more conserved than either non-essential domains or essential genes. Comparing essential domains in prokaryotes and eukaryotes revealed an evolutionary distance consistent with that inferred from ribosomal RNA. When utilizing these essential domains to reproduce the annotation of essential genes, we received accurate results that suggest protein domains are more basic units for the essentiality of genes. Furthermore, we presented several examples to illustrate how the combination of essential and non-essential domains can lead to genes with divergent essentiality. In summary, we have described the first systematic analysis on gene essentiality on the level of domains. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yulan Lu, Jingyuan Deng, Hai Peng, Hui Lu 0004, Long Jason Lu
Bioinform.5
2012 A structure-based protocol for learning the family-specific mechanisms of membrane-binding domains
abstract
MOTIVATION: Peripheral membrane-targeting domain (MTD) families, such as C1-, C2- and PH domains, play a key role in signal transduction and membrane trafficking by dynamically translocating their parent proteins to specific plasma membranes when changes in lipid composition occur. It is, however, difficult to determine the subset of domains within families displaying this property, as sequence motifs signifying the membrane binding properties are not well defined. For this reason, procedures based on sequence similarity alone are often insufficient in computational identification of MTDs within families (yielding less than 65% accuracy even with a sequence identity of 70%). RESULTS: We present a machine learning protocol for determining membrane-targeting properties achieving 85-90% accuracy in separating binding and non-binding domains within families. Our model is based on features from both sequence and structure, thereby incorporation statistics obtained from the entire domain family and domain-specific physical quantities such as surface electrostatics. In addition, by using the enriched rules in alternating decision tree classifiers, we are able to determine the meaning of the assigned function labels in terms of biological mechanisms. CONCLUSIONS: The high accuracy of the learned models and good agreement between the rules discovered using the ADtree classifier and mechanisms reported in the literature reflect the value of machine learning protocols in both prediction and biological knowledge discovery. Our protocol can thus potentially be used as a general function annotation and knowledge mining tool for other protein domains. AVAILABILITY: metador.bioengr.uic.edu CONTACT: [email protected].
Morten Källberg, Nitin Bhardwaj, Robert E. Langlois, Hui Lu 0004
Bioinform.4
2010 Genome-wide sequence-based prediction of peripheral proteins using a novel semi-supervised learning technique
abstract
BACKGROUND: In supervised learning, traditional approaches to building a classifier use two sets of examples with pre-defined classes along with a learning algorithm. The main limitation of this approach is that examples from both classes are required which might be infeasible in certain cases, especially those dealing with biological data. Such is the case for membrane-binding peripheral domains that play important roles in many biological processes, including cell signaling and membrane trafficking by reversibly binding to membranes. For these domains, a well-defined positive set is available with domains known to bind membrane along with a large unlabeled set of domains whose membrane binding affinities have not been measured. The aforementioned limitation can be addressed by a special class of semi-supervised machine learning called positive-unlabeled (PU) learning that uses a positive set with a large unlabeled set. METHODS In this study, we implement the first application of PU-learning to a protein function prediction problem: identification of peripheral domains. PU-learning starts by identifying reliable negative (RN) examples iteratively from the unlabeled set until convergence and builds a classifier using the positive and the final RN set. A data set of 232 positive cases and ~3750 unlabeled ones were used to construct and validate the protocol. RESULTS: Holdout evaluation of the protocol on a left-out positive set showed that the accuracy of prediction reached up to 95% during two independent implementations. CONCLUSION: These results suggest that our protocol can be used for predicting membrane-binding properties of a wide variety of modular domains. Protocols like the one presented here become particularly useful in the case of availability of information from one class only.
Nitin Bhardwaj, Mark Gerstein, Hui Lu 0004
BMC Bioinform.3
2010 An improved machine learning protocol for the identification of correct Sequest search results
abstract
BACKGROUND: Mass spectrometry has become a standard method by which the proteomic profile of cell or tissue samples is characterized. To fully take advantage of tandem mass spectrometry (MS/MS) techniques in large scale protein characterization studies robust and consistent data analysis procedures are crucial. In this work we present a machine learning based protocol for the identification of correct peptide-spectrum matches from Sequest database search results, improving on previously published protocols. RESULTS: The developed model improves on published machine learning classification procedures by 6% as measured by the area under the ROC curve. Further, we show how the developed model can be presented as an interpretable tree of additive rules, thereby effectively removing the 'black-box' notion often associated with machine learning classifiers, allowing for comparison with expert rule-of-thumb. Finally, a method for extending the developed peptide identification protocol to give probabilistic estimates of the presence of a given protein is proposed and tested. CONCLUSIONS: We demonstrate the construction of a high accuracy classification model for Sequest search results from MS/MS spectra obtained by using the MALDI ionization. The developed model performs well in identifying correct peptide-spectrum matches and is easily extendable to the protein identification problem. The relative ease with which additional experimental parameters can be incorporated into the classification framework, to give additional discriminatory power, allows for future tailoring of the model to take advantage of information from specific instrument set-ups.
Morten Källberg, Hui Lu 0004
BMC Bioinform.2
2010 Analysis of Combinatorial Regulation: Scaling of Partnerships between Regulators with the Number of Governed Targets
abstract
Through combinatorial regulation, regulators partner with each other to control common targets and this allows a small number of regulators to govern many targets. One interesting question is that given this combinatorial regulation, how does the number of regulators scale with the number of targets? Here, we address this question by building and analyzing co-regulation (co-transcription and co-phosphorylation) networks that describe partnerships between regulators controlling common genes. We carry out analyses across five diverse species: Escherichia coli to human. These reveal many properties of partnership networks, such as the absence of a classical power-law degree distribution despite the existence of nodes with many partners. We also find that the number of co-regulatory partnerships follows an exponential saturation curve in relation to the number of targets. (For E. coli and Bacillus subtilis, only the beginning linear part of this curve is evident due to arrangement of genes into operons.) To gain intuition into the saturation process, we relate the biological regulation to more commonplace social contexts where a small number of individuals can form an intricate web of connections on the internet. Indeed, we find that the size of partnership networks saturates even as the complexity of their output increases. We also present a variety of models to account for the saturation phenomenon. In particular, we develop a simple analytical model to show how new partnerships are acquired with an increasing number of target genes; with certain assumptions, it reproduces the observed saturation. Then, we build a more general simulation of network growth and find agreement with a wide range of real networks. Finally, we perform various down-sampling calculations on the observed data to illustrate the robustness of our conclusions.
Nitin Bhardwaj, Matthew B. Carson, Alexej Abyzov, Koon-Kiu Yan, Hui Lu 0004, Mark Gerstein
PLoS Comput. Biol.5
2007 MeTaDoR: a comprehensive resource for membrane targeting domains and their host proteins
abstract
MOTIVATION: Protein-lipid interactions play a central role in cellular signaling and membrane trafficking and at the core of these interactions are domains specialized in lipid binding and membrane targeting. Considering the importance of these domains, we have created MeTaDoR, a comprehensive resource dedicated to membrane targeting domains (MTDs). RESULT: MeTaDoR begins with a brief introduction about all the important MTDs including their subcellular localization and structural features. Sequences of all known MTDs are then provided in two formats: standard Prosite format and a parsed tab-delimited format that provides a manually curated classification into binding or non-binding. Structures of all MTDs and host proteins known so far are provided with links to PDB and Pfam databases. Membrane-binding orientation of these proteins, whether experimentally determined or proposed, is also provided with links to the appropriate literature. To facilitate molecular dynamics studies of these proteins, the force-field parameters for many non-standard lipids that commonly interact with these proteins are also provided. Finally, an online server for predicting membrane-binding proteins and a search function with various search fields are included. The resource is publicly available and will be updated on a regular basis.
Nitin Bhardwaj, Robert V. Stahelin, Guijun Zhao, Wonhwa Cho, Hui Lu 0004
Bioinform.5
2007 Finding new structural and sequence attributes to predict possible disease association of single amino acid polymorphism (SAP)
abstract
MOTIVATION: The rapid accumulation of single amino acid polymorphisms (SAPs), also known as non-synonymous single nucleotide polymorphisms (nsSNPs), brings the opportunities and needs to understand and predict their disease association. Currently published attributes are limited, the detailed mechanisms governing the disease association of a SAP remain unclear and thus, further investigation of new attributes and improvement of the prediction are desired. RESULTS: A SAP dataset was compiled from the Swiss-Prot variant pages. We extracted and demonstrated the effectiveness of several new biologically informative attributes including the structural neighbor profiles that describe the SAP's microenvironment, nearby functional sites that measure the structure-based and sequence-based distances between the SAP site and its nearby functional sites, aggregation properties that measure the likelihood of protein aggregation and disordered regions that consider whether the SAP is located in structurally disordered regions. The new attributes provided insights into the mechanisms of the disease association of SAPs. We built a support vector machines (SVMs) classifier employing a carefully selected set of new and previously published attributes. Through a strict protein-level 5-fold cross-validation, we attained an overall accuracy of 82.61%, and an MCC of 0.60. Moreover, a web server was developed to provide a user-friendly interface for biologists. AVAILABILITY: The web server is available at http://sapred.cbi.pku.edu.cn/
Zhi-Qiang Ye, Shuqi Zhao, Xiao-Qiao Liu, Robert E. Langlois, Hui Lu 0004, Liping Wei
Bioinform.6
2005 Correlation between gene expression profiles and protein-protein interactions within and across genomes
abstract
MOTIVATION: Function annotation of an unclassified protein on the basis of its interaction partners is well documented in the literature. Reliable predictions of interactions from other data sources such as gene expression measurements would provide a useful route to function annotation. We investigate the global relationship of protein-protein interactions with gene expression. This relationship is studied in four evolutionarily diverse species, for which substantial information regarding their interactions and expression is available: human, mouse, yeast and Escherichia coli. RESULTS: In E.coli the expression of interacting pairs is highly correlated in comparison to random pairs, while in the other three species, the correlation of expression of interacting pairs is only slightly stronger than that of random pairs. To strengthen the correlation, we developed a protocol to integrate ortholog information into the interaction and expression datasets. In all four genomes, the likelihood of predicting protein interactions from highly correlated expression data is increased using our protocol. In yeast, for example, the likelihood of predicting a true interaction, when the correlation is > 0.9, increases from 1.4 to 9.4. The improvement demonstrates that protein interactions are reflected in gene expression and the correlation between the two is strengthened by evolution information. The results establish that co-expression of interacting protein pairs is more conserved than that of random ones.
Nitin Bhardwaj, Hui Lu 0004
Bioinform.2