VLDB 2026 Research / reviewers in the wild / expert
Ruth Nussinov
dblp:39/5856
· DBLP profile ↗
68ranked-venue papers
14as first author
5since 2021 · last 2025
0000-0002-8115-6415ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 63 · 13 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-authorArtificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mutations in tumor signaling, metastases, and synthetic lethality establish distinct patternsabstractEffective identification of oncogenic mutations is essential for diagnosis, forecasting resistance, and metastasis in remission. It is required for an optimal drug regimen. We develop a framework to discover mutations that co-exist in different oncoproteins, and those that are excluded, likely encoding oncogene-induced senescence. First, mapping the proteins onto pathways assists combinatorial drug selections and helps to detect metastases. Second, it provides the molecular basis for synthetic lethality, to date investigated at the genome level. Our pan-cancer profiles of ~60,000 tumor sequences, detect 3424 co-existing tumor-specific mutations. Mapping them onto pathways indicates that they preferentially promote specific primary tumors. We uncover metastatic mutations and provide metastatic breast-cancer markers. This work not only clarifies the mechanistic basis of intratumor mutational diversity but usefully reveals markers for metastasis in patients' genomes and introduces a novel computational framework for detecting metastasis based on tumor mutational profiles. Mapping the mutations onto pathways provides an invaluable metastasis-targeting resource, guiding drug combinations. Bengi Ruken Yavuz, Ugur Sahin, Hyunbum Jang, Ruth Nussinov, Nurcan Tuncbag |
PLoS Comput. Biol. | 4 |
| 2022 | HMI-PRED 2.0: a biologist-oriented web application for prediction of host-microbe protein-protein interaction by interface mimicryabstractSUMMARY: HMI-PRED 2.0 is a publicly available web service for the prediction of host-microbe protein-protein interaction by interface mimicry that is intended to be used without extensive computational experience. A microbial protein structure is screened against a database covering the entire available structural space of complexes of known human proteins. AVAILABILITY AND IMPLEMENTATION: HMI-PRED 2.0 provides user-friendly graphic interfaces for predicting, visualizing and analyzing host-microbe interactions. HMI-PRED 2.0 is available at https://hmipred.org/. Hansaim Lim, Chung-Jung Tsai, Ozlem Keskin, Ruth Nussinov, Attila Gürsoy |
Bioinform. | 4 |
| 2022 | SARS-CoV-2 Interactome 3D: A Web interface for 3D visualization and analysis of SARS-CoV-2-human mimicry and interactionsabstractSUMMARY: We present a web-based server for navigating and visualizing possible interactions between SARS-CoV-2 and human host proteins. The interactions are obtained from HMI_Pred which relies on the rationale that virus proteins mimic host proteins. The structural alignment of the viral protein with one side of the human protein-protein interface determines the mimicry. The mimicked human proteins and predicted interactions, and the binding sites are presented. The user can choose one of the 18 SARS-CoV-2 protein structures and visualize the potential 3D complexes it forms with human proteins. The mimicked interface is also provided. The user can superimpose two interacting human proteins in order to see whether they bind to the same site or different sites on the viral protein. The server also tabulates all available mimicked interactions together with their match scores and number of aligned residues. This is the first server listing and cataloging all interactions between SARS-CoV-2 and human protein structures, enabled by our innovative interface mimicry strategy. AVAILABILITY AND IMPLEMENTATION: The server is available at https://interactome.ku.edu.tr/sars/. Damla Ovek, Ameer Taweel, Zeynep Abali, Ece Tezsezen, Yunus Emre Koroglu, Chung-Jung Tsai, Ruth Nussinov, Ozlem Keskin, Attila Gürsoy |
Bioinform. | 7 |
| 2021 | Antigen Binding Reshapes Antibody Energy Landscape and Conformation DynamicsabstractThis study elucidates the conformation dynamics of the free and antigen-bound antibody. Previous work has verified that antigen binding allosterically promotes Fc receptor recognition. Analysis of extensive molecular dynamics simulations finds that the energy landscape may play a decisive role in coordinating conformation changes but does not provide connections between the various conformational states. Here we provide such a connection. To obtain a detailed understanding of the impact of antigen binding on antibody conformation dynamics, this study utilizes Markov State Models to summarize the conformation dynamics probed in silico. We additionally equip these models with the ability to directly exploit the energy landscape view of dynamics via a computational method that detects energy basins and so allows utilizing detected basins as macrostates for the Markov State Model. Our study reveals many interesting findings and suggests that the antigen-bound form with high energy may provide many dynamic processes to further enhance co-factor binding of the antibody in the next step. Kazi Lutful Kabir, Ruth Nussinov, Buyong Ma, Amarda Shehu |
BIBM | 2 |
| 2021 | A network-based deep learning methodology for stratification of tumor mutationsabstractMOTIVATION: Tumor stratification has a wide range of biomedical and clinical applications, including diagnosis, prognosis and personalized treatment. However, cancer is always driven by the combination of mutated genes, which are highly heterogeneous across patients. Accurately subdividing the tumors into subtypes is challenging. RESULTS: We developed a network-embedding based stratification (NES) methodology to identify clinically relevant patient subtypes from large-scale patients' somatic mutation profiles. The central hypothesis of NES is that two tumors would be classified into the same subtypes if their somatic mutated genes located in the similar network regions of the human interactome. We encoded the genes on the human protein-protein interactome with a network embedding approach and constructed the patients' vectors by integrating the somatic mutation profiles of 7344 tumor exomes across 15 cancer types. We firstly adopted the lightGBM classification algorithm to train the patients' vectors. The AUC value is around 0.89 in the prediction of the patient's cancer type and around 0.78 in the prediction of the tumor stage within a specific cancer type. The high classification accuracy suggests that network embedding-based patients' features are reliable for dividing the patients. We conclude that we can cluster patients with a specific cancer type into several subtypes by using an unsupervised clustering algorithm to learn the patients' vectors. Among the 15 cancer types, the new patient clusters (subtypes) identified by the NES are significantly correlated with patient survival across 12 cancer types. In summary, this study offers a powerful network-based deep learning methodology for personalized cancer medicine. AVAILABILITY AND IMPLEMENTATION: Source code and data can be downloaded from https://github.com/ChengF-Lab/NES. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chuang Liu 0001, Zi-Ke Zhang, Ruth Nussinov, Feixiong Cheng |
Bioinform. | 4 |
| 2020 | Network-based prediction of drug-target interactions using an arbitrary-order proximity embedded deep forestabstractMOTIVATION: Systematic identification of molecular targets among known drugs plays an essential role in drug repurposing and understanding of their unexpected side effects. Computational approaches for prediction of drug-target interactions (DTIs) are highly desired in comparison to traditional experimental assays. Furthermore, recent advances of multiomics technologies and systems biology approaches have generated large-scale heterogeneous, biological networks, which offer unexpected opportunities for network-based identification of new molecular targets among known drugs. RESULTS: In this study, we present a network-based computational framework, termed AOPEDF, an arbitrary-order proximity embedded deep forest approach, for prediction of DTIs. AOPEDF learns a low-dimensional vector representation of features that preserve arbitrary-order proximity from a highly integrated, heterogeneous biological network connecting drugs, targets (proteins) and diseases. In total, we construct a heterogeneous network by uniquely integrating 15 networks covering chemical, genomic, phenotypic and network profiles among drugs, proteins/targets and diseases. Then, we build a cascade deep forest classifier to infer new DTIs. Via systematic performance evaluation, AOPEDF achieves high accuracy in identifying molecular targets among known drugs on two external validation sets collected from DrugCentral [area under the receiver operating characteristic curve (AUROC) = 0.868] and ChEMBL (AUROC = 0.768) databases, outperforming several state-of-the-art methods. In a case study, we showcase that multiple molecular targets predicted by AOPEDF are associated with mechanism-of-action of substance abuse disorder for several marketed drugs (such as aripiprazole, risperidone and haloperidol). AVAILABILITY AND IMPLEMENTATION: Source code and data can be downloaded from https://github.com/ChengF-Lab/AOPEDF. Xiangxiang Zeng, Siyi Zhu, Yuan Hou, Pengyue Zhang, Lang Li 0001, L. Frank Huang, Stephen J. Lewis, Ruth Nussinov, Feixiong Cheng |
Bioinform. | 9 |
| 2020 | Individualized genetic network analysis reveals new therapeutic vulnerabilities in 6, 700 cancer genomesabstractTumor-specific genomic alterations allow systematic identification of genetic interactions that promote tumorigenesis and tumor vulnerabilities, offering novel strategies for development of targeted therapies for individual patients. We develop an Individualized Network-based Co-Mutation (INCM) methodology by inspecting over 2.5 million nonsynonymous somatic mutations derived from 6,789 tumor exomes across 14 cancer types from The Cancer Genome Atlas. Our INCM analysis reveals a higher genetic interaction burden on the significantly mutated genes, experimentally validated cancer genes, chromosome regulatory factors, and DNA damage repair genes, as compared to human pan-cancer essential genes identified by CRISPR-Cas9 screenings on 324 cancer cell lines. We find that genes involved in the cancer type-specific genetic subnetworks identified by INCM are significantly enriched in established cancer pathways, and the INCM-inferred putative genetic interactions are correlated with patient survival. By analyzing drug pharmacogenomics profiles from the Genomics of Drug Sensitivity in Cancer database, we show that the network-predicted putative genetic interactions (e.g., BRCA2-TP53) are significantly correlated with sensitivity/resistance of multiple therapeutic agents. We experimentally validated that afatinib has the strongest cytotoxic activity on BT474 (IC50 = 55.5 nM, BRCA2 and TP53 co-mutant) compared to MCF7 (IC50 = 7.7 μM, both BRCA2 and TP53 wild type) and MDA-MB-231 (IC50 = 7.9 μM, BRCA2 wild type but TP53 mutant). Finally, drug-target network analysis reveals several potential druggable genetic interactions by targeting tumor vulnerabilities. This study offers a powerful network-based methodology for identification of candidate therapeutic pathways that target tumor vulnerabilities and prioritization of potential pharmacogenomics biomarkers for development of personalized cancer medicine. Chuang Liu 0001, Junfei Zhao, Weiqiang Lu, Yao Dai, Jennifer Hockings, Yadi Zhou, Ruth Nussinov, Charis Eng, Feixiong Cheng |
PLoS Comput. Biol. | 7 |
| 2019 | deepDR: a network-based deep learning approach to in silico drug repositioningabstractMOTIVATION: Traditional drug discovery and development are often time-consuming and high risk. Repurposing/repositioning of approved drugs offers a relatively low-cost and high-efficiency approach toward rapid development of efficacious treatments. The emergence of large-scale, heterogeneous biological networks has offered unprecedented opportunities for developing in silico drug repositioning approaches. However, capturing highly non-linear, heterogeneous network structures by most existing approaches for drug repositioning has been challenging. RESULTS: In this study, we developed a network-based deep-learning approach, termed deepDR, for in silico drug repurposing by integrating 10 networks: one drug-disease, one drug-side-effect, one drug-target and seven drug-drug networks. Specifically, deepDR learns high-level features of drugs from the heterogeneous networks by a multi-modal deep autoencoder. Then the learned low-dimensional representation of drugs together with clinically reported drug-disease pairs are encoded and decoded collectively via a variational autoencoder to infer candidates for approved drugs for which they were not originally approved. We found that deepDR revealed high performance [the area under receiver operating characteristic curve (AUROC) = 0.908], outperforming conventional network-based or machine learning-based approaches. Importantly, deepDR-predicted drug-disease associations were validated by the ClinicalTrials.gov database (AUROC = 0.826) and we showcased several novel deepDR-predicted approved drugs for Alzheimer's disease (e.g. risperidone and aripiprazole) and Parkinson's disease (e.g. methylphenidate and pergolide). AVAILABILITY AND IMPLEMENTATION: Source code and data can be downloaded from https://github.com/ChengF-Lab/deepDR. SUPPLEMENTARY INFORMATION: Supplementary data are available online at Bioinformatics. Xiangxiang Zeng, Siyi Zhu, Xiangrong Liu, Yadi Zhou, Ruth Nussinov, Feixiong Cheng |
Bioinform. | 5 |
| 2019 | Review: Precision medicine and driver mutations: Computational methods, functional assays and conformational principles for interpreting cancer driversabstractAt the root of the so-called precision medicine or precision oncology, which is our focus here, is the hypothesis that cancer treatment would be considerably better if therapies were guided by a tumor's genomic alterations. This hypothesis has sparked major initiatives focusing on whole-genome and/or exome sequencing, creation of large databases, and developing tools for their statistical analyses-all aspiring to identify actionable alterations, and thus molecular targets, in a patient. At the center of the massive amount of collected sequence data is their interpretations that largely rest on statistical analysis and phenotypic observations. Statistics is vital, because it guides identification of cancer-driving alterations. However, statistics of mutations do not identify a change in protein conformation; therefore, it may not define sufficiently accurate actionable mutations, neglecting those that are rare. Among the many thematic overviews of precision oncology, this review innovates by further comprehensively including precision pharmacology, and within this framework, articulating its protein structural landscape and consequences to cellular signaling pathways. It provides the underlying physicochemical basis, thereby also opening the door to a broader community. Ruth Nussinov, Hyunbum Jang, Chung-Jung Tsai, Feixiong Cheng |
PLoS Comput. Biol. | 1 |
| 2019 | Correction: Review: Precision medicine and driver mutations: Computational methods, functional assays and conformational principles for interpreting cancer driversabstract[This corrects the article DOI: 10.1371/journal.pcbi.1006658.]. Ruth Nussinov, Hyunbum Jang, Chung-Jung Tsai, Feixiong Cheng |
PLoS Comput. Biol. | 1 |
| 2019 | Protein ensembles link genotype to phenotypeabstractClassically, phenotype is what is observed, and genotype is the genetic makeup. Statistical studies aim to project phenotypic likelihoods of genotypic patterns. The traditional genotype-to-phenotype theory embraces the view that the encoded protein shape together with gene expression level largely determines the resulting phenotypic trait. Here, we point out that the molecular biology revolution at the turn of the century explained that the gene encodes not one but ensembles of conformations, which in turn spell all possible gene-associated phenotypes. The significance of a dynamic ensemble view is in understanding the linkage between genetic change and the gained observable physical or biochemical characteristics. Thus, despite the transformative shift in our understanding of the basis of protein structure and function, the literature still commonly relates to the classical genotype-phenotype paradigm. This is important because an ensemble view clarifies how even seemingly small genetic alterations can lead to pleiotropic traits in adaptive evolution and in disease, why cellular pathways can be modified in monogenic and polygenic traits, and how the environment may tweak protein function. Ruth Nussinov, Chung-Jung Tsai, Hyunbum Jang |
PLoS Comput. Biol. | 1 |
| 2019 | A component overlapping attribute clustering (COAC) algorithm for single-cell RNA sequencing data analysis and potential pathobiological implicationsabstractRecent advances in next-generation sequencing and computational technologies have enabled routine analysis of large-scale single-cell ribonucleic acid sequencing (scRNA-seq) data. However, scRNA-seq technologies have suffered from several technical challenges, including low mean expression levels in most genes and higher frequencies of missing data than bulk population sequencing technologies. Identifying functional gene sets and their regulatory networks that link specific cell types to human diseases and therapeutics from scRNA-seq profiles are daunting tasks. In this study, we developed a Component Overlapping Attribute Clustering (COAC) algorithm to perform the localized (cell subpopulation) gene co-expression network analysis from large-scale scRNA-seq profiles. Gene subnetworks that represent specific gene co-expression patterns are inferred from the components of a decomposed matrix of scRNA-seq profiles. We showed that single-cell gene subnetworks identified by COAC from multiple time points within cell phases can be used for cell type identification with high accuracy (83%). In addition, COAC-inferred subnetworks from melanoma patients' scRNA-seq profiles are highly correlated with survival rate from The Cancer Genome Atlas (TCGA). Moreover, the localized gene subnetworks identified by COAC from individual patients' scRNA-seq data can be used as pharmacogenomics biomarkers to predict drug responses (The area under the receiver operating characteristic curves ranges from 0.728 to 0.783) in cancer cell lines from the Genomics of Drug Sensitivity in Cancer (GDSC) database. In summary, COAC offers a powerful tool to identify potential network-based diagnostic and pharmacogenomics biomarkers from large-scale scRNA-seq profiles. COAC is freely available at https://github.com/ChengF-Lab/COAC. He Peng, Xiangxiang Zeng, Yadi Zhou, Ruth Nussinov, Feixiong Cheng |
PLoS Comput. Biol. | 5 |
| 2017 | Structural host-microbiota interaction networksabstractHundreds of different species colonize multicellular organisms making them "metaorganisms". A growing body of data supports the role of microbiota in health and in disease. Grasping the principles of host-microbiota interactions (HMIs) at the molecular level is important since it may provide insights into the mechanisms of infections. The crosstalk between the host and the microbiota may help resolve puzzling questions such as how a microorganism can contribute to both health and disease. Integrated superorganism networks that consider host and microbiota as a whole-may uncover their code, clarifying perhaps the most fundamental question: how they modulate immune surveillance. Within this framework, structural HMI networks can uniquely identify potential microbial effectors that target distinct host nodes or interfere with endogenous host interactions, as well as how mutations on either host or microbial proteins affect the interaction. Furthermore, structural HMIs can help identify master host cell regulator nodes and modules whose tweaking by the microbes promote aberrant activity. Collectively, these data can delineate pathogenic mechanisms and thereby help maximize beneficial therapeutics. To date, challenges in experimental techniques limit large-scale characterization of HMIs. Here we highlight an area in its infancy which we believe will increasingly engage the computational community: predicting interactions across kingdoms, and mapping these on the host cellular networks to figure out how commensal and pathogenic microbiota modulate the host signaling and broadly cross-species consequences. Emine Guven-Maiorov, Chung-Jung Tsai, Ruth Nussinov |
PLoS Comput. Biol. | 3 |
| 2017 | Network approaches and applications in biologyabstractDOAJ is a unique and extensive index of diverse open access journals from around the world, driven by a growing community, committed to ensuring quality content is freely available online for everyone. Trey Ideker, Ruth Nussinov |
PLoS Comput. Biol. | 2 |
| 2017 | How can computation advance microbiome research?abstractMicrobiome science is already in the fast lane, and computations share the credit.How can computations further speed up the pace of microbiome research as it dashes farther, scaling higher summits?What can computations accomplish?What questions can computations help address and what types of algorithms and software should computational biologists aim to develop?These are tantalizing questions that many of us are facing, probably in particular new groups aiming to venture into a relatively nascent and fast-developing area that promises rapid accumulation of data, emergence of alluring new concepts, and significant discoveries.It almost seems like every day, inspiring observations come to light.Consider, for example, the possible link that was observed between gut bacteria of mice and Parkinson disease, in which changes in the bacteria populating the gut apparently appear to be associated with a decline in motor skills.How can this be?It turns out that 70% of all neurons in the peripheral nervous system are located in the gut, and these are directly connected to the central nervous system through the vagus nerve [1].A remarkable earlier study suggested that Parkinson disease may start in the stomach, because people who had their vagus nerve cut to treat gastric ulcers exhibited a lower risk of Parkinson disease than those whose treatment involved only a partial dissection [2].In another striking discovery, it was found that gut bacteria may have a role in autism and, curiously, there is evidence that a single species of gut bacteria can reverse autism-related social behavior in mice [3].In fact, there is emerging evidence for relationships between disruptions in the human microbiome and cancer [4], cardiovascular disease [5], obesity ([6] and more, e.g., [7]), food allergies [8], and asthma [9], among many other diseases.Beyond these emerging examples, there are well-established links (yet still quite recent!) between the diversity of a healthy gut microbiome and protection against Clostridium difficile infection, which has led to therapeutic interventions with remarkable success [10].The microbiota also has important roles in cancer therapy [11].Tumor growth can be suppressed by biofilm-producing bacteria [12].So how can computation accelerate research in the examples above?Broadly, computations can make headway in problems ranging from characterizing taxonomic diversity, classification of microbial species, and tracing their evolution.Computational methods are emerging to facilitate the detection and quantification of diverse patterns among these data, as well as the construction of microbial networks and cross interactions between members of microbial communities.Computation can tackle complex data (e.g., genomic, transcriptomic, proteomic, and metabolomic) on the interactions between microbial communities and their hosts, towards the most challenging question of the quantification of the impact of the human microbiome on our health, as in the examples above. Ruth Nussinov, Jason A. Papin |
PLoS Comput. Biol. | 1 |
| 2017 | Computing the Dynamic Supramolecular Structural ProteomeabstractCells execute their functions through protein interactions.The pathways they link are neither discrete nor spatially separated, as typically depicted in cellular diagrams.Cellular diagrams are useful; however, they neglect the physical structure of cell signaling.In reality, functions are shaped by molecular transitions between small-and large-supramolecular assemblies.Even though they are preorganized, they consist of clusters that are loose and dynamic.Importantly too, they are often anchored in the membrane and interact with scaffolding proteins and the cytoskeleton.Their continuum may physically span the cell.Indeed, efficient, productive, and reliable cell signaling can only take place through transient and cooperative protein-protein interactions, not through stochastic, diffusion-controlled processes.Despite this, current computational approaches to the modeling of the structural proteome still do not fully account for the in vivo, real physical cell organization.The enigmas of the assembly sizes and dynamic conformational distributions-and the diverse cellular environments that influence thempresent daunting challenges, which we have only begun to address.How will we then compute the realistic structural proteome in the next decade?How will we overcome the challenges that we confront, and which methods should we develop to meet them?Clearly, our views of protein structure and function have undergone a revolution.We no longer believe that a protein exists in only two distinct (active and inactive) states.We now recognize that even though a specific function is executed by a distinct active state, proteins (and other bio-macromolecules) exist in ensembles of states.The structure-function paradigm that now dominates molecular biology was inspired by physics and chemistry, which stipulate that even living things must abide by the laws of quantum mechanics and structural chemistry.This paradigm argues that biomolecules should be viewed-and described-statistically, not statically.Though challenging, eventually, to realistically capture the functional versatility and model the working proteome, we must consider conformational ensembles and their allosteric shifts, which result in changes to the populations of the conformations.Moreover, within this framework, the heterogeneous cellular environments, as well as allosteric covalent post-translational modifications, cannot be overlooked.Have we indeed treated the structural proteome as such in our computations?Determining the structures of protein assemblies has long been a vastly important aim of structural biology.The problem is challenging: a pair of protein structures can interact by complementary patches of surfaces.The patch size and identity are unknown, and it is difficult to assess which patches on one protein interact with which patches on the other.In principle, Ruth Nussinov, Jason A. Papin, Ilya A. Vakser |
PLoS Comput. Biol. | 1 |
| 2016 | Genome Landscapes of Disease: Strategies to Predict the Phenotypic Consequences of Human Germline and Somatic VariationabstractComputational biology can marry disciplines to help solve some of the most pressing problems in medical research.When built on fundamental evolutionary, biological, and/or physics principles and provided with large quantities of diverse experimental data, computing power, and rigorous statistics tools, efficient and effective computational strategies can help unlock the secrets of the genome to cure disease.Computational biology has undertaken this challenge, spearheading efforts to decode the genetic blueprints encrypted in the human germline and in somatic variations to decipher complex phenotypic consequences.It devises new and creative approaches to sift through the immense-and rapidly growing-assembled human genetic material, its products, the proteome and the metabolome, and their possible clinical ramifications.The cell is the basic unit of life; the nucleus houses the genetic material that is believed to have provided a distinct advantage to the evolving cell.The organization of the genome varies; it depends on cell type, stage of development, differentiation, disease status, and more.The higher-order spatial and temporal organization of genomes-which itself is a function of conditions and environment-is a driver of biological function in differentiation, development, and disease.Studies of genomics, epigenetics, big data analysis, imaging, and clinical cell and molecular biology all benefit from rigorous computational biology analyses and modeling.They profit from the testable hypotheses that computational biology provides.These link the genetic material to the physiological cell state; however, in-depth understanding of-and predicting-the phenotypic relationship to genome landscapes also requires effective algorithms to unravel the outcomes of small-and large-scale alterations on DNA, RNA, and protein molecules.Diseases can be viewed as perturbed states of molecular systems.They can be of different types, including single-gene (monogenic) diseases and multifactorial diseases, such as cancers, immune system diseases, neurodegenerative diseases, cardiovascular diseases, and metabolic diseases.They may involve genetic alterations or be more complex.They can also be infectious diseases where interacting molecular networks of both pathogens and humans are involved, with the pathogen protein subverting the cell's machinery.The rapid development of next-generation methods for whole-genome, whole-exome, and targeted sequencing, complemented by earlier microarray technologies, including comparative Rachel Karchin, Ruth Nussinov |
PLoS Comput. Biol. | 2 |
| 2016 | Allostery: An Overview of Its History, Concepts, Methods, and ApplicationsabstractThe concept of allostery has evolved in the past century. In this Editorial, we briefly overview the history of allostery, from the pre-allostery nomenclature era starting with the Bohr effect (1904) to the birth of allostery by Monod and Jacob (1961). We describe the evolution of the allostery concept, from a conformational change in a two-state model (1965, 1966) to dynamic allostery in the ensemble model (1999); from multi-subunit (1965) proteins to all proteins (2004). We highlight the current available methods to study allostery and their applications in studies of conformational mechanisms, disease, and allosteric drug discovery. We outline the challenges and future directions that we foresee. Altogether, this Editorial narrates the history of this fundamental concept in the life sciences, its significance, methodologies to detect and predict it, and its application in a broad range of living systems. Jin Liu 0005, Ruth Nussinov |
PLoS Comput. Biol. | 2 |
| 2016 | Principles and Overview of Sampling Methods for Modeling Macromolecular Structure and DynamicsabstractInvestigation of macromolecular structure and dynamics is fundamental to understanding how macromolecules carry out their functions in the cell. Significant advances have been made toward this end in silico, with a growing number of computational methods proposed yearly to study and simulate various aspects of macromolecular structure and dynamics. This review aims to provide an overview of recent advances, focusing primarily on methods proposed for exploring the structure space of macromolecules in isolation and in assemblies for the purpose of characterizing equilibrium structure and dynamics. In addition to surveying recent applications that showcase current capabilities of computational methods, this review highlights state-of-the-art algorithmic techniques proposed to overcome challenges posed in silico by the disparate spatial and time scales accessed by dynamic macromolecules. This review is not meant to be exhaustive, as such an endeavor is impossible, but rather aims to balance breadth and depth of strategies for modeling macromolecular structure and dynamics for a broad audience of novices and experts. Tatiana Maximova, Ryan Moffatt, Buyong Ma, Ruth Nussinov, Amarda Shehu |
PLoS Comput. Biol. | 4 |
| 2016 | Computing BiologyabstractIs Computational Biology increasingly-and steadily-progressing toward addressing the mammoth challenge of actually computing biology?That is, have we reached the stage where we do not support biological research but drive it?This question is vitally important for allyoung and established computational biologists.Even though forecasting future research can be risky, we still venture to predict that the future will see considerably more research projects drifting toward this ambitious aspiration.Computational Biology is powerful for abstracting signatures of disease, for predicting it, and for proposing medications.It is effective in figuring out disease mechanisms and forceful in bridging experimental disciplines to obtain testable predictions.However, perhaps its biggest challenges lie in putting together the available broad and disparate information, devising tools to efficiently and effectively carry out these tasks while sifting through noise and recognizing cell specificity, and most importantly coming up with sound, coherent, and testable schemes.Among the examples of the complexity and the type of questions that we will increasingly face are the vast potential implications of findings that underscore the role of commensal microbiota in disease treatments.Further, the mechanisms which are involved-on the molecular level-are not understood and neither do we fully understand in detail how pathogens can modulate the host immune response and subvert it to their own advantage.We would like to identify-and understand-signatures of chronic inflammatory diseases; recurrent viral sequences in multiple patients and multiple cancers; we would like to map disease risks and to figure out what are the mechanisms for the distinct signaling of specific isoforms of oncogenic proteins in specific cancers.Analyses of cancer genomes points to driver mutations; however, we are baffled by the higher frequencies of specific mutations in certain tissues.We are also mystified by the complexity of the cellular network and its apparent redundancy which results in drug resistance.These are mere examples demonstrating the enormity of the questions which are facing us.Perhaps most tellingly is the fact that we are often even struggling to articulate specific objectives, to delineate the available knowledge and to evaluate apparent successes.Within the grand challenge to compute biology, efforts to compute human health fosters multidirectional and multidisciplinary integration of basic, patient-oriented, and populationbased research at different levels and across scales.Diverse data accumulate at increasingly high rates and quantities.We do not have a crystal ball for the future of research; we expect, however, that human health is going to be at the center.Human health is complex and encompasses multiple areas; thus, focusing on human health does not necessarily imply a narrower direction.Indeed, we expect that it will drive expansion, algorithm and tool development, and integration, all merging with experiments.Tools involving pattern recognition in images may Ruth Nussinov, Jason A. Papin |
PLoS Comput. Biol. | 1 |
| 2015 | Mapping the Conformation Space of Wildtype and Mutant H-Ras with a Memetic, Cellular, and Multiscale Evolutionary AlgorithmabstractAn important goal in molecular biology is to understand functional changes upon single-point mutations in proteins. Doing so through a detailed characterization of structure spaces and underlying energy landscapes is desirable but continues to challenge methods based on Molecular Dynamics. In this paper we propose a novel algorithm, SIfTER, which is based instead on stochastic optimization to circumvent the computational challenge of exploring the breadth of a protein's structure space. SIfTER is a data-driven evolutionary algorithm, leveraging experimentally-available structures of wildtype and variant sequences of a protein to define a reduced search space from where to efficiently draw samples corresponding to novel structures not directly observed in the wet laboratory. The main advantage of SIfTER is its ability to rapidly generate conformational ensembles, thus allowing mapping and juxtaposing landscapes of variant sequences and relating observed differences to functional changes. We apply SIfTER to variant sequences of the H-Ras catalytic domain, due to the prominent role of the Ras protein in signaling pathways that control cell proliferation, its well-studied conformational switching, and abundance of documented mutations in several human tumors. Many Ras mutations are oncogenic, but detailed energy landscapes have not been reported until now. Analysis of SIfTER-computed energy landscapes for the wildtype and two oncogenic variants, G12V and Q61L, suggests that these mutations cause constitutive activation through two different mechanisms. G12V directly affects binding specificity while leaving the energy landscape largely unchanged, whereas Q61L has pronounced, starker effects on the landscape. An implementation of SIfTER is made available at http://www.cs.gmu.edu/~ashehu/?q=OurTools. We believe SIfTER is useful to the community to answer the question of how sequence mutations affect the function of a protein, when there is an abundance of experimental structures that can be exploited to reconstruct an energy landscape that would be computationally impractical to do via Molecular Dynamics. Rudy Clausen, Buyong Ma, Ruth Nussinov, Amarda Shehu |
PLoS Comput. Biol. | 3 |
| 2015 | How to Write a Presubmission InquiryabstractLike many other journals, the journal PLOS Computational Biology admits and in some cases requires presubmission inquiries to be submitted before the submission of a full paper. Presubmission inquiries serve the purpose of informing the journal’s Editorial Board of the essence of the intended submission. Based on the information in the inquiry, the editors can make a quick assessment of its contribution with respect to the criteria for publication in the journal. This assessment is then communicated to the authors. This enables fast turnaround to the authors about the basic suitability of a submission for processing by the journal and spares the editors and reviewers the effort of detailed inspection of submissions that clearly do not meet the criteria of the journal.
In this Editorial, we give suggestions for preparing presubmission inquiries for journal submissions. We exemplify these suggestions with reference to presubmission inquiries for the Methods section of PLOS Computational Biology. However, our suggestions generalize to presubmission inquiries of other kinds and for other journals, and in places, we will make specific comments to that effect.
Over two years ago, PLOS Computational Biology opened a special section dedicated to Methods papers. As the scope statement spells out,
Methods papers should describe outstanding methods of exceptional importance that have been shown, or have the promise to provide new biological insights. The method must already be widely adopted, or have the promise of wide adoption by a broad community of users. Enhancements to existing published methods will only be considered if those enhancements bring exceptional new capabilities.
Since Methods papers are different from other research papers in PLOS Computational Biology and also differ from typical papers on bioinformatics methods published in other journals, a mandatory presubmission stage has been introduced for the submission of Methods papers to PLOS Computational Biology. (Note that a presubmission inquiry is not mandatory for general research papers in PLOS Computational Biology, though it is also mandatory for submission to the Software papers category.)
Since the Methods section was launched in October 2012, we have received 334 presubmission inquiries. For roughly half of them (159), we encouraged submission and received full papers, of which 41 papers were published, so far, as Methods papers. We find that, while many presubmission inquiries are informative enough to make an educated decision on the submission, we also receive a number of presubmission inquiries that are not sufficiently informative, such that submission may be discouraged not on the basis of the quality or scope of the paper but on that of the presubmission inquiry. We generally do not allow revisions of presubmission inquiries. In order to minimize the number of papers that fail to get a chance to be published in PLOS Computational Biology merely because of the inadequacy of the presubmission inquiry, here we give a number of suggestions for preparing such an inquiry.
The goal of a presubmission inquiry is to make the statement to the Editorial Board that the paper to be submitted reasonably satisfies the criteria detailed in the scope statement. The presubmission inquiry must be detailed enough to convincingly make that point. Most of the insufficiently informative inquiries that we get are either too terse, i.e., they do not give enough detail, or they are not specific enough. Therefore, we suggest a way of structuring a presubmission inquiry. These suggestions are the result of our experience with presubmission inquiries over the past couple of years.
What is the problem? Please summarize the problem domain and statement and the relevance of the problem to the general readership of the journal—in the case of PLOS Computational Biology, the biological research community or a substantial subcommunity. What is that subcommunity? How relevant is the problem to them?
What is the innovation? Here it is important that you give enough detail on your contribution to allow the Editorial Board to form an image of the substance and relevance of the advance over the state of the art in the field and over your previous work. If your paper presents material that rests on or is related to your own previous publications, this entails addressing dual publication issues. In order to argue your point, you have to summarize the state of the art on which you base your contribution and give the essential ingredients of your innovation. Depending on the journal, the innovation can take different shapes: a contribution to technology or experimental design, a biological finding, a methodical or theoretical piece of work, etc. For papers in the Methods section of PLOS Computational Biology, the computational method is expected to be at the center of the innovation. This section is not for papers whose methodical core has been published elsewhere and for which you present a—possibly extended or modified—application scenario. Also, studies presenting a comparative assessment of existing methods on an application domain are not within the scope of a Methods paper. We expect a concrete and specific relationship to underlying biological issues. This is why general methods on statistical learning that find their application in biology as well as in other fields of science are typically not considered in scope, unless the paper focuses on sufficiently deep issues of the configuration of the method that are specific to biology. Finally, the method must be the major innovation of the paper. However interesting it may be, a biological finding that has been obtained with methods that are prepublished or only minor modification of prepublished methods is not within the scope of the Methods papers category. (On the other hand, it may be a suitable General research paper for the journal.)
How is the method validated? Validation can take manifold forms but is a key element in most scientific papers. For a theoretical paper, the validation often takes the form of a proof. In contrast, methods in computational biology have manifold forms of validation and, usually, a single form is not sufficient to make the point. For instance, for the Methods section of PLOS Computational Biology, we expect more than an anecdotal validation based on a couple of biological use cases. A validation purely on synthetic data is not sufficient either. Rather, the validation must make a convincing argument for the general applicability of the method in a substantial biological problem domain. Please note, however, that papers that center on the validation of a method that has been published elsewhere are not considered in scope either. The paper has to contain both the method and its application.
How is the method being made available? Availability of research results becomes an increasingly desired and often required aspect of a publication. For papers that are based on experimental data, making the data and the protocols of the experimental design available is a prerequisite for making the research reproducible. For methods papers, reusability of the method also becomes an issue. For the Methods section of PLOS Computational Biology, we only accept papers on methods that are useful to and can be readily applied by other scientists. The best way of satisfying this criterion is to make the software implementation of the methods openly available. For methods that are not based on software, a workable protocol for how to use the methods must be provided.
If you have the full submission ready at the time of the presubmission inquiry, we encourage you to attach it to the inquiry as an optional supplement and mention that you have done this in your cover letter. However, your presubmission inquiry should be worded such that the editors do not need to inspect the complete paper for making their assessment.
We wish you much success with your future submission to the Methods section of PLOS Computational Biology. Thomas Lengauer, Ruth Nussinov |
PLoS Comput. Biol. | 2 |
| 2015 | Advancements and Challenges in Computational BiologyabstractComputational biology has soared from being an auxiliary discipline to being a crucial element for progress in practically all aspects of the biological sciences. In this annual Editorial, I would like to step back, consider significant computational biology advances of the last decade, and reflect on some key challenges ahead. The timing is particularly appropriate. PLOS Computational Biology, the premier journal in computational biology, is approaching its tenth anniversary. The task is daunting; not only has the field come a long way in ten years but it is broad with many advances to consider. In addition, since computational biology has become closely tied to experimental research, progress is not purely computational; it is tied to experiment. And that's as it should be. Ten years ago, computational biology was not entirely trusted by experimental biologists. By contrast, today computational biology is integrated in the community. It's easier for computational biologists to collaborate across disciplines. Laboratory scientists have a better understanding of the merit of computational models for hypothesis generation as well as the need to iterate between modeling and laboratory testing [1].
We have witnessed huge leaps in biological computing [2]. We now have at our disposal large information-rich resources, and we are increasingly able to integrate and understand the vast quantities of data that they encompass. We have also made big strides toward multiscale biological modeling, and we have a vastly more networked world of researchers and their data. Analysis of massive gene expression and proteomic data permitted the construction of comprehensive and predictive models for cellular pathways, as well as software for inferring interaction networks, and steps toward modeling of cells. Genes susceptible to disease have been identified and, on a different level, the electrical behavior of neurons has been modeled. Molecules have been imaged in action and networks that regulate cell functions untangled. Matching targets for selective cancer therapy is difficult. Nonetheless, recent strategies have been proposed to restrict the combinatorial space, minimize toxicity, and increase the precision and power of such restrictive combinations, altogether leading to drugs that could be tested in clinical trials. Leveraging the enhanced identification of drug targets, including repertoires of redundant pathway combinations, has been helped by such innovative concepts [3].
Formidable challenges include: the establishment of computer networks for surveillance of disease; mapping the pathways and biological networks associated with the initiation, growth and spread of cancer; predicting function and mutational dysfunction in disease from the structure of complex molecules; resolving the mechanisms of oncogenic mutations and the cellular network which is rewired in cancer; achieving accurate, efficient, and comprehensive dynamic models; and moving from artificial intelligence to the “connectome”—the connections among all of the neurons of the brain. Multiscale biological modeling—an area where vast progress has been made during the last decade—still faces major challenges. To tackle this aim, hybrid methods across disciplines, scales, and sources are essential. Hybrid methods integrate data from, for example, serial crystallography and time-resolved wide-angle X-ray scattering, micro- and nano-crystals for (future) free-electron lasers, electron microscopy, fluorescence resonance energy transfer (FRET), cross-linking data, small-angle X-ray scattering, crystallography, nuclear magnetic resonance (NMR), and more. Equally important is the development of protocols for model validation. We may expect an influx of models based on experimental data integration. If these are to be deposited in a public archival system, which is now a community aim, such clear protocols are essential for maintaining quality control. Finally, studying the dynamics of large integrated models is increasingly used to improve our understanding of how large complexes function in the cell and how they are regulated. The dynamics of such large associations provides an additional hugely complex layer; to date, we are still struggling to comprehend the dynamics of single molecules and their associations. This is compounded by the fact that large regions of the molecules can be disordered, and multiple temporal post-translational modifications take place, with different combinations spelling distinct functions. On a different level, improved tumor mutational analysis platforms and knowledge of the redundant pathways, which can take over in cancer, may not only supplement known actionable findings but forecast possible cancer progression and resistance. Such forward-looking can be powerful, endowing the oncologist with mechanistic insight and cancer prognosis, and consequently more informed treatment options.
Lastly, the community faces the global challenge of linking genetics to phenotype, including the genetics of cancer. Genetics is mediated by dynamic conformational ensembles. Powerful ideas such as that of the free energy landscape [4], imported from physics and chemistry, can help solve the mysteries of life. Biomolecules are not static sculptures; they are dynamic objects that are always interconverting between structures with varying energies. Such ideas help to understand how and why one-dimensionally connected biomolecules can organize themselves into functionally relevant ensembles of three-dimensional conformation [5]. Designing high affinity drugs that work is yet another highly significant aim.
The significance of any research advance and challenge—achieved or aspired to—is a matter of opinion. The list above is partial, incomplete, and possibly biased toward structural biology and cancer. Nonetheless, this list does indicate the magnitude of the tasks confronting computational biology as a discipline. In the absence of a meaningful way to quantify a journal's contribution to a field, it is unclear whether, and to what extent, PLOS Computational Biology has contributed to each advance and challenge. Manuscripts can be declined, for example, because of the absence of substantiating experimental data at the time, lack of sufficient rigor, or if the manuscripts included new experimental data, the authors may have opted for alternative journals. At the same time, it may also suggest that PLOS Computational Biology needs to be more open and receptive to new concepts. Differentiating between novel ideas that may lead to key advances and speculative propositions can, however, be challenging.
PLOS Computational Biology aims to serve the biological community and welcomes manuscripts addressing all areas of computational biology. We encourage submission of research papers describing novel results that provide significant new insights into biological processes and of methods papers presenting new protocols for tackling key problems that have been shown, or have the promise to provide, new biological insights. We aspire to be the journal that will publish key computational advances in the next decade with the rigor that PLOS Computational Biology is known for. The PLOS Computational Biology editorial team seeks to identify and publish only the most outstanding papers, aiming to consider only those that are of exceptional quality. Our goal of furthering our understanding of living systems through the application of computational methods is shared with the International Society for Computational Biology (ISCB); together, we hope to meet the challenge.
Finally, for 2015, our tenth anniversary year, PLOS Computational Biology plans to publish a series of “Focus Features” addressing key areas of computational biology. We welcome suggestions from our community. Ruth Nussinov |
PLoS Comput. Biol. | 1 |
| 2015 | From "What Is?" to "What Isn't?" Computational BiologyabstractISSN:1553-734X Ruth Nussinov, Sebastian Bonhoeffer, Jason A. Papin, Olaf Sporns |
PLoS Comput. Biol. | 1 |
| 2015 | Computational Methods for Exploration and Analysis of Macromolecular Structure and DynamicsabstractAll processes that maintain and replicate a living cell involve fluctuating biological macromolecules. As computational biologists, our aim is to discern the behavior of macromolecules in a way that experimental biology is not able to achieve. No single technique—experimental or computational—can capture all the relevant scales of cellular functional behavior. In principle, computations are the tools that can integrate different kinds of experimental and computational characterizations at different resolutions to obtain a more complete description of the processes of life. Computer simulations can act as a bridge between the microscopic length and time scales, and the macroscopic world of the laboratory. They can start from a macroscopic experiment-based guess of interactions between molecules, and obtain “exact” predictions of bulk and detailed properties subject to limitations. They are able to test a theory by constructing and simulating the model, and comparing the results with experimental measurements; and they are able to provide models that experiments can test. Computations can provide leads by processing large sets of data, predicting molecular behaviors, and supplying the mechanistic underpinning that experiments alone may not be able to achieve.
Macromolecules play a vital role in countless biological processes, including DNA replication, transcription, genome reorganization in development and in disease, protein synthesis, protein folding, and active transport with molecular motors. Cell signaling, a multistep pathway on length scales from nanometers to micrometers, provides another inclusive example, incorporating all of the above over time and space. Signals are relayed from the extracellular space to the nucleus through dynamic shifts of molecular ensembles. Macromolecular fluctuations underlie signal amplification; they result in a large number of activated molecules across the cell, creating multiple reactions and producing a major cellular response. Changes in fluctuations through binding second messengers can regulate catalysis, and dynamic shifts in conformational ensembles can also take place through binding to membrane lipids. Key hub proteins that govern cell behaviors are often membrane-anchored. Helped by experimental data, computations can model the components of the systems and their transient interactions to provide a useful, integrated view of the flow of information, its regulation, and its deregulation.
Ultimately, we want to make direct quantitative comparisons with experimental data. We would like to reduce the amount of fitting and guesswork; but at the same time we may also be interested in phenomena of a generic nature, or in discriminating between good and bad theories. Doing this well is challenging. To understand the dynamic interplay across multiple scales, to link it to the atomic-scale physicochemical basis of the conformational behavior of single molecules and their interactions, and, ultimately, to relate it to cellular function, we need efficient and reliable methods to sample the macromolecular fluctuations and identify the biological and disease-related states and their transitions. Inspiration may come from a combination of biology and other fields that model dynamic systems.
Macromolecules move, and their movements are needed for a complete picture of life. Computational biology, with concepts imported from physics and chemistry, increasingly plays a major role, which has recently been recognized by the Nobel Committee [1]. The energy landscape underscores the inherent nature of biomolecules, which are dynamical objects that are always interconverting between structures with varying energies. It affirms that biomolecules must be described statistically, not statically. Macromolecules are not static objects; rather, they populate ensembles of conformations. The transitions between these states occur on length scales from tenths of an Angstrom to nanometers, and time scales that can vary from nanoseconds to seconds. These are linked to functionally relevant phenomena such as allosteric signaling and enzyme catalysis.
Computational methods also include those for molecular modeling and refinement of three-dimensional structures, de novo design of proteins, prediction and modeling of protein-ligand interactions and development of docking protocols, and prediction of macromolecular interactions at varying spatial resolutions and timescales. They further encompass methods of ligand screening in drug development and protein-protein docking, methods for assessing sequence-structure-function relationships and prediction of macromolecular function, protocols for molecular visualization and annotation, and geometric and topological characterization of proteins and polynucleotides. This list is still far from complete.
To celebrate its tenth anniversary [2], PLOS Computational Biology presents a special collection of manuscripts focusing on methods exploring macromolecular structure and dynamics. This collection does not aim to cover all methods; it does, however, aim to provide a taste of currently available approaches and strategies toward these aims. Altogether, the collection covers a broad ground: from sampling, detection of rare events, and exposing hidden alternative backbone conformations in X-ray crystallography, to multi-scale visualization of molecular architecture using real-time ambient occlusion; from discrimination between obligatory and non-obligatory protein-protein interactions based on the dynamics of the complex to binding free energies of inhibitors, to predicting the effect of mutations on protein-protein interactions by exploiting interface profiles; from a virtual mixture approach to the study of multistate equilibrium, to identification of misfolded intermediates; from multiscale estimation of binding kinetics using a combination of Brownian dynamics, molecular dynamics, and mile-stoning, to mapping the protein fold universe. This collection underscores the breadth of computational methods in structural biology, and only some of them made it into this special PLOS Computational Biology collection.
Computational structural biology has made tremendous progress over the last two decades. Computational methods were developed for protein structure prediction, macromolecular function and protein design, as well as for drug discovery. It has also undertaken computational challenges related to experimental approaches in structural biology. Along with new experimental tools, higher resolution, and the rising efficiency of experimental approaches leading to huge amounts of accumulating data, computational biology is pushing the frontiers to meet its challenges. We expect that, in the future, the focus of our methods and tools may shift and more integrative tools will be developed, along with methodological adaptation to massively larger quantities of data.
Experimentally, the biological and chemical sciences are now attempting to push boundaries in drug discovery. We may expect that a translational direction will also prevail in computational structural biology. This, however, does not mean only direct drug discovery; for these efforts to be successful, the underlying mechanistic basis of diseases needs to be understood as well. In addition, we expect even stronger emphasis on the human microbiome and its relationship to human health. PLOS Computational Biology aims to meet this challenge and place a larger focus on this important and certain-to-become-central area in the biological sciences. Clearly, many challenges remain for computational biologists in the coming years.
PLOS Computational Biology, along with the computational biology community and the International Society for Computational Biology (ISCB), are poised to take on this challenge. We view this collection as the first in this direction, helping the community toward this aim. Amarda Shehu, Ruth Nussinov |
PLoS Comput. Biol. | 2 |
| 2014 | Making Biomolecular Simulations Accessible in the Post-Nobel Prize EraabstractIn 2013, three pioneers of computational biophysics and structural biology, Martin Karplus, Arieh Warshel, and Michael Levitt, were awarded the Nobel Prize in Chemistry. Although the citation focused on their innovative efforts on integrating quantum mechanical and classical mechanical models to study reactive processes in proteins, the award has also been seen by many researchers in the biomolecular simulation field as recognizing the tremendous value of computations for the investigation of biomolecules in general. From the days when proteins were modeled at the picosecond timescale using a united atom representation [1], or even as coarse-grained beads [2], in vacuum, to modern simulations that approach the millisecond timescale for a fully solvated protein [3], the biomolecular simulation field has, indeed, come a long way. Much of the progress has been due to the efforts of the three laureates, their contemporaries, and many others (e.g., their students) who were inspired by their dream of understanding life by studying “the jiggling and wiggling of atoms” [4]. One could only admire the tremendous courage, imagination, and vision that drove these three scientists to start pursuing their dream in an era when theoretical and computational chemistry largely focused on understanding the interactions and reactivity of small molecules.
Just as the Nobel Prize in 1998 to John Pople and Walter Kohn highlighted both the impact and emerging challenges of quantum chemistry, the 2013 Chemistry Prize should also further inspire us to ponder about the future of computational biology. Clearly, developing methodologies that further enhance the quantitative accuracy and/or complexity of computational models are important and being actively pursued by many researchers. On the quantitative aspect, several community-wide blind tests on observables such as solvation free energies, binding affinities, and pKa values are being held. Provided that the results are disseminated in a constructive manner, these blind tests are highly valuable for helping the community converge towards the most robust and efficient computational algorithms and protocols. On the other hand, it is valuable to bear in mind that in many (certainly not necessarily all) investigations, quantitative computations represent a means to validate the model rather than the ultimate goal, which ought to focus on revealing the physical and chemical principles that govern the biological problem at hand. In other words, understanding qualitative trends is equally important. Therefore, building models with different levels of complexity and identifying robust features relevant to the biological problem remains an important research strategy. After all, in many mechanistic studies, whether at the molecular or cellular scale, the ultimate goal is to establish a conceptual framework to guide the development of novel mechanistic hypotheses and to stimulate new experiments to evaluate them.
Another important issue worth emphasizing in this “Post-Nobel Prize era” concerns making high-quality biomolecular simulation protocols available to the bioscience community, especially to young researchers who have just entered the field and perhaps even researchers who are primarily experimentalists. Such efforts will be essential to further enhancing the impact of biomolecular simulations while maintaining a high level of integrity in the result. In this issue of PLOS Computational Biology, Woodcock and coworkers have made a major step in this direction by describing a set of web-based tutorials and tools for the simulation package Chemistry at HARvard Molecular Mechanics (CHARMM) [5]–[7]; the tools are fittingly and playfully referred to as “CHARMMing.” The web-based tools make it straightforward to set up complex biomolecular simulations, including reduction potential computation for proteins and molecular dynamics simulations using a coarse-grained model. For even an expert in biomolecular simulation, it is often cumbersome to set up a new simulation that requires the generation of force field parameters for cofactors; CHARMMing is helpful in this context by providing an easy access to several automated small molecule force field generation services (e.g., the ParamChem web-server, the MATCH toolkit).
Importantly, CHARMMing goes beyond simply facilitating the set-up of biomolecular simulations by including carefully designed lessons on topics that range from basic simulation tutorials to advanced protocols such as quantum mechanical (QM)/molecular mechanical (MM) calculations and enhanced sampling techniques. The graphic interface allows the “students,” who take those lessons, to understand and modify CHARMM input scripts as well as visualize simulation results. Therefore, CHARMMing is valuable not only as a research tool, but also an educational module that can easily be incorporated into curriculum at both the undergraduate and early graduate level. As a result, CHARMMing is complementary to another valuable web-based research tool, CHARMM-GUI [8], which features a number of sophisticated functionalities, such as setting up membrane simulations [9] and absolute ligand binding affinity calculations [10]. We hope that the set of CHARMMing papers will help stimulate additional efforts in bringing advanced simulations, good computational practices, and thorough analysis of simulation results to the broader biological research community. Although pushing the limit of computational research via method development is always essential, an equally important goal is, to paraphrase what Martin Karplus once stated [11], that experimental (structural) biologists, who know their systems better than anyone else, will make increasing use of molecular dynamics simulations for obtaining a deeper understanding of particular biological systems.
This Editorial was first published as a blog post on PLOS Biologue on July 25, 2014. Qiang Cui 0003, Ruth Nussinov |
PLoS Comput. Biol. | 2 |
| 2014 | The Structural Basis of ATP as an Allosteric ModulatorabstractAdenosine-5'-triphosphate (ATP) is generally regarded as a substrate for energy currency and protein modification. Recent findings uncovered the allosteric function of ATP in cellular signal transduction but little is understood about this critical behavior of ATP. Through extensive analysis of ATP in solution and proteins, we found that the free ATP can exist in the compact and extended conformations in solution, and the two different conformational characteristics may be responsible for ATP to exert distinct biological functions: ATP molecules adopt both compact and extended conformations in the allosteric binding sites but conserve extended conformations in the substrate binding sites. Nudged elastic band simulations unveiled the distinct dynamic processes of ATP binding to the corresponding allosteric and substrate binding sites of uridine monophosphate kinase, and suggested that in solution ATP preferentially binds to the substrate binding sites of proteins. When the ATP molecules occupy the allosteric binding sites, the allosteric trigger from ATP to fuel allosteric communication between allosteric and functional sites is stemmed mainly from the triphosphate part of ATP, with a small number from the adenine part of ATP. Taken together, our results provide overall understanding of ATP allosteric functions responsible for regulation in biological systems. Shaoyong Lu, Wenkang Huang, Qiancheng Shen, Ruth Nussinov, Jian Zhang 0037 |
PLoS Comput. Biol. | 6 |
| 2014 | The Significance of the 2013 Nobel Prize in Chemistry and the Challenges AheadabstractLast week, the 2013 Nobel Prize in Chemistry was awarded to Martin Karplus, Michael Levitt, and Arieh Warshal for “the development of multiscale models for complex chemical systems”. As the Royal Swedish Academy of Sciences noted, “Chemists used to create models of molecules using plastic balls and sticks. Today, the modelling is carried out in computers. In the 1970s, Martin Karplus, Michael Levitt and Arieh Warshel laid the foundation for the powerful programs that are used to understand and predict chemical processes. Computer models mirroring real life have become crucial for most advances made in chemistry today.” Furthermore, “Today the computer is just as important a tool for chemists as the test tube. Simulations are so realistic that they predict the outcome of traditional experiments.” [1]
This event is a milestone for the broad community that PLOS Computational Biology represents. Along with Philip E. Bourne, the Founding Editor-in-Chief, and our Editorial Board, which proudly lists Michael Levitt among its members, I extend the warmest congratulations to the winners. Beyond the specific, personal scientific achievements that have already been widely discussed, we must consider the more general and broader context of this unique prize. Here, I would like to present this Nobel Prize within this framework, emphasizing its magnitude and far-reaching implications not only for computational biology, but for the biological community at large.
In recent decades, molecular biology has progressed by leaps and bounds. Huge technological advances have taken place in sequencing; in mapping structure and dynamics via electron microscopy (EM), X-ray, and nuclear magnetic resonance (NMR); in manipulating imaging of nuclei and cells; in sequencing single biomolecules; and more. These have led to fundamental new insights; biology and medicine have soared to new heights with the DNA double helix providing the molecular basis for genetics and Darwinism. Many steps were required to identify and untangle DNA-RNA-protein sequence-structure-function and reverse transcription processes; RNA enzymes; the importance of key multi-partnered scaffolding molecules under normal physiological conditions and in disease; their structures, mutations, and the principles and mechanisms of their dynamic regulation; and other landmark developments. These involved technological breakthroughs and greater understanding of the specific mechanisms involved. Most of the Nobel prizes in chemistry and medicine in recent years have been awarded at these junctures.
Vast amounts of information on sequences and structures are yet to be explained and pose a challenge for computational biology. Recently, this has been compounded by interdisciplinary studies of the nervous system, posing questions such as how it is structured, how it develops, how it works, the mechanisms of signal processing, and more, all at multiple levels, ranging from the molecular and cellular levels to the systems and cognitive levels. Thus, even if we gain in-depth insight into static properties such as the genomic data and structural snapshots of proteins (DNA and RNA) at different levels of resolution, the truly monumental challenge of understanding their dynamics still looms ahead. And eventually, it is the dynamics of molecules that provides the basis for cells, tissues, and organisms' development and work.
The systems in question operate at all scales: force fields and free energy landscapes relevant for protein folding and function, large complexes, biomolecular recognition involving proteins, DNA, RNA, lipids, post-translational (and DNA) modifications, and interactions with small molecules. On a larger scale we see cellular locomotion, cell division and trafficking, and cell-cell recognition. Furthermore, beyond these lurks the working of the complex cell as a cohesive unit: the cellular network controls metabolism and regulation, intra- and inter-cellular signaling, and the neural circuits of nerve cells, where the activity of one cell directly influences many others. All are dynamic, all change with the cellular environment, and all present a daunting challenge. The relevant timescales range from femtosecond for simple chemical reactions to the eons of evolution; however, all operate with the same underlying physical principles of conformational variability and selection.
At each timescale and corresponding physical size we strive to identify the relevant moving parts and degrees of freedom and to formulate effective—though often approximate—rules for their mutual interactions and resulting motion. Solving, understanding, and computing the dynamic behavior at any given scale is of great interest in its own right and provides approximate dynamical input for the next scale, which is one rung above it. Only at the lowest, most basic scale of individual atoms and electrons are the dynamical rules (electrostatics and Schrodinger's equation) completely well defined. And the all-important work cited by the Nobel Prize Committee and which is carried out by our community is roughly at the first/second level, making it of fundamental importance.
This Nobel Prize is the first given to work in computational biology, indicating that the field has matured and is on a par with experimental biology. It may also be the very first prize given in any area of the exact sciences for calculations. What is different in the present case? I believe that the answer is simple: the present calculations are of much greater interest to a much broader community. In endeavoring to imitate the basic processes of life in silico, great strides are being made toward understanding the secret of life. Computational biology, and simulations, for which Martin Karplus, Michael Levitt, and Arieh Warshal shared the Nobel Prize, can carry the torch leading the sciences to decipher the elemental processes and help alleviate human suffering.
What are the challenges ahead? Are simulations with timescales of microseconds, milliseconds, or beyond, under the current force field framework, capable of producing results in agreement with experiments also for large and complex proteins like membrane receptors? Do the challenges also lie in the type of questions which are asked, for which such long timescale simulations can be useful in providing answers? Or is it the biology behind the questions that is also the key? Ultimately, as in experimental biology which also exploits methods and machines, it is likely to be all of the above. Computations are our treasured tool; they are not our aim. Merely running long molecular dynamics trajectories is unlikely to advance science.
PLOS Computational Biology joins the International Society of Computational Biology (ISCB) and our computational biology community in congratulating the awardees and celebrating this momentous event.
This Editorial was first published as a blog post on PLOS Biologue on October 18, 2013. Ruth Nussinov |
PLoS Comput. Biol. | 1 |
| 2014 | The Structural Pathway of Interleukin 1 (IL-1) Initiated Signaling Reveals Mechanisms of Oncogenic Mutations and SNPs in Inflammation and CancerabstractInterleukin-1 (IL-1) is a large cytokine family closely related to innate immunity and inflammation. IL-1 proteins are key players in signaling pathways such as apoptosis, TLR, MAPK, NLR and NF-κB. The IL-1 pathway is also associated with cancer, and chronic inflammation increases the risk of tumor development via oncogenic mutations. Here we illustrate that the structures of interfaces between proteins in this pathway bearing the mutations may reveal how. Proteins are frequently regulated via their interactions, which can turn them ON or OFF. We show that oncogenic mutations are significantly at or adjoining interface regions, and can abolish (or enhance) the protein-protein interaction, making the protein constitutively active (or inactive, if it is a repressor). We combine known structures of protein-protein complexes and those that we have predicted for the IL-1 pathway, and integrate them with literature information. In the reconstructed pathway there are 104 interactions between proteins whose three dimensional structures are experimentally identified; only 15 have experimentally-determined structures of the interacting complexes. By predicting the protein-protein complexes throughout the pathway via the PRISM algorithm, the structural coverage increases from 15% to 71%. In silico mutagenesis and comparison of the predicted binding energies reveal the mechanisms of how oncogenic and single nucleotide polymorphism (SNP) mutations can abrogate the interactions or increase the binding affinity of the mutant to the native partner. Computational mapping of mutations on the interface of the predicted complexes may constitute a powerful strategy to explain the mechanisms of activation/inhibition. It can also help explain how an oncogenic mutation or SNP works. Saliha Ece Acuner, Attila Gürsoy, Ruth Nussinov, Ozlem Keskin |
PLoS Comput. Biol. | 3 |
| 2014 | A Unified View of "How Allostery Works"abstractThe question of how allostery works was posed almost 50 years ago. Since then it has been the focus of much effort. This is for two reasons: first, the intellectual curiosity of basic science and the desire to understand fundamental phenomena, and second, its vast practical importance. Allostery is at play in all processes in the living cell, and increasingly in drug discovery. Many models have been successfully formulated, and are able to describe allostery even in the absence of a detailed structural mechanism. However, conceptual schemes designed to qualitatively explain allosteric mechanisms usually lack a quantitative mathematical model, and are unable to link its thermodynamic and structural foundations. This hampers insight into oncogenic mutations in cancer progression and biased agonists' actions. Here, we describe how allostery works from three different standpoints: thermodynamics, free energy landscape of population shift, and structure; all with exactly the same allosteric descriptors. This results in a unified view which not only clarifies the elusive allosteric mechanism but also provides structural grasp of agonist-mediated signaling pathways, and guides allosteric drug discovery. Of note, the unified view reasons that allosteric coupling (or communication) does not determine the allosteric efficacy; however, a communication channel is what makes potential binding sites allosteric. Chung-Jung Tsai, Ruth Nussinov |
PLoS Comput. Biol. | 2 |
| 2013 | New Methods Section in PLOS Computational BiologyabstractPLOS Computational Biology is the Public Library of Science journal that targets new biology that is facilitated by computational methods. Since the inception of the journal, biology has been at the center of the scope of PLOS Computational Biology. Thus, each submission is required to put advances in biology into its focus in a very concrete manner and not only focus on computation. This is in contrast to other, more technically oriented journals in the area of computational biology.
The new Methods section of the journal acknowledges the fact that a major methodical advance can, in itself, be so relevant that it deserves transcending the technological domain and being presented to a broader readership including not only computational biologists, but also biologists, the targeted readership of this journal. For this reason, PLOS Computational Biology has installed a special type of submission, the Methods paper. As the scope statement of the journal spells out, “Methods papers should describe outstanding methods of exceptional importance that have been shown, or have the promise to provide new biological insights. The method must already be widely adopted, or have the promise of wide adoption by a broad community of users. Enhancements to existing published methods will only be considered if those enhancements bring exceptional new capabilities.”
In order to render the processing of Methods papers as effective as possible and to limit the effort required on the part of authors and reviewers, a presubmission inquiry is mandatory for Methods papers. In such an inquiry, the authors are requested to present a concise abstract-like statement on what the manuscript they plan to submit entails and why the authors feel that it fits the Methods section of the journal. Presubmission inquiries are given top priority in the paper handling process; the median processing time should be about a week. Within that time, submission is either encouraged or discouraged.
The Methods section of the journal is handled by Deputy Editor Thomas Lengauer. Thomas Lengauer, Ruth Nussinov |
PLoS Comput. Biol. | 2 |
| 2013 | How Can PLOS Computational Biology Help the Biological Sciences?abstractA year has passed since I took over as the Editor-in-Chief of PLOS Computational Biology, making it a good time to reflect on the journal, reassess the needs of our community, and, broadly, the direction of computational biology within the framework of the biological sciences. The rapid growth in computational power, the masses of data that need to be understood, and the advances in the biological sciences all point to a need to reevaluate directions, merging these with the fast pace of discoveries made by experimental studies. Are we, as computational biologists, contributing as much as we can to the advancement of the sciences in key areas? Since PLOS Computational Biology is a broad, community-based open-access journal, the areas of focus that we deem central to the field may impact the future of computational biology.
The biological sciences increasingly shift to research projects that aim to understand the causes and the mechanisms of diseases and to discover therapeutics. This shift over recent years is fueled both by the desire to alleviate human suffering and by the preferences of funding agencies. Within this framework, computational biology can effectively contribute to the understanding of a broad range of biological processes under normal physiological conditions and in disease; and it can do this across a range of scales, from the molecular and biochemical to the organismal and population levels. As a leading journal in the computational sciences, PLOS Computational Biology can, and, I believe, should, foster a highly stimulating and interactive environment among computational and experimental scientists in our community toward these aims.
Computational biology increasingly gains the scientific center stage. It is an exciting interdisciplinary field that draws scientists from different fields: physics, chemistry, mathematics, engineering, computer science, and biology. Computational biology aims to organize and make sense of huge amounts of data and large-scale cellular processes and pathways, and, at the same time, to understand biological phenomena on the atomic scale. Experiments obtain information at multiple levels. Computational biology is a quantitative field. The goals of computational biology are to distinguish between noise and signals, obtain and quantify trends, and put these together, so that we are able to figure out how the information flows, how the processes are regulated, and what goes wrong in disease. We would like to understand the mechanisms through which mutations lead to cell proliferation, how the signals transmit in the cellular network and how external stimuli translate to cellular differentiation and to turning genes on and off, and how viruses enter cells. However, beyond these, computational biology aims to use the information to make predictions and obtain experimentally testable models.
PLOS Computational Biology embraces all areas of computational biology. However, I believe that to help the advancement of the biological sciences, the journal should consider papers that address biologically relevant questions that are of broad interest and provide new concepts that can guide the design of experiments. This has been in our scope since the inception of the journal, and we shall continue to follow these guidelines. It is emboldening to think of open questions in the biological sciences, particularly those relating to diseases, where computational biology can drive progress.
PLOS Computational Biology is highly regarded by the scientific community. We should aim to not only retain this appreciation but to further it. In response to requests by our community, the first step that I took after taking over as the Editor-in-Chief was to introduce a Methods section, led by Thomas Lengauer. Under Thomas' leadership, this section is flourishing, and is fast becoming highly popular, with an increasing number of submissions. The section strives to publish only the most outstanding methods, aiming to consider only those that are of exceptional quality. We are now also actively engaged in reducing the response times to authors, publication processing, and in simplifying the submission process.
The content of PLOS Computational Biology is important not only for scientific advancement; it also bears on our students who should be exposed to a variety of ways to define, analyze, question, and solve relevant scientific research problems computationally, relying on experimental data. For PLOS Computational Biology to help and inspire the biological sciences, we, as a community, should remain tuned-in to research trends and new data, and take action. These goals are shared with the International Society for Computational Biology (ISCB), with whom we are proud to partner. Together, we hope to make a difference.
Finally, the achievements of PLOS Computational Biology would not have been possible had it not been for the devoted journal staff, who are always there to help, guide, suggest, and follow on initiatives, and our extended Editorial Board. Armed with these, and embracing our community, our open-access PLOS Computational Biology journal will continue to thrive. Ruth Nussinov |
PLoS Comput. Biol. | 1 |
| 2012 | A Review of 2011 for PLoS Computational BiologyabstractIn 2011, during discussions at various conferences, as well as informally with authors, readers, reviewers, and editors, we were struck by one resonating theme: the view that PLoS Computational Biology has helped to create a sense of community amongst a broad group of scientists and educators. While the journal labels itself as a PLoS “community” journal, if that label has any true meaning, it must come from the community itself. We feel that, after six years, we are indeed serving the community well, but as always you can disagree at any time, either publicly with a comment in response to this article or by email (gro.solp@loibpmocsolp).
That service comes first and foremost from the research we publish, but also from our desire to educate, report on open-source software, provide a history of the field, capture the vision of our editors, and move beyond the boundaries of traditional publishing to inform people within and outside of our community. Before we take a look at developments in each of these areas, and what is to come in 2012, let us first review how we served the community in 2011.
According to Google Analytics, 2011 saw over 553,000 unique visitors to our website and more than two million article views (not including access statistics from PubMed Central). Visitors came from 211 countries/territories, which was undoubtedly helped by the fact that the journal is open access. India, Spain, Russia, and Iran each showed over a 40% increase in visitors from the previous year. From the journal website, the most accessed Research Article was “Effect of Promoter Architecture on the Cell-to-Cell Variability in Gene Expression” by Sanchez et al. [1], published in March 2011 (8,954 views at the time of writing); the most accessed article overall was “Ten Simple Rules for Building and Maintaining a Scientific Reputation” by Bourne and Barbour [2], published in June 2011 (15,255 views at the time of writing).
Also in 2011, 1,623 research articles from 57 countries were submitted, up 16% from 2010, and 384 were published (down 2% from 2010). Receiving more but publishing about the same number in real terms should reflect the increasing quality of our content. We are very grateful to our Associate Editors, Guest Editors, reviewers (a list of Guest Editors and reviewers from 2011 is available in Table S1), and, of course, our Deputy Editors – Patricia Babbitt, Joel Bader, Sebastian Bonhoeffer, Lyle J. Graham, Konrad Kording, Douglas Lauffenburger, Uwe Ohler, Nathan Price, Burkhard Rost, Olaf Sporns, Wyeth Wasserman, and Weixiong Zhang – for helping us to handle this growth. With this growth, we have not met our goal of reducing the times to first decision, even with the addition of new editors, but we will continue to work on this in 2012. Our median decision before review time in 2011 was 8 days, and our median decision after review time was 47 days.
A number of our Research Articles were featured in blogs and the popular press. Notably, Mitra Hartman's paper on the morphology of the rat vibrissal array [3] was covered extensively, including two videos, by National Public Radio and Science Bytes.
Our Software section was launched in August 2011, and we have so far published one article, with six more either accepted or under review. Uptake has been relatively slow, based on, we believe, the open source and stringent documentation requirements we have imposed. We believe it is better to publish only a few, but high-quality, software articles, and that this will highlight the lack of rigor of software otherwise in the field.
Our Education section has continued to flourish, in part because of the journal's relationship with the International Society for Computational Biology (ISCB). This year we introduced a collection, Bioinformatics: Starting Early, which takes the notion of biology as a computational science into secondary schools. We are hoping for more articles from those involved in secondary teaching in 2012. Open science removes all boundaries not only to reading the latest science, but also to contributing to that science. We have even seen secondary school students as authors and expect to see more in the future.
In July 2011 we began the Editors Outlook series, with five published [4]–[8] and more on the way. These mini-reviews already broach subjects from ontologies to genome organization, and from evolution to data and privacy. They speak to the breadth of our field and editorial board, and collectively will form a vision from our many expert editors of what is being, and will be, accomplished in the coming years.
That our journal is fully open access provides opportunities for maximizing the use and reuse of our scholarship; we intend to explore this further in 2012. Early in 2012 we will launch our first Topic Page on circular permutations in proteins. Wikipedia is a valuable resource for knowledge dissemination, yet Wikipedia pages are lacking in coverage of computational biology. In part this is because authors gain little career-based reward for creating Wikipedia pages. We aim to bridge the gap. Topic Page articles, which will be published in the journal and will each receive a PubMed identifier and DOI, will become the copy of record, thereby crediting the author(s). At the same time the Topic Page will be used to seed a Wikipedia article and become a living version of the same material–a viable option thanks to our Creative Commons license. Look for an announcement of this development in the new year, but in the interim if you have ideas for Topic Pages you would like to contribute, please do get in touch for further information (gro.solp@loibpmocsolp).
We are also contemplating a new article type: Data Pages. Data Pages would be brief publications about datasets, in which the data are not already well described in other papers yet are considered of great value to the community. Such brief publications would bring a traditional reward to the producers of these shared datasets. Which is more valuable: a dataset downloaded and used by 100 investigators, who in turn publish research based on these data, or a paper that is cited only by the authors who wrote it? Data Pages would, from our point of view, help to answer this question.
If you want to provide feedback on our plans for Data Pages later in 2012, please do so by commenting on this article. Feel free to comment in public or to us privately on anything we are doing, or ideas that you have for the future of the journal. After all, PLoS Computational Biology is a community journal, and if you have read this far, you should consider yourself an important part of our ever-broadening community. Rosemary Dickin, Chris James Hall, Laura K. Taylor, Andrew M. Collings, Ruth Nussinov, Philip E. Bourne |
PLoS Comput. Biol. | 5 |
| 2012 | Conformational Control of the Binding of the Transactivation Domain of the MLL Protein and c-Myb to the KIX Domain of CREBabstractThe KIX domain of CBP is a transcriptional coactivator. Concomitant binding to the activation domain of proto-oncogene protein c-Myb and the transactivation domain of the trithorax group protein mixed lineage leukemia (MLL) transcription factor lead to the biologically active ternary MLL∶KIX∶c-Myb complex which plays a role in Pol II-mediated transcription. The binding of the activation domain of MLL to KIX enhances c-Myb binding. Here we carried out molecular dynamics (MD) simulations for the MLL∶KIX∶c-Myb ternary complex, its binary components and KIX with the goal of providing a mechanistic explanation for the experimental observations. The dynamic behavior revealed that the MLL binding site is allosterically coupled to the c-Myb binding site. MLL binding redistributes the conformational ensemble of KIX, leading to higher populations of states which favor c-Myb binding. The key element in the allosteric communication pathways is the KIX loop, which acts as a control mechanism to enhance subsequent binding events. We tested this conclusion by in silico mutations of loop residues in the KIX∶MLL complex and by comparing wild type and mutant dynamics through MD simulations. The loop assumed MLL binding conformation similar to that observed in the KIX∶c-Myb state which disfavors the allosteric network. The coupling with c-Myb binding site faded, abolishing the positive cooperativity observed in the presence of MLL. Our major conclusion is that by eliciting a loop-mediated allosteric switch between the different states following the binding events, transcriptional activation can be regulated. The KIX system presents an example how nature makes use of conformational control in higher level regulation of transcriptional activity and thus cellular events. Elif Nihal Korkmaz, Ruth Nussinov, Turkan Haliloglu |
PLoS Comput. Biol. | 2 |
| 2012 | A Future Vision for PLOS Computational BiologyabstractWith much trepidation I accepted the great honor and responsibility of becoming Editor-in-Chief of PLOS Computational Biology. I am fully aware of how hard it will be to step into the shoes of Philip Bourne, the Editor-in-Chief of the journal for the last seven years, since it was founded by him, Steven Brenner, and Michael Eisen. We are all deeply appreciative and thankful to Phil for his unique role; and we are grateful that he will be continuing his association with the journal in the future in the role of Founding Editor-in-Chief, aiding and inspiring us and the PLOS Computational Biology community. As a Deputy Editor-in-Chief, I became aware of the true family relationship in the broad PLOS organization and of the devoted and gifted editorial force so nicely fostered by Phil in PLOS Computational Biology. These played a crucial role in helping me decide to accept the invitation to become Editor-in-Chief.
I am deeply committed to the underlying principle of free public access to scientific information. In particular, what is truly unique and special about PLOS Computational Biology is that it fulfills this mission while maintaining the highest standards of scientific rigor, originality, and clear biological relevance. As Editor-in-Chief I will do my best to have PLOS Computational Biology continue these traditions.
Computational biology is often perceived as a single field; this however is not the case. Like experimental biology, computational biology is enormously broad; the only distinction from experimental biology is the means. This has disadvantages and advantages: the main disadvantage is that conclusions based on computations are often treated by biologists as less conclusive than those based on experiments; the main advantages are that computations allow researchers to analyze vast amounts of data and make testable predications, and they allow researchers to address problems that current experimental methods may not be able to treat. The high quality of papers published in PLOS Computational Biology indicates that the apparent disadvantage is not necessarily there. While they may still need further experimental and computational validation, conclusions based on rigorous computations applied to carefully assembled and curated data, which are backed up by available experimental results, can be as reliable, insightful, and biologically significant as those based on experiments.
PLOS Computational Biology is broad; it addresses diverse biological problems. We look forward to further expanding its scope through the inclusion of outstanding methods and resource papers, opening up significant new research directions while retaining and enhancing the strong scientific merit of PLOS Computational Biology publications. We further plan on improving the pace of submissions processing. PLOS Computational Biology is recognized by the community as the premier journal in computational biology. We will strive to continue in this tradition.
PLOS Computational Biology is tightly associated with the International Society of Computational Biology (ISCB). We cherish and will continue fostering this association. The Society, its meetings, and the journal all have a common goal: enhancing and promoting excellence in computational biology. Outstanding research with clear biological relevance, which leads to fundamental new insights into important biological problems, is the hallmark of future contributions by our community to biology. As the Editor-in-Chief I shall do my utmost to achieve these goals. Ruth Nussinov |
PLoS Comput. Biol. | 1 |
| 2011 | GOSSIP: a method for fast and accurate global alignment of protein structuresabstractMOTIVATION: The database of known protein structures (PDB) is increasing rapidly. This results in a growing need for methods that can cope with the vast amount of structural data. To analyze the accumulating data, it is important to have a fast tool for identifying similar structures and clustering them by structural resemblance. Several excellent tools have been developed for the comparison of protein structures. These usually address the task of local structure alignment, an important yet computationally intensive problem due to its complexity. It is difficult to use such tools for comparing a large number of structures to each other at a reasonable time. RESULTS: Here we present GOSSIP, a novel method for a global all-against-all alignment of any set of protein structures. The method detects similarities between structures down to a certain cutoff (a parameter of the program), hence allowing it to detect similar structures at a much higher speed than local structure alignment methods. GOSSIP compares many structures in times which are several orders of magnitude faster than well-known available structure alignment servers, and it is also faster than a database scanning method. We evaluate GOSSIP both on a dataset of short structural fragments and on two large sequence-diverse structural benchmarks. Our conclusions are that for a threshold of 0.6 and above, the speed of GOSSIP is obtained with no compromise of the accuracy of the alignments or of the number of detected global similarities. AVAILABILITY: A server, as well as an executable for download, are available at http://bioinfo3d.cs.tau.ac.il/gossip/. Ilona Kifer, Ruth Nussinov, Haim J. Wolfson |
Bioinform. | 2 |
| 2011 | A Formal MIM Specification and Tools for the Common Exchange of MIM Diagrams: an XML-Based Format, an API, and a Validation MethodabstractBACKGROUND: The Molecular Interaction Map (MIM) notation offers a standard set of symbols and rules on their usage for the depiction of cellular signaling network diagrams. Such diagrams are essential for disseminating biological information in a concise manner. A lack of software tools for the notation restricts wider usage of the notation. Development of software is facilitated by a more detailed specification regarding software requirements than has previously existed for the MIM notation. RESULTS: A formal implementation of the MIM notation was developed based on a core set of previously defined glyphs. This implementation provides a detailed specification of the properties of the elements of the MIM notation. Building upon this specification, a machine-readable format is provided as a standardized mechanism for the storage and exchange of MIM diagrams. This new format is accompanied by a Java-based application programming interface to help software developers to integrate MIM support into software projects. A validation mechanism is also provided to determine whether MIM datasets are in accordance with syntax rules provided by the new specification. CONCLUSIONS: The work presented here provides key foundational components to promote software development for the MIM notation. These components will speed up the development of interoperable tools supporting the MIM notation and will aid in the translation of data stored in MIM diagrams to other standardized formats. Several projects utilizing this implementation of the notation are outlined herein. The MIM specification is available as an additional file to this publication. Source code, libraries, documentation, and examples are available at http://discover.nci.nih.gov/mim. Augustin Luna, Evrim I. Karac, Margot Sunshine, Lucas Chang, Ruth Nussinov, Mirit I. Aladjem, Kurt W. Kohn |
BMC Bioinform. | 5 |
| 2011 | A Review of 2010 for PLoS Computational BiologyabstractComputational Biology celebrated its fifth anniversary in 2010, and all in our community, either as readers, authors, or editors, should take pride in what has been accomplished in such a short space of time.In the past year we received 1,403 new Research Articles, a 295% increase from our first year of operation in 2005-2006 and a 17% increase over 2009.Of the articles submitted in 2010, 875 (62%) were rejected, and 70% of these were before review.We have seen growth not only in submissions, but in readership as well.Currently, around 16,000 readers receive the electronic table of contents, a 14% increase over the previous year.We published 392 Research Articles this year, along with 23 ''front section'' articles (Reviews, Perspectives, Education), down from 33 in the previous year.Eighty Associate Editors handled the combined submissions, with a total of 26 new editors joining this past year and six departing.We are proud to say that virtually every editor we asked to join accepted, a testament to how our community values the journal.These editors worked with more than 180 guest editors and 1,800 reviewers to handle the submissions, and we are of course very grateful for their support (Table S1). Rosemary Dickin, Cecy Marden, Andrew M. Collings, Ruth Nussinov, Philip E. Bourne |
PLoS Comput. Biol. | 4 |
| 2011 | The Role of Response Elements Organization in Transcription Factor Selectivity: The IFN-β Enhanceosome ExampleabstractWhat is the mechanism through which transcription factors (TFs) assemble specifically along the enhancer DNA? The IFN-β enhanceosome provides a good model system: it is small; its components' crystal structures are available; and there are biochemical and cellular data. In the IFN-β enhanceosome, there are few protein-protein interactions even though consecutive DNA response elements (REs) overlap. Our molecular dynamics (MD) simulations on different motif combinations from the enhanceosome illustrate that cooperativity is achieved via unique organization of the REs: specific binding of one TF can enhance the binding of another TF to a neighboring RE and restrict others, through overlap of REs; the order of the REs can determine which complexes will form; and the alternation of consensus and non-consensus REs can regulate binding specificity by optimizing the interactions among partners. Our observations offer an explanation of how specificity and cooperativity can be attained despite the limited interactions between neighboring TFs on the enhancer DNA. To date, when addressing selective TF binding, attention has largely focused on RE sequences. Yet, the order of the REs on the DNA and the length of the spacers between them can be a key factor in specific combinatorial assembly of the TFs on the enhancer and thus in function. Our results emphasize cooperativity via RE binding sites organization. Ruth Nussinov |
PLoS Comput. Biol. | 2 |
| 2010 | Lysine120 Interactions with p53 Response Elements can Allosterically Direct p53 Organizationabstractp53 can serve as a paradigm in studies aiming to figure out how allosteric perturbations in transcription factors (TFs) triggered by small changes in DNA response element (RE) sequences, can spell selectivity in co-factor recruitment. p53-REs are 20-base pair (bp) DNA segments specifying diverse functions. They may be located near the transcription start sites or thousands of bps away in the genome. Their number has been estimated to be in the thousands, and they all share a common motif. A key question is then how does the p53 protein recognize a particular p53-RE sequence among all the similar ones? Here, representative p53-REs regulating diverse functions including cell cycle arrest, DNA repair, and apoptosis were simulated in explicit solvent. Among the major interactions between p53 and its REs involving Lys120, Arg280 and Arg248, the bps interacting with Lys120 vary while the interacting partners of other residues are less so. We observe that each p53-RE quarter site sequence has a unique pattern of interactions with p53 Lys120. The allosteric, DNA sequence-induced conformational and dynamic changes of the altered Lys120 interactions are amplified by the perturbation of other p53-DNA interactions. The combined subtle RE sequence-specific allosteric effects propagate in the p53 and in the DNA. The resulting amplified allosteric effects far away are reflected in changes in the overall p53 organization and in the p53 surface topology and residue fluctuations which play key roles in selective co-factor recruitment. As such, these observations suggest how similar p53-RE sequences can spell the preferred co-factor binding, which is the key to the selective gene transactivation and consequently different functional effects. Ruth Nussinov |
PLoS Comput. Biol. | 2 |
| 2010 | A Mechanistic View of the Role of E3 in SumoylationabstractSumoylation, the covalent attachment of SUMO (Small Ubiquitin-Like Modifier) to proteins, differs from other Ubl (Ubiquitin-like) pathways. In sumoylation, E2 ligase Ubc9 can function without E3 enzymes, albeit with lower reaction efficiency. Here, we study the mechanism through which E3 ligase RanBP2 triggers target recognition and catalysis by E2 Ubc9. Two mechanisms were proposed for sumoylation. While in both the first step involves Ubc9 conjugation to SUMO, the subsequent sequence of events differs: in the first E2-SUMO forms a complex with the target and E3, followed by SUMO transfer to the target. In the second, Ubc9-SUMO binds to the target and facilitates SUMO transfer without E3. Using dynamic correlations obtained from explicit solvent molecular dynamic simulations we illustrate the key roles played by allostery in both mechanisms. Pre-existence of conformational states explains the experimental observations that sumoylation can occur without E3, even though at a reduced rate. Furthermore, we propose a mechanism for enhancement of sumoylation by E3. Analysis of the conformational ensembles of the complex of E2 conjugated to SUMO illustrates that the E2 enzyme is already largely pre-organized for target binding and catalysis; E3 binding shifts the equilibrium and enhances these pre-existing populations. We further observe that E3 binding regulates allosterically the key residues in E2, Ubc9 Asp100/Lys101 E2, for the target recognition. Melda Tozluoglu, Ezgi Karaca, Ruth Nussinov, Turkan Haliloglu |
PLoS Comput. Biol. | 3 |
| 2009 | A survey of available tools and web servers for analysis of protein-protein interactions and interfacesabstractThe unanimous agreement that cellular processes are (largely) governed by interactions between proteins has led to enormous community efforts culminating in overwhelming information relating to these proteins; to the regulation of their interactions, to the way in which they interact and to the function which is determined by these interactions. These data have been organized in databases and servers. However, to make these really useful, it is essential not only to be aware of these, but in particular to have a working knowledge of which tools to use for a given problem; what are the tool advantages and drawbacks; and no less important how to combine these for a particular goal since usually it is not one tool, but some combination of tool-modules that is needed. This is the goal of this review. Nurcan Tuncbag, Gozde Kar, Ozlem Keskin, Attila Gürsoy, Ruth Nussinov |
Briefings Bioinform. | 5 |
| 2009 | The Mechanism of Ubiquitination in the Cullin-RING E3 Ligase Machinery: Conformational Control of Substrate OrientationabstractIn cullin-RING E3 ubiquitin ligases, substrate binding proteins, such as VHL-box, SOCS-box or the F-box proteins, recruit substrates for ubiquitination, accurately positioning and orienting the substrates for ubiquitin transfer. Yet, how the E3 machinery precisely positions the substrate is unknown. Here, we simulated nine substrate binding proteins: Skp2, Fbw7, beta-TrCP1, Cdc4, Fbs1, TIR1, pVHL, SOCS2, and SOCS4, in the unbound form and bound to Skp1, ASK1 or Elongin C. All nine proteins have two domains: one binds to the substrate; the other to E3 ligase modules Skp1/ASK1/Elongin C. We discovered that in all cases the flexible inter-domain linker serves as a hinge, rotating the substrate binding domain, optimally and accurately positioning it for ubiquitin transfer. We observed a conserved proline in the linker of all nine proteins. In all cases, the prolines pucker substantially and the pucker is associated with the backbone rotation toward the E2/ubiquitin. We further observed that the linker flexibility could be regulated allosterically by binding events associated with either domain. We conclude that the flexible linker in the substrate binding proteins orients the substrate for the ubiquitin transfer. Our findings provide a mechanism for ubiquitination and polyubiquitination, illustrating that these processes are under conformational control. Jin Liu 0005, Ruth Nussinov |
PLoS Comput. Biol. | 2 |
| 2009 | Cooperativity Dominates the Genomic Organization of p53-Response Elements: A Mechanistic Viewabstractp53-response elements (p53-REs) are organized as two repeats of a palindromic DNA segment spaced by 0 to 20 base pairs (bp). Several experiments indicate that in the vast majority of the human p53-REs there are no spacers between the two repeats; those with spacers, particularly with sizes beyond two nucleotides, are rare. This raises the question of what it indicates about the factors determining the p53-RE genomic organization. Clearly, given the double helical DNA conformation, the orientation of two p53 core domain dimers with respect to each other will vary depending on the spacer size: a small spacer of 0 to 2 bps will lead to the closest p53 dimer-dimer orientation; a 10-bp spacer will locate the p53 dimers on the same DNA face but necessitate DNA looping; while a 5-bp spacer will position the p53 dimers on opposite DNA faces. Here, via conformational analysis we show that when there are 0-2 bp spacers, p53-DNA binding is cooperative; however, cooperativity is greatly diminished when there are spacers with sizes beyond 2 bp. Cooperative binding is broadly recognized to be crucial for biological processes, including transcriptional regulation. Our results clearly indicate that cooperativity of the p53-DNA association dominates the genomic organization of the p53-REs, raising questions of the structural organization and functional roles of p53-REs with larger spacers. We further propose that a dynamic landscape scenario of p53 and p53-REs can better explain the selectivity of the degenerate p53-REs. Our conclusions bear on the evolutionary preference of the p53-RE organization and as such, are expected to have broad implications to other multimeric transcription factor response element organization. Ruth Nussinov |
PLoS Comput. Biol. | 2 |
| 2008 | Bio-geometry: challenges, approaches, and future opportunities in proteomics and drug discoveryabstractBiology has been an experimental science until the recent prominence of Bioinformatics and Computational Biology. With the discovery of the DNA sequence the protein structure determination is now an emerging challenge. The protein structure is closely coupled to the function. Today, with the given increase in available computing power and no-cost storage, the ability to do computational experiments is emerging as a core competence necessary for rapid discovery in the future. The ability to include various complex physics such as electrostatics and hydrophobic interactions in realistic simulations has increased. The discovery of structure of proteins is the next frontier for a number of convergent areas in science. Hence the combination of geometry and physics becomes very critical to do realistic computational experiments. Ruth Nussinov, Talapady Bhat, Jack Snoeyink, Karthik Ramani |
Symposium on Solid and Physical Modeling | 1 |
| 2007 | Deterministic Pharmacophore Detection Via Multiple Flexible Alignment of Drug-Like Molecules
Yuval Inbar, Dina Schneidman, Oranit Dror, Ruth Nussinov, Haim J. Wolfson |
RECOMB | 4 |
| 2007 | Ligand Binding and Circular Permutation Modify Residue Interaction Network in DHFRabstractResidue interaction networks and loop motions are important for catalysis in dihydrofolate reductase (DHFR). Here, we investigate the effects of ligand binding and chain connectivity on network communication in DHFR. We carry out systematic network analysis and molecular dynamics simulations of the native DHFR and 19 of its circularly permuted variants by breaking the chain connections in ten folding element regions and in nine nonfolding element regions as observed by experiment. Our studies suggest that chain cleavage in folding element areas may deactivate DHFR due to large perturbations in the network properties near the active site. The protein active site is near or coincides with residues through which the shortest paths in the residue interaction network tend to go. Further, our network analysis reveals that ligand binding has "network-bridging effects" on the DHFR structure. Our results suggest that ligand binding leads to a modification, with most of the interaction networks now passing through the cofactor, shortening the average shortest path. Ligand binding at the active site has profound effects on the network centrality, especially the closeness. Zengjian Hu, Donnell Bowen, William M. Southerland, Antonio del Sol, Ruth Nussinov, Buyong Ma |
PLoS Comput. Biol. | 6 |
| 2007 | EMatch: Discovery of High Resolution Structural Homologues of Protein Domains in Intermediate Resolution Cryo-EM MapsabstractCryo-EM has become an increasingly powerful technique for elucidating the structure, dynamics, and function of large flexible macromolecule assemblies that cannot be determined at atomic resolution. However, due to the relatively low resolution of cryo-EM data, a major challenge is to identify components of complexes appearing in cryo-EM maps. Here, we describe EMatch, a novel integrated approach for recognizing structural homologues of protein domains present in a 6-10 A resolution cryo-EM map and constructing a quasi-atomic structural model of their assembly. The method is highly efficient and has been successfully validated on various simulated data. The strength of the method is demonstrated by a domain assembly of an experimental cryo-EM map of native GroEL at 6 A resolution. Keren Lasker, Oranit Dror, Maxim Shatsky, Ruth Nussinov, Haim J. Wolfson |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2006 | A permissive secondary structure-guided superposition tool for clustering of protein fragments toward protein structure prediction via fragment assemblyabstractMOTIVATION: Secondary-Structure Guided Superposition tool (SSGS) is a permissive secondary structure-based algorithm for matching of protein structures and in particular their fragments. The algorithm was developed towards protein structure prediction via fragment assembly. RESULTS: In a fragment-based structural prediction scheme, a protein sequence is cut into building blocks (BBs). The BBs are assembled to predict their relative 3D arrangement. Finally, the assemblies are refined. To implement this prediction scheme, a clustered structural library representing sequence patterns for protein fragments is essential. To create a library, BBs generated by cutting proteins from the PDB are compared and structurally similar BBs are clustered. To allow structural comparison and clustering of the BBs, which are often relatively short with flexible loops, we have devised SSGS. SSGS maintains high similarity between cluster members and is highly efficient. When it comes to comparing BBs for clustering purposes, the algorithm obtains better results than other, non-secondary structure guided protein superimposition algorithms. Gilad Wainreb, Nurit Haspel, Haim J. Wolfson, Ruth Nussinov |
Bioinform. | 4 |
| 2006 | Designing a Nanotube Using Naturally Occurring Protein Building BlocksabstractHere our goal is to carry out nanotube design using naturally occurring protein building blocks. Inspection of the protein structural database reveals the richness of the conformations of proteins, their parts, and their chemistry. Given target functional protein nanotube geometry, our strategy involves scanning a library of candidate building blocks, combinatorially assembling them into the shape and testing its stability. Since self-assembly takes place on time scales not affordable for computations, here we propose a strategy for the very first step in protein nanotube design: we map the candidate building blocks onto a planar sheet and wrap the sheet around a cylinder with the target dimensions. We provide examples of three nanotubes, two peptide and one protein, in atomistic model detail for which there are experimental data. The nanotube models can be used to verify a nanostructure observed by low-resolution experiments, and to study the mechanism of tube formation. Chung-Jung Tsai, Jie Zheng 0003, Ruth Nussinov |
PLoS Comput. Biol. | 3 |
| 2005 | Recognition of Binding Patterns Common to a Set of Protein Structures
Maxim Shatsky, Alexandra Shulman-Peleg, Ruth Nussinov, Haim J. Wolfson |
RECOMB | 3 |
| 2005 | Discovery of Protein Substructures in EM Maps
Keren Lasker, Oranit Dror, Ruth Nussinov, Haim J. Wolfson |
WABI | 3 |
| 2004 | Protein-Protein Interfaces: Recognition of Similar Spatial and Chemical Organizations
Alexandra Shulman-Peleg, Shira Mintz, Ruth Nussinov, Haim J. Wolfson |
WABI | 3 |
| 2002 | Efficient Unbound Docking of Rigid Molecules
Dina Schneidman, Ruth Nussinov, Haim J. Wolfson |
WABI | 2 |
| 2002 | MultiProt - A Multiple Protein Structural Alignment Algorithm
Maxim Shatsky, Ruth Nussinov, Haim J. Wolfson |
WABI | 2 |
| 2000 | Alignment of Flexible Protein Structures
Maxim Shatsky, Zipora Y. Fligelman, Ruth Nussinov, Haim J. Wolfson |
ISMB | 3 |
| 1999 | Multiple Structural Alignment and Core Detection by Geometric Hashing
Nathaniel Leibowitz, Zipora Y. Fligelman, Ruth Nussinov, Haim J. Wolfson |
ISMB | 3 |
| 1996 | Docking of Conformationally Flexible Proteins
Bilha Sandak, Ruth Nussinov, Haim J. Wolfson |
CPM | 2 |
| 1995 | An automated computer vision and robotics-based technique for 3-D flexible biomolecular docking and matchingabstractThe generation of binding modes between two molecules, also known as molecular docking, is a key problem in rational drug design and biomolecular recognition. Docking a ligand, e.g., a drug molecule or a protein molecule, to a protein receptor, involves recognition of molecular surfaces as molecules interact at their surface. Recent studies report that the activity of many molecules induces conformational transitions by 'hinge-bending', which involves movements of relatively rigid parts with respect to each other. In ligand-receptor binding, relative rotational movements of molecular substructures about their common hinges have been observed. For automatically predicting flexible molecular interactions, we adapt a new technique developed in Computer Vision and Robotics for the efficient recognition of partially occluded articulated objects. These type of objects consist of rigid parts which are connected by rotary joints (hinges). Our approach is based on an extension and generalization of the Geometric Hashing and Generalized Hough Transform paradigm for rigid object recognition. Unlike other techniques which match each part individually, our approach exploits forcefully and efficiently enough the fact that the different rigid parts do belong to the same flexible molecule. We show experimental results obtained by an implementation of the algorithm for rigid and flexible docking. While the 'correct', crystal-bound complex is obtained with a small RMSD, additional, predictive 'high scoring' binding modes are generated as well. The diverse applications and implications of this general, powerful tool are discussed. Bilha Sandak, Ruth Nussinov, Haim J. Wolfson |
Comput. Appl. Biosci. | 2 |
| 1994 | Docking of protein moleculesabstractThe problem of receptor-ligand recognition and binding is encountered in a very large number of biological processes, The behavior of the molecules depends both on their geometric shape and on the chemical interactions among their atoms. This work addresses only the geometrical (key-in-lock) aspect of the problem, where acceptable solutions should exhibit shape complementarity. The problem we are faced with here is reminiscent of partial 3-D surface matching problems in computer vision. Here we present a new 3-D molecular surface representation by (hundreds of) sparse interest (critical) points with associated normals and a subsequent matching approach which is based on the geometric hashing paradigm originally developed for computer vision motivated object recognition applications. Potential solutions resulting in the interpenetration of the molecules are discarded by a subsequent verification procedure. Numerous examples of the successful geometric prediction of our technique are presented. In all cases the performance of our algorithm has been by several orders of magnitude faster than of other state-of-the-art docking algorithms. Daniel Fischer 0001, Shuo L. Lin, Ruth Nussinov, Haim J. Wolfson |
ICPR (2) | 3 |
| 1993 | 3-D Docking of Protein Molecules
Daniel Fischer 0001, Raquel Norel, Ruth Nussinov, Haim J. Wolfson |
CPM | 3 |
| 1992 | 3-D Substructure Matching in Protein Molecules
Daniel Fischer 0001, Ruth Nussinov, Haim J. Wolfson |
CPM | 2 |
| 1991 | Compositional variations in DNA sequencesabstractBiologically occurring nucleotide sequences differ from randomly generated ones. Here we describe general patterns found in prokaryotic and in eukaryotic DNA. In the accompanying paper (Nussinov, 1991) we also describe DNA signals recognized by their corresponding protein factors. In particular, we focus on modes of searches for such patterns and signals and on the potential properties such sequences may possess. Ruth Nussinov |
Comput. Appl. Biosci. | 1 |
| 1991 | Signals in DNA sequences and their potential propertiesabstractDNA and RNA molecules contain signals which are recognized by regulatory proteins or enzymes either directly, through their nucleotide sequences or indirectly, through induced structural changes on their neighboring sequences. To date, most signal searches have been focused on specific recurrences of nucleotide sequences. Much less attention has been directed towards the structure, flexibility and hydrogen-bonding patterns that recognition elements may possess. Here we review the various methods involved in such searches. In particular, however, we also address the searches for potential properties. In this regard it is of interest to inspect the asymmetry in the distributions of complementary oligomers near biological features. Upstream of transcription initiation the frequencies of G-rich oligomers are particularly high (Nussinov, 1987a; 1990). The frequencies of C-rich oligomers are lower. A-tracts are also very frequent in these regions. This may correlate with the recent finding that guanine, but not cytosine, tracts enhance A-tract directed bends (Milton et al., 1990). Presumably A-tracts near G-tracts on the same strand may induce a structural change in the G-tracts which may enhance the bend. G-tracts may have the potential for participating in a DNA bend due to their compression of the major groove. Thus, proteins may not always be necessary to induce DNA conformational changes. This example illustrates the importance of studies of the properties of DNA oligomers in regulatory regions, and of algorithms for their detection. Ruth Nussinov |
Comput. Appl. Biosci. | 1 |
| 1989 | RNA secondary structures: comparison and determination of frequently recurring substructures by consensusabstractA method for assessing the preserved stem-loops of RNA secondary structures is presented. Frequently recurring helical stems in a set of secondary structures resulting from the simulated folding process of a given RNA are assessed and consensus structural motifs can then be selected to construct a secondary structure of the RNA. Alternatively, it can be applied to a series of 'optimal' and 'suboptimal' secondary structures computed using the dynamic program developed by Williams and Tinoco. To demonstrate the power and the usefulness of the program we give examples of this procedure. Shu-Yun Le, John Owens, Ruth Nussinov, Jih-Hsiang Chen, Bruce A. Shapiro, Jacob V. Maizel Jr. |
Comput. Appl. Biosci. | 3 |
| 1988 | Locating alignments with k differences for nucleotide and amino acid sequencesabstractGiven two sequences, a pattern of length m, a text of length n and a positive integer k, we give two algorithms. The first finds all occurrences of the pattern in the text as long as these do not differ from each other by more than k differences. It runs in O(nk) time. The second algorithm finds all subsequence alignments between the pattern and the test with at most k differences. This algorithm runs in O(nmk) time, is very simple and easy to program. Gad M. Landau, Uzi Vishkin, Ruth Nussinov |
Comput. Appl. Biosci. | 3 |
| 1988 | An improved secondary structure computation method and its application to intervening sequence in the human alpha-likeglobin mRNA precursorsabstractCurrent secondary structure prediction computations have a serious drawback. The calculated thermodynamically most stable structure often differs from that observed in solution or in crystal form. In this paper we suggest a way to partially over-come some of these limitations by simulating the RNA folding process and calculating the frequencies of occurrence of the various substructures obtained. The frequently recurring substructures are then selected to construct the secondary structure of the whole RNA. 142 tRNA molecules and an E. coli 16S rRNA molecule have been examined by this method. The percentage of successful prediction of the correct helices are significantly higher than those calculated previously. The secondary structures of intervening sequences (IVSs) excised from human alpha-like globin pre-mRNAs are also computed. Thus, in this method the secondary structures obtained are composed of the statistically more significant substructures. This has also been demonstrated by using randomly shuffled sequences. The secondary structures of each of the randomized sequences are computed and their mean and standard deviations are used in evaluating the significance of the substructures obtained in the folding of the biological sequence. Some potentially appealing structural features aligning adjacent exons for ligation have been found. Shu-Yun Le, Jih-Hsiang Chen, Ruth Nussinov, Jacob V. Maizel Jr. |
Comput. Appl. Biosci. | 3 |
| 1988 | A fixed-point alignment technique for detection of recurrent and common sequence motifs associated with biological featuresabstractA fixed-point alignment analysis technique is presented which is designed to locate common sequence motifs in collections of proteins or nucleic acids. Initially a program aligns a collection of sequences by a common sequence pattern or known biological feature. The common pattern or feature (fixed-point) may be a user-specified sequence string or a known sequence position like mRNA start site, which may be taken directly from the annotated feature table of GenBank. Once all alignment markers are located, the sequences are scanned for occurrences of given oligomers within a specified span both upstream and downstream of the fixed-point. The occurrences may then be plotted as a function of the position relative to the fixed-point, displayed as an actual sequence alignment or selectively summarized via various program options. Applications of the technique are discussed. John Owens, Devjani Chatterjee, Ruth Nussinov, Andrzej K. Konopka, Jacob V. Maizel Jr. |
Comput. Appl. Biosci. | 3 |