Ivet Bahar

dblp:57/2788 · DBLP profile ↗
← Back
36ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0001-9959-4176ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 36 · 6 since 2021
YearPublicationVenuePosition
2026 DruGUI 2.0: mapping protein druggability with probe-based molecular dynamics
abstract
SUMMARY: We introduce DruGUI 2.0, a drug discovery tool for assessing the druggability of proteins, integrated into the ProDy application programming interface (API). DruGUI 2.0 is developed to facilitate the search for druggable sites while allowing for proteins' conformational flexibility. Simulations in explicit solvent, with an option to include membrane, are carried out in the presence of probe molecules selected from an expanded library of small molecules containing drug-like fragments. Druggable sites beyond orthosteric sites are identifiable, as well as the probes that show high affinity to bind to those sites. Characterization of the composition and position of the probes helps build pharmacophore models and estimate relative binding affinities. As a Python module with enhanced visualization features, DruGUI 2.0 complements, and benefits from, the vast collection of protein sequence, structure, and dynamics analyses modules accessible in ProDy. Case studies in the Supplemental Material showcase the utility of DruGUI 2.0 applied to both soluble targets and membrane proteins. AVAILABILITY: ProDy is open-sourced and freely available under MIT License from https://github.com/prody/ProDy. The code version of DruGUI 2.0 used for simulations is available on Zenodo : 10.5281/zenodo.20511357.
Carlos Ventura, Anthony T. Bogetti, Anupam Banerjee, Matthew Licht, Ivet Bahar
Bioinform.6
2024 WatFinder: a ProDy tool for protein-water interactions
abstract
SUMMARY: We introduce WatFinder, a tool designed to identify and visualize protein-water interactions (water bridges, water-mediated associations, or water channels, fluxes, and clusters) relevant to protein stability, dynamics, and function. WatFinder is integrated into ProDy, a Python API broadly used for structure-based prediction of protein dynamics. WatFinder provides a suite of functions for generating raw data as well as outputs from statistical analyses. The ProDy framework facilitates comprehensive automation and efficient analysis of the ensembles of structures resolved for a given protein or the time-evolved conformations from simulations in explicit water, as illustrated in five case studies presented in the Supplementary Material. AVAILABILITY AND IMPLEMENTATION: ProDy is open-source and freely available under MIT License from https://github.com/ProDy/ProDy.
James Krieger, Frane Doljanin, Anthony T. Bogetti, Thiliban Manivarma, Ivet Bahar, Karolina Mikulska-Ruminska
Bioinform.6
2022 Elastic network modeling of cellular networks unveils sensor and effector genes that control information flow
abstract
The high-level organization of the cell is embedded in indirect relationships that connect distinct cellular processes. Existing computational approaches for detecting indirect relationships between genes typically consist of propagating abstract information through network representations of the cell. However, the selection of genes to serve as the source of propagation is inherently biased by prior knowledge. Here, we sought to derive an unbiased view of the high-level organization of the cell by identifying the genes that propagate and receive information most effectively in the cell, and the indirect relationships between these genes. To this aim, we adapted a perturbation-response scanning strategy initially developed for identifying allosteric interactions within proteins. We deployed this strategy onto an elastic network model of the yeast genetic interaction profile similarity network. This network revealed a superior propensity for information propagation relative to simulated networks with similar topology. Perturbation-response scanning identified the major distributors and receivers of information in the network, named effector and sensor genes, respectively. Effectors formed dense clusters centrally integrated into the network, whereas sensors formed loosely connected antenna-shaped clusters and contained genes with previously characterized involvement in signal transduction. We propose that indirect relationships between effector and sensor clusters represent major paths of information flow between distinct cellular processes. Genetic similarity networks for fission yeast and human displayed similarly strong propensities for information propagation and clusters of effector and sensor genes, suggesting that the global architecture enabling indirect relationships is evolutionarily conserved across species. Our results demonstrate that elastic network modeling of cellular networks constitutes a promising strategy to probe the high-level organization and cooperativity in the cell.
Omer Acar, She Zhang, Ivet Bahar, Anne-Ruxandra Carvunis
PLoS Comput. Biol.3
2021 ClustENMD: efficient sampling of biomolecular conformational space at atomic resolution
abstract
SUMMARY: Efficient sampling of conformational space is essential for elucidating functional/allosteric mechanisms of proteins and generating ensembles of conformers for docking applications. However, unbiased sampling is still a challenge especially for highly flexible and/or large systems. To address this challenge, we describe a new implementation of our computationally efficient algorithm ClustENMD that is integrated with ProDy and OpenMM softwares. This hybrid method performs iterative cycles of conformer generation using elastic network model for deformations along global modes, followed by clustering and short molecular dynamics simulations. ProDy framework enables full automation and analysis of generated conformers and visualization of their distributions in the essential subspace. AVAILABILITY AND IMPLEMENTATION: ClustENMD is open-source and freely available under MIT License from https://github.com/prody/ProDy. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Burak T. Kaynak, She Zhang, Ivet Bahar, Pemra Doruker
Bioinform.3
2021 ProDy 2.0: increased scale and scope after 10 years of protein dynamics modelling with Python
abstract
SUMMARY: ProDy, an integrated application programming interface developed for modelling and analysing protein dynamics, has significantly evolved in recent years in response to the growing data and needs of the computational biology community. We present major developments that led to ProDy 2.0: (i) improved interfacing with databases and parsing new file formats, (ii) SignDy for signature dynamics of protein families, (iii) CryoDy for collective dynamics of supramolecular systems using cryo-EM density maps and (iv) essential site scanning analysis for identifying sites essential to modulating global dynamics. AVAILABILITY AND IMPLEMENTATION: ProDy is open-source and freely available under MIT License from https://github.com/prody/ProDy. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
She Zhang, James Krieger, Cihan Kaya, Burak T. Kaynak, Karolina Mikulska-Ruminska, Pemra Doruker, Hongchun Li, Ivet Bahar
Bioinform.9
2021 Coupled mixed model for joint genetic analysis of complex disorders with two independently collected data sets
abstract
BACKGROUND: In the last decade, Genome-wide Association studies (GWASs) have contributed to decoding the human genome by uncovering many genetic variations associated with various diseases. Many follow-up investigations involve joint analysis of multiple independently generated GWAS data sets. While most of the computational approaches developed for joint analysis are based on summary statistics, the joint analysis based on individual-level data with consideration of confounding factors remains to be a challenge. RESULTS: In this study, we propose a method, called Coupled Mixed Model (CMM), that enables a joint GWAS analysis on two independently collected sets of GWAS data with different phenotypes. The CMM method does not require the data sets to have the same phenotypes as it aims to infer the unknown phenotypes using a set of multivariate sparse mixed models. Moreover, CMM addresses the confounding variables due to population stratification, family structures, and cryptic relatedness, as well as those arising during data collection such as batch effects that frequently appear in joint genetic studies. We evaluate the performance of CMM using simulation experiments. In real data analysis, we illustrate the utility of CMM by an application to evaluating common genetic associations for Alzheimer's disease and substance use disorder using datasets independently collected for the two complex human disorders. Comparison of the results with those from previous experiments and analyses supports the utility of our method and provides new insights into the diseases. The software is available at https://github.com/HaohanWang/CMM .
Haohan Wang, Fen Pei, Michael M. Vanyukov, Ivet Bahar, Wei Wu 0023, Eric P. Xing
BMC Bioinform.4
2020 QuartataWeb: Integrated Chemical-Protein-Pathway Mapping for Polypharmacology and Chemogenomics
abstract
SUMMARY: QuartataWeb is a user-friendly server developed for polypharmacological and chemogenomics analyses. Users can easily obtain information on experimentally verified (known) and computationally predicted (new) interactions between 5494 drugs and 2807 human proteins in DrugBank, and between 315 514 chemicals and 9457 human proteins in the STITCH database. In addition, QuartataWeb links targets to KEGG pathways and GO annotations, completing the bridge from drugs/chemicals to function via protein targets and cellular pathways. It allows users to query a series of chemicals, drug combinations or multiple targets, to enable multi-drug, multi-target, multi-pathway analyses, toward facilitating the design of polypharmacological treatments for complex diseases. AVAILABILITY AND IMPLEMENTATION: QuartataWeb is freely accessible at http://quartata.csb.pitt.edu. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hongchun Li, Fen Pei, D. Lansing Taylor, Ivet Bahar
Bioinform.4
2020 Rhapsody: predicting the pathogenicity of human missense variants
abstract
MOTIVATION: The biological effects of human missense variants have been studied experimentally for decades but predicting their effects in clinical molecular diagnostics remains challenging. Available computational tools are usually based on the analysis of sequence conservation and structural properties of the mutant protein. We recently introduced a new machine learning method that demonstrated for the first time the significance of protein dynamics in determining the pathogenicity of missense variants. RESULTS: Here, we present a new interface (Rhapsody) that enables fully automated assessment of pathogenicity, incorporating both sequence coevolution data and structure- and dynamics-based features. Benchmarked against a dataset of about 20 000 annotated variants, the methodology is shown to outperform well-established and/or advanced prediction tools. We illustrate the utility of Rhapsody by in silico saturation mutagenesis studies of human H-Ras, phosphatase and tensin homolog and thiopurine S-methyltransferase. AVAILABILITY AND IMPLEMENTATION: The new tool is available both as an online webserver at http://rhapsody.csb.pitt.edu and as an open-source Python package (GitHub repository: https://github.com/prody/rhapsody; PyPI package installation: pip install prody-rhapsody). Links to additional resources, tutorials and package documentation are provided in the 'Python package' section of the website. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Luca Ponzoni, Daniel A. Peñaherrera, Zoltán N. Oltvai, Ivet Bahar
Bioinform.4
2020 Complementary computational and experimental evaluation of missense variants in the ROMK potassium channel
abstract
The renal outer medullary potassium (ROMK) channel is essential for potassium transport in the kidney, and its dysfunction is associated with a salt-wasting disorder known as Bartter syndrome. Despite its physiological significance, we lack a mechanistic understanding of the molecular defects in ROMK underlying most Bartter syndrome-associated mutations. To this end, we employed a ROMK-dependent yeast growth assay and tested single amino acid variants selected by a series of computational tools representative of different approaches to predict each variants' pathogenicity. In one approach, we used in silico saturation mutagenesis, i.e. the scanning of all possible single amino acid substitutions at all sequence positions to estimate their impact on function, and then employed a new machine learning classifier known as Rhapsody. We also used two additional tools, EVmutation and Polyphen-2, which permitted us to make consensus predictions on the pathogenicity of single amino acid variants in ROMK. Experimental tests performed for selected mutants in different classes validated the vast majority of our predictions and provided insights into variants implicated in ROMK dysfunction. On a broader scope, our analysis suggests that consolidation of data from complementary computational approaches provides an improved and facile method to predict the severity of an amino acid substitution and may help accelerate the identification of disease-causing mutations in any protein.
Luca Ponzoni, Nga H. Nguyen, Ivet Bahar, Jeffrey L. Brodsky
PLoS Comput. Biol.3
2015 BalestraWeb: efficient online evaluation of drug-target interactions
abstract
SUMMARY: BalestraWeb is an online server that allows users to instantly make predictions about the potential occurrence of interactions between any given drug-target pair, or predict the most likely interaction partners of any drug or target listed in the DrugBank. It also permits users to identify most similar drugs or most similar targets based on their interaction patterns. Outputs help to develop hypotheses about drug repurposing as well as potential side effects. AVAILABILITY AND IMPLEMENTATION: BalestraWeb is accessible at http://balestra.csb.pitt.edu/. The tool is built using a probabilistic matrix factorization method and DrugBank v3, and the latent variable models are trained using the GraphLab collaborative filtering toolkit. The server is implemented using Python, Flask, NumPy and SciPy.
Murat Can Cobanoglu, Zoltán N. Oltvai, D. Lansing Taylor, Ivet Bahar
Bioinform.4
2015 The anisotropic network model web server at 2015 (ANM 2.0)
abstract
SUMMARY: The anisotropic network model (ANM) is one of the simplest yet powerful tools for exploring protein dynamics. Its main utility is to predict and visualize the collective motions of large complexes and assemblies near their equilibrium structures. The ANM server, introduced by us in 2006 helped making this tool more accessible to non-sophisticated users. We now provide a new version (ANM 2.0), which allows inclusion of nucleic acids and ligands in the network model and thus enables the investigation of the collective motions of protein-DNA/RNA and -ligand systems. The new version offers the flexibility of defining the system nodes and the interaction types and cutoffs. It also includes extensive improvements in hardware, software and graphical interfaces. AVAILABILITY AND IMPLEMENTATION: ANM 2.0 is available at http://anm.csb.pitt.edu CONTACT: [email protected], [email protected].
Eran Eyal, Gengkon Lum, Ivet Bahar
Bioinform.3
2015 Comparative study of the effectiveness and limitations of current methods for detecting sequence coevolution
abstract
MOTIVATION: With rapid accumulation of sequence data on several species, extracting rational and systematic information from multiple sequence alignments (MSAs) is becoming increasingly important. Currently, there is a plethora of computational methods for investigating coupled evolutionary changes in pairs of positions along the amino acid sequence, and making inferences on structure and function. Yet, the significance of coevolution signals remains to be established. Also, a large number of false positives (FPs) arise from insufficient MSA size, phylogenetic background and indirect couplings. RESULTS: Here, a set of 16 pairs of non-interacting proteins is thoroughly examined to assess the effectiveness and limitations of different methods. The analysis shows that recent computationally expensive methods designed to remove biases from indirect couplings outperform others in detecting tertiary structural contacts as well as eliminating intermolecular FPs; whereas traditional methods such as mutual information benefit from refinements such as shuffling, while being highly efficient. Computations repeated with 2,330 pairs of protein families from the Negatome database corroborated these results. Finally, using a training dataset of 162 families of proteins, we propose a combined method that outperforms existing individual methods. Overall, the study provides simple guidelines towards the choice of suitable methods and strategies based on available MSA size and computing resources. AVAILABILITY AND IMPLEMENTATION: Software is freely available through the Evol component of ProDy API.
Wenzhi Mao, Cihan Kaya, Anindita Dutta, Amnon Horovitz, Ivet Bahar
Bioinform.5
2015 The center for causal discovery of biomedical knowledge from big data
abstract
The Big Data to Knowledge (BD2K) Center for Causal Discovery is developing and disseminating an integrated set of open source tools that support causal modeling and discovery of biomedical knowledge from large and complex biomedical datasets. The Center integrates teams of biomedical and data scientists focused on the refinement of existing and the development of new constraint-based and Bayesian algorithms based on causal Bayesian networks, the optimization of software for efficient operation in a supercomputing environment, and the testing of algorithms and software developed using real data from 3 representative driving biomedical projects: cancer driver mutations, lung disease, and the functional connectome of the human brain. Associated training activities provide both biomedical and data scientists with the knowledge and skills needed to apply and extend these tools. Collaborative activities with the BD2K Consortium further advance causal discovery tools and integrate tools and resources developed by other centers.
Gregory F. Cooper, Ivet Bahar, Michael J. Becich, Panayiotis V. Benos, Jeremy M. Berg, Jeremy U. Espino, Clark Glymour, Rebecca S. Jacobson, Michelle Kienholz, Adrian V. Lee, Xinghua Lu 0001, Richard Scheines
J. Am. Medical Informatics Assoc.2
2014 Evol and ProDy for bridging protein sequence evolution and structural dynamics
abstract
UNLABELLED: Correlations between sequence evolution and structural dynamics are of utmost importance in understanding the molecular mechanisms of function and their evolution. We have integrated Evol, a new package for fast and efficient comparative analysis of evolutionary patterns and conformational dynamics, into ProDy, a computational toolbox designed for inferring protein dynamics from experimental and theoretical data. Using information-theoretic approaches, Evol coanalyzes conservation and coevolution profiles extracted from multiple sequence alignments of protein families with their inferred dynamics. AVAILABILITY AND IMPLEMENTATION: ProDy and Evol are open-source and freely available under MIT License from http://prody.csb.pitt.edu/.
Ahmet Bakan, Anindita Dutta, Wenzhi Mao, S. Chakra Chennubhotla, Timothy R. Lezon, Ivet Bahar
Bioinform.7
2014 Complete Mapping of Substrate Translocation Highlights the Role of LeuT N-terminal Segment in Regulating Transport Cycle
abstract
Neurotransmitter: sodium symporters (NSSs) regulate neuronal signal transmission by clearing excess neurotransmitters from the synapse, assisted by the co-transport of sodium ions. Extensive structural data have been collected in recent years for several members of the NSS family, which opened the way to structure-based studies for a mechanistic understanding of substrate transport. Leucine transporter (LeuT), a bacterial orthologue, has been broadly adopted as a prototype in these studies. This goal has been elusive, however, due to the complex interplay of global and local events as well as missing structural data on LeuT N-terminal segment. We provide here for the first time a comprehensive description of the molecular events leading to substrate/Na+ release to the postsynaptic cell, including the structure and dynamics of the N-terminal segment using a combination of molecular simulations. Substrate and Na+-release follows an influx of water molecules into the substrate/Na+-binding pocket accompanied by concerted rearrangements of transmembrane helices. A redistribution of salt bridges and cation-π interactions at the N-terminal segment prompts substrate release. Significantly, substrate release is followed by the closure of the intracellular gate and a global reconfiguration back to outward-facing state to resume the transport cycle. Two minimally hydrated intermediates, not structurally resolved to date, are identified: one, substrate-bound, stabilized during the passage from outward- to inward-facing state (holo-occluded), and another, substrate-free, along the reverse transition (apo-occluded).
Mary Hongying Cheng, Ivet Bahar
PLoS Comput. Biol.2
2014 Exploring the Conformational Transitions of Biomolecular Systems Using a Simple Two-State Anisotropic Network Model
abstract
Biomolecular conformational transitions are essential to biological functions. Most experimental methods report on the long-lived functional states of biomolecules, but information about the transition pathways between these stable states is generally scarce. Such transitions involve short-lived conformational states that are difficult to detect experimentally. For this reason, computational methods are needed to produce plausible hypothetical transition pathways that can then be probed experimentally. Here we propose a simple and computationally efficient method, called ANMPathway, for constructing a physically reasonable pathway between two endpoints of a conformational transition. We adopt a coarse-grained representation of the protein and construct a two-state potential by combining two elastic network models (ENMs) representative of the experimental structures resolved for the endpoints. The two-state potential has a cusp hypersurface in the configuration space where the energies from both the ENMs are equal. We first search for the minimum energy structure on the cusp hypersurface and then treat it as the transition state. The continuous pathway is subsequently constructed by following the steepest descent energy minimization trajectories starting from the transition state on each side of the cusp hypersurface. Application to several systems of broad biological interest such as adenylate kinase, ATP-driven calcium pump SERCA, leucine transporter and glutamate transporter shows that ANMPathway yields results in good agreement with those from other similar methods and with data obtained from all-atom molecular dynamics simulations, in support of the utility of this simple and efficient approach. Notably the method provides experimentally testable predictions, including the formation of non-native contacts during the transition which we were able to detect in two of the systems we studied. An open-access web server has been created to deliver ANMPathway results.
Avisek Das, Mert Gur, Mary Hongying Cheng, Sunhwan Jo, Ivet Bahar, Benoît Roux
PLoS Comput. Biol.5
2014 ATPase Subdomain IA Is a Mediator of Interdomain Allostery in Hsp70 Molecular Chaperones
abstract
The versatile functions of the heat shock protein 70 (Hsp70) family of molecular chaperones rely on allosteric interactions between their nucleotide-binding and substrate-binding domains, NBD and SBD. Understanding the mechanism of interdomain allostery is essential to rational design of Hsp70 modulators. Yet, despite significant progress in recent years, how the two Hsp70 domains regulate each other's activity remains elusive. Covariance data from experiments and computations emerged in recent years as valuable sources of information towards gaining insights into the molecular events that mediate allostery. In the present study, conservation and covariance properties derived from both sequence and structural dynamics data are integrated with results from Perturbation Response Scanning and in vivo functional assays, so as to establish the dynamical basis of interdomain signal transduction in Hsp70s. Our study highlights the critical roles of SBD residues D481 and T417 in mediating the coupled motions of the two domains, as well as that of G506 in enabling the movements of the α-helical lid with respect to the β-sandwich. It also draws attention to the distinctive role of the NBD subdomains: Subdomain IA acts as a key mediator of signal transduction between the ATP- and substrate-binding sites, this function being achieved by a cascade of interactions predominantly involving conserved residues such as V139, D148, R167 and K155. Subdomain IIA, on the other hand, is distinguished by strong coevolutionary signals (with the SBD) exhibited by a series of residues (D211, E217, L219, T383) implicated in DnaJ recognition. The occurrence of coevolving residues at the DnaJ recognition region parallels the behavior recently observed at the nucleotide-exchange-factor recognition region of subdomain IIB. These findings suggest that Hsp70 tends to adapt to co-chaperone recognition and activity via coevolving residues, whereas interdomain allostery, critical to chaperoning, is robustly enabled by conserved interactions.
Ignacio J. General, Mandy E. Blackburn, Wenzhi Mao, Lila M. Gierasch, Ivet Bahar
PLoS Comput. Biol.6
2012 Coupling between Catalytic Loop Motions and Enzyme Global Dynamics
abstract
Catalytic loop motions facilitate substrate recognition and binding in many enzymes. While these motions appear to be highly flexible, their functional significance suggests that structure-encoded preferences may play a role in selecting particular mechanisms of motions. We performed an extensive study on a set of enzymes to assess whether the collective/global dynamics, as predicted by elastic network models (ENMs), facilitates or even defines the local motions undergone by functional loops. Our dataset includes a total of 117 crystal structures for ten enzymes of different sizes and oligomerization states. Each enzyme contains a specific functional/catalytic loop (10-21 residues long) that closes over the active site during catalysis. Principal component analysis (PCA) of the available crystal structures (including apo and ligand-bound forms) for each enzyme revealed the dominant conformational changes taking place in these loops upon substrate binding. These experimentally observed loop reconfigurations are shown to be predominantly driven by energetically favored modes of motion intrinsically accessible to the enzyme in the absence of its substrate. The analysis suggests that robust global modes cooperatively defined by the overall enzyme architecture also entail local components that assist in suitable opening/closure of the catalytic loop over the active site.
Zeynep Kurkcuoglu, Ahmet Bakan, Duygu Kocaman, Ivet Bahar, Pemra Doruker
PLoS Comput. Biol.4
2011 ProDy: Protein Dynamics Inferred from Theory and Experiments
abstract
SUMMARY: We developed a Python package, ProDy, for structure-based analysis of protein dynamics. ProDy allows for quantitative characterization of structural variations in heterogeneous datasets of structures experimentally resolved for a given biomolecular system, and for comparison of these variations with the theoretically predicted equilibrium dynamics. Datasets include structural ensembles for a given family or subfamily of proteins, their mutants and sequence homologues, in the presence/absence of their substrates, ligands or inhibitors. Numerous helper functions enable comparative analysis of experimental and theoretical data, and visualization of the principal changes in conformations that are accessible in different functional states. ProDy application programming interface (API) has been designed so that users can easily extend the software and implement new methods. AVAILABILITY: ProDy is open source and freely available under GNU General Public License from http://www.csb.pitt.edu/ProDy/.
Ahmet Bakan, Lidio M. C. Meireles, Ivet Bahar
Bioinform.3
2011 Changes in Dynamics upon Oligomerization Regulate Substrate Binding and Allostery in Amino Acid Kinase Family Members
abstract
Oligomerization is a functional requirement for many proteins. The interfacial interactions and the overall packing geometry of the individual monomers are viewed as important determinants of the thermodynamic stability and allosteric regulation of oligomers. The present study focuses on the role of the interfacial interactions and overall contact topology in the dynamic features acquired in the oligomeric state. To this aim, the collective dynamics of enzymes belonging to the amino acid kinase family both in dimeric and hexameric forms are examined by means of an elastic network model, and the softest collective motions (i.e., lowest frequency or global modes of motions) favored by the overall architecture are analyzed. Notably, the lowest-frequency modes accessible to the individual subunits in the absence of multimerization are conserved to a large extent in the oligomer, suggesting that the oligomer takes advantage of the intrinsic dynamics of the individual monomers. At the same time, oligomerization stiffens the interfacial regions of the monomers and confers new cooperative modes that exploit the rigid-body translational and rotational degrees of freedom of the intact monomers. The present study sheds light on the mechanism of cooperative inhibition of hexameric N-acetyl-L-glutamate kinase by arginine and on the allosteric regulation of UMP kinases. It also highlights the significance of the particular quaternary design in selectively determining the oligomer dynamics congruent with required ligand-binding and allosteric activities.
Enrique Marcos, Ramon Crehuet, Ivet Bahar
PLoS Comput. Biol.3
2010 Using Entropy Maximization to Understand the Determinants of Structural Dynamics beyond Native Contact Topology
abstract
Comparison of elastic network model predictions with experimental data has provided important insights on the dominant role of the network of inter-residue contacts in defining the global dynamics of proteins. Most of these studies have focused on interpreting the mean-square fluctuations of residues, or deriving the most collective, or softest, modes of motions that are known to be insensitive to structural and energetic details. However, with increasing structural data, we are in a position to perform a more critical assessment of the structure-dynamics relations in proteins, and gain a deeper understanding of the major determinants of not only the mean-square fluctuations and lowest frequency modes, but the covariance or the cross-correlations between residue fluctuations and the shapes of higher modes. A systematic study of a large set of NMR-determined proteins is analyzed using a novel method based on entropy maximization to demonstrate that the next level of refinement in the elastic network model description of proteins ought to take into consideration properties such as contact order (or sequential separation between contacting residues) and the secondary structure types of the interacting residues, whereas the types of amino acids do not play a critical role. Most importantly, an optimal description of observed cross-correlations requires the inclusion of destabilizing, as opposed to exclusively stabilizing, interactions, stipulating the functional significance of local frustration in imparting native-like dynamics. This study provides us with a deeper understanding of the structural basis of experimentally observed behavior, and opens the way to the development of more accurate models for exploring protein dynamics.
Timothy R. Lezon, Ivet Bahar
PLoS Comput. Biol.2
2010 Role of Hsp70 ATPase Domain Intrinsic Dynamics and Sequence Evolution in Enabling its Functional Interactions with NEFs
abstract
Catalysis of ADP-ATP exchange by nucleotide exchange factors (NEFs) is central to the activity of Hsp70 molecular chaperones. Yet, the mechanism of interaction of this family of chaperones with NEFs is not well understood in the context of the sequence evolution and structural dynamics of Hsp70 ATPase domains. We studied the interactions of Hsp70 ATPase domains with four different NEFs on the basis of the evolutionary trace and co-evolution of the ATPase domain sequence, combined with elastic network modeling of the collective dynamics of the complexes. Our study reveals a subtle balance between the intrinsic (to the ATPase domain) and specific (to interactions with NEFs) mechanisms shared by the four complexes. Two classes of key residues are distinguished in the Hsp70 ATPase domain: (i) highly conserved residues, involved in nucleotide binding, which mediate, via a global hinge-bending, the ATPase domain opening irrespective of NEF binding, and (ii) not-conserved but co-evolved and highly mobile residues, engaged in specific interactions with NEFs (e.g., N57, R258, R262, E283, D285). The observed interplay between these respective intrinsic (pre-existing, structure-encoded) and specific (co-evolved, sequence-dependent) interactions provides us with insights into the allosteric dynamics and functional evolution of the modular Hsp70 ATPase domain.
Lila M. Gierasch, Ivet Bahar
PLoS Comput. Biol.3
2010 On the Conservation of the Slow Conformational Dynamics within the Amino Acid Kinase Family: NAGK the Paradigm
abstract
N-acetyl-L-glutamate kinase (NAGK) is the structural paradigm for examining the catalytic mechanisms and dynamics of amino acid kinase family members. Given that the slow conformational dynamics of the NAGK (at the microseconds time scale or slower) may be rate-limiting, it is of importance to assess the mechanisms of the most cooperative modes of motion intrinsically accessible to this enzyme. Here, we present the results from normal mode analysis using an elastic network model representation, which shows that the conformational mechanisms for substrate binding by NAGK strongly correlate with the intrinsic dynamics of the enzyme in the unbound form. We further analyzed the potential mechanisms of allosteric signalling within NAGK using a Markov model for network communication. Comparative analysis of the dynamics of family members strongly suggests that the low-frequency modes of motion and the associated intramolecular couplings that establish signal transduction are highly conserved among family members, in support of the paradigm sequence-->structure-->dynamics-->function.
Enrique Marcos, Ramon Crehuet, Ivet Bahar
PLoS Comput. Biol.3
2009 Principal component analysis of native ensembles of biomolecular structures (PCA_NEST): insights into functional dynamics
abstract
MOTIVATION: To efficiently analyze the 'native ensemble of conformations' accessible to proteins near their folded state and to extract essential information from observed distributions of conformations, reliable mathematical methods and computational tools are needed. RESULT: Examination of 24 pairs of structures determined by both NMR and X-ray reveals that the differences in the dynamics of the same protein resolved by the two techniques can be tracked to the most robust low frequency modes elucidated by principal component analysis (PCA) of NMR models. The active sites of enzymes are found to be highly constrained in these PCA modes. Furthermore, the residues predicted to be highly immobile are shown to be evolutionarily conserved, lending support to a PCA-based identification of potential functional sites. An online tool, PCA_NEST, is designed to derive the principal modes of conformational changes from structural ensembles resolved by experiments or generated by computations. AVAILABILITY: http://ignm.ccbb.pitt.edu/oPCA_Online.htm
Lee-Wei Yang, Eran Eyal, Ivet Bahar, Akio Kitao
Bioinform.3
2009 Principal component analysis of native ensembles of biomolecular structures (PCA_NEST): insights into functional dynamics
abstract
Bioinformatics 25(5), 606–614
Lee-Wei Yang, Eran Eyal, Ivet Bahar, Akio Kitao
Bioinform.3
2009 Global Motions of the Nuclear Pore Complex: Insights from Elastic Network Models
abstract
The nuclear pore complex (NPC) is the gate to the nucleus. Recent determination of the configuration of proteins in the yeast NPC at approximately 5 nm resolution permits us to study the NPC global dynamics using coarse-grained structural models. We investigate these large-scale motions by using an extended elastic network model (ENM) formalism applied to several coarse-grained representations of the NPC. Two types of collective motions (global modes) are predicted by the ENMs to be intrinsically favored by the NPC architecture: global bending and extension/contraction from circular to elliptical shapes. These motions are shown to be robust against tested variations in the representation of the NPC, and are largely captured by a simple model of a toroid with axially varying mass density. We demonstrate that spoke multiplicity significantly affects the accessible number of symmetric low-energy modes of motion; the NPC-like toroidal structures composed of 8 spokes have access to highly cooperative symmetric motions that are inaccessible to toroids composed of 7 or 9 spokes. The analysis reveals modes of motion that may facilitate macromolecular transport through the NPC, consistent with previous experimental observations.
Timothy R. Lezon, Andrej Sali, Ivet Bahar
PLoS Comput. Biol.3
2009 Allosteric Transitions of Supramolecular Systems Explored by Network Models: Application to Chaperonin GroEL
abstract
Identification of pathways involved in the structural transitions of biomolecular systems is often complicated by the transient nature of the conformations visited across energy barriers and the multiplicity of paths accessible in the multidimensional energy landscape. This task becomes even more challenging in exploring molecular systems on the order of megadaltons. Coarse-grained models that lend themselves to analytical solutions appear to be the only possible means of approaching such cases. Motivated by the utility of elastic network models for describing the collective dynamics of biomolecular systems and by the growing theoretical and experimental evidence in support of the intrinsic accessibility of functional substates, we introduce a new method, adaptive anisotropic network model (aANM), for exploring functional transitions. Application to bacterial chaperonin GroEL and comparisons with experimental data, results from action minimization algorithm, and previous simulations support the utility of aANM as a computationally efficient, yet physically plausible, tool for unraveling potential transition pathways sampled by large complexes/assemblies. An important outcome is the assessment of the critical inter-residue interactions formed/broken near the transition state(s), most of which involve conserved residues.
Peter Májek, Ivet Bahar
PLoS Comput. Biol.3
2008 Analysis of correlated mutations in HIV-1 protease using spectral clustering
abstract
MOTIVATION: The ability of human immunodeficiency virus-1 (HIV-1) protease to develop mutations that confer multi-drug resistance (MDR) has been a major obstacle in designing rational therapies against HIV. Resistance is usually imparted by a cooperative mechanism that can be elucidated by a covariance analysis of sequence data. Identification of such correlated substitutions of amino acids may be obscured by evolutionary noise. RESULTS: HIV-1 protease sequences from patients subjected to different specific treatments (set 1), and from untreated patients (set 2) were subjected to sequence covariance analysis by evaluating the mutual information (MI) between all residue pairs. Spectral clustering of the resulting covariance matrices disclosed two distinctive clusters of correlated residues: the first, observed in set 1 but absent in set 2, contained residues involved in MDR acquisition; and the second, included those residues differentiated in the various HIV-1 protease subtypes, shortly referred to as the phylogenetic cluster. The MDR cluster occupies sites close to the central symmetry axis of the enzyme, which overlap with the global hinge region identified from coarse-grained normal-mode analysis of the enzyme structure. The phylogenetic cluster, on the other hand, occupies solvent-exposed and highly mobile regions. This study demonstrates (i) the possibility of distinguishing between the correlated substitutions resulting from neutral mutations and those induced by MDR upon appropriate clustering analysis of sequence covariance data and (ii) a connection between global dynamics and functional substitution of amino acids.
Eran Eyal, Ivet Bahar
Bioinform.3
2007 Rapid assessment of correlated amino acids from pair-to-pair (P2P) substitution matrices
abstract
UNLABELLED: Identification of correlated amino acids in proteins has been a topic of broad interest in view of its functional implications and importance in protein design. A new set of pair-to-pair (P2P) substitution matrices for amino acids was recently introduced as a useful tool for inferring information on such correlated sites. We present a website developed for automated application of these matrices for analysis of query sequences. The site offers options for graphical analysis of correlations, as well as visualization of correlated amino acids on representative, structurally characterized, members of the examined family of sequences. AVAILABILITY: http://www.ccbb.pitt.edu/p2p.
Eran Eyal, Shmuel Pietrokovski, Ivet Bahar
Bioinform.3
2007 Recruitment of rare 3-grams at functional sites: Is this a mechanism for increasing enzyme specificity?
abstract
BACKGROUND: A wealth of unannotated and functionally unknown protein sequences has accumulated in recent years with rapid progresses in sequence genomics, giving rise to ever increasing demands for developing methods to efficiently assess functional sites. Sequence and structure conservations have traditionally been the major criteria adopted in various algorithms to identify functional sites. Here, we focus on the distributions of the 203 different types of 3-grams (or triplets of sequentially contiguous amino acid) in the entire space of sequences accumulated to date in the UniProt database, and focus in particular on the rare 3-grams distinguished by their high entropy-based information content. RESULTS: Comparison of the UniProt distributions with those observed near/at the active sites on a non-redundant dataset of 59 enzyme/ligand complexes shows that the active sites preferentially recruit 3-grams distinguished by their low frequency in the UniProt. Three cases, Src kinase, hemoglobin, and tyrosyl-tRNA synthetase, are discussed in details to illustrate the biological significance of the results. CONCLUSION: The results suggest that recruitment of rare 3-grams may be an efficient mechanism for increasing specificity at functional sites. Rareness/scarcity emerges as a feature that may assist in identifying key sites for proteins function, providing information complementary to that derived from sequence alignments. In addition it provides us (for the first time) with a means of identifying potentially functional sites from sequence information alone, when sequence conservation properties are not available.
Dror Tobi, Ivet Bahar
BMC Bioinform.2
2007 Signal Propagation in Proteins and Relation to Equilibrium Fluctuations
abstract
Elastic network (EN) models have been widely used in recent years for describing protein dynamics, based on the premise that the motions naturally accessible to native structures are relevant to biological function. We posit that equilibrium motions also determine communication mechanisms inherent to the network architecture. To this end, we explore the stochastics of a discrete-time, discrete-state Markov process of information transfer across the network of residues. We measure the communication abilities of residue pairs in terms of hit and commute times, i.e., the number of steps it takes on an average to send and receive signals. Functionally active residues are found to possess enhanced communication propensities, evidenced by their short hit times. Furthermore, secondary structural elements emerge as efficient mediators of communication. The present findings provide us with insights on the topological basis of communication in proteins and design principles for efficient signal transduction. While hit/commute times are information-theoretic concepts, a central contribution of this work is to rigorously show that they have physical origins directly relevant to the equilibrium fluctuations of residues predicted by EN models.
S. Chakra Chennubhotla, Ivet Bahar
PLoS Comput. Biol.2
2007 Correction: Signal Propagation in Proteins and Relation to Equilibrium Fluctuations
abstract
Correction for: Chennubhotla C, Bahar I (2007) Signal propagation in proteins and relation to equilibrium fluctuations. PLoS Comput Biol 3(9): e172. 10.1371/journal.pcbi.0030172 Four mathematical expressions appeared incorrectly. The correct expressions follow. The hitting time expression Equation 14 involves three different types of contributions: a one-body term that depends on the destination node, ; a two-body term that depends on the initial and final nodes, ; and a series of three-body terms that depend on intermediate nodes, in addition to the two end points, . Derivation of Equation 14. The discussion below borrows from results in [12,13]. Deriving from is a three-step process: …
S. Chakra Chennubhotla, Ivet Bahar
PLoS Comput. Biol.2
2006 Markov Methods for Hierarchical Coarse-Graining of Large Protein Dynamics
S. Chakra Chennubhotla, Ivet Bahar
RECOMB2
2006 Anisotropic network model: systematic evaluation and a new web interface
abstract
MOTIVATION: The Anisotropic Network Model (ANM) is a simple yet powerful model for normal mode analysis of proteins. Despite its broad use for exploring biomolecular collective motions, ANM has not been systematically evaluated to date. A lack of a convenient interface has been an additional obstacle for easy usage. RESULTS: ANM has been evaluated on a large set of proteins to establish the optimal model parameters that achieve the highest correlation with experimental data and its limits of accuracy and applicability. Residue fluctuations in globular proteins are shown to be more accurately predicted than those in nonglobular proteins, and core residues are more accurately described than solvent-exposed ones. Significant improvement in agreement with experiments is observed with increase in the resolution of the examined structure. A new server for ANM calculations is presented, which offers flexible options for controlling model parameters and output formats, interactive animation of collective modes and advanced graphical features. AVAILABILITY: ANM server (http://www.ccbb.pitt.edu/anm)
Eran Eyal, Lee-Wei Yang, Ivet Bahar
Bioinform.3
2006 Topological basis of signal integration in the transcriptional-regulatory network of the yeast, Saccharomyces cerevisiae
abstract
BACKGROUND: Signal recognition and information processing is a fundamental cellular function, which in part involves comprehensive transcriptional regulatory (TR) mechanisms carried out in response to complex environmental signals in the context of the cell's own internal state. However, the network topological basis of developing such integrated responses remains poorly understood. RESULTS: By studying the TR network of the yeast Saccharomyces cerevisiae we show that an intermediate layer of transcription factors naturally segregates into distinct subnetworks. In these topological units transcription factors are densely interlinked in a largely hierarchical manner and respond to external signals by utilizing a fraction of these subnets. CONCLUSION: As transcriptional regulation represents the 'slow' component of overall information processing, the identified topology suggests a model in which successive waves of transcriptional regulation originating from distinct fractions of the TR network control robust integrated responses to complex stimuli.
Illés J. Farkas, S. Chakra Chennubhotla, Ivet Bahar, Zoltán N. Oltvai
BMC Bioinform.4
2005 iGNM: a database of protein functional motions based on Gaussian Network Model
abstract
MOTIVATION: The knowledge of protein structure is not sufficient for understanding and controlling its function. Function is a dynamic property. Although protein structural information has been rapidly accumulating in databases, little effort has been invested to date toward systematically characterizing protein dynamics. The recent success of analytical methods based on elastic network models, and in particular the Gaussian Network Model (GNM), permits us to perform a high-throughput analysis of the collective dynamics of proteins. RESULTS: We computed the GNM dynamics for 20 058 structures from the Protein Data Bank, and generated information on the equilibrium dynamics at the level of individual residues. The results are stored on a web-based system called iGNM and configured so as to permit the users to visualize or download the results through a standard web browser using a simple search engine. Static and animated images for describing the conformational mobility of proteins over a broad range of normal modes are accessible, along with an online calculation engine available for newly deposited structures. A case study of the dynamics of 20 non-homologous hydrolases is presented to illustrate the utility of the iGNM database for identifying key residues that control the cooperative motions and revealing the connection between collective dynamics and catalytic activity.
Lee-Wei Yang, Xiong Liu 0001, Christopher Jon Jursa, Mark Holliman, A. J. Rader, Hassan A. Karimi, Ivet Bahar
Bioinform.7