VLDB 2026 Research / reviewers in the wild / expert
Andrej Sali
dblp:59/1163
· DBLP profile ↗
28ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 28 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
13 papers |
Bioinformatics and computational biology · 100% |
Topics — the 19 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
protein structure prediction |
0.3 | 4 | 2013 | Optimized atomic statistical potentials: assessment of protein interfaces and loops · Bioinform. 2013 ModLink+: improving fold recognition by using protein-protein interactions · Bioinform. 2009 ModLoop: automated modeling of loops in protein structures · Bioinform. 2003 |
Bioinformatics and computational biology
structural bioinformatics |
0.3 | 4 | 2012 | A method for integrative structure determination of protein-protein complexes · Bioinform. 2012 PIBASE: a comprehensive database of structurally defined protein interfaces · Bioinform. 2005 LigBase: a database of families of aligned ligand binding sites in known protein sequences and structures · Bioinform. 2002 |
Bioinformatics and computational biology › protein structure prediction
loop modeling |
0.2 | 2 | 2013 | Optimized atomic statistical potentials: assessment of protein interfaces and loops · Bioinform. 2013 ModLoop: automated modeling of loops in protein structures · Bioinform. 2003 |
Bioinformatics and computational biology › protein structure analysis
protein structure alignment |
0.2 | 3 | 2012 | SALIGN: a web server for alignment of multiple protein sequences and structures · Bioinform. 2012 DBAli: a database of protein structure alignments · Bioinform. 2001 LigBase: a database of families of aligned ligand binding sites in known protein sequences and structures · Bioinform. 2002 |
Bioinformatics and computational biology › protein structure prediction
protein-protein docking |
0.2 | 1 | 2013 | Optimized atomic statistical potentials: assessment of protein interfaces and loops · Bioinform. 2013 |
Bioinformatics and computational biology › molecular informatics › molecular modeling
scoring function |
0.2 | 1 | 2013 | Optimized atomic statistical potentials: assessment of protein interfaces and loops · Bioinform. 2013 |
Bioinformatics and computational biology › protein structure analysis
statistical potential |
0.2 | 1 | 2013 | Optimized atomic statistical potentials: assessment of protein interfaces and loops · Bioinform. 2013 |
Bioinformatics and computational biology
multiple sequence alignment |
0.2 | 2 | 2012 | SALIGN: a web server for alignment of multiple protein sequences and structures · Bioinform. 2012 ModView, visualization of multiple protein sequences and structures · Bioinform. 2003 |
Bioinformatics and computational biology
proteomics |
0.1 | 1 | 2010 | Prediction of protease substrates using sequence and structure features · Bioinform. 2010 |
Bioinformatics and computational biology › protein structure prediction › template-based modeling
fold recognition |
0.1 | 1 | 2009 | ModLink+: improving fold recognition by using protein-protein interactions · Bioinform. 2009 |
Bioinformatics and computational biology › protein function prediction
functional impact prediction |
0.1 | 1 | 2005 | LS-SNP: large-scale annotation of coding non-synonymous SNPs based on multiple information sources · Bioinform. 2005 |
Bioinformatics and computational biology › genome annotation
genomic variant annotation |
0.1 | 1 | 2005 | LS-SNP: large-scale annotation of coding non-synonymous SNPs based on multiple information sources · Bioinform. 2005 |
Bioinformatics and computational biology › structural bioinformatics › protein modeling
comparative protein structure modeling |
0.0 | 1 | 2012 | SALIGN: a web server for alignment of multiple protein sequences and structures · Bioinform. 2012 |
Bioinformatics and computational biology › molecular informatics › molecular modeling
molecular docking |
0.0 | 1 | 2012 | A method for integrative structure determination of protein-protein complexes · Bioinform. 2012 |
Bioinformatics and computational biology › structural bioinformatics › protein structure representation
protein structure visualization |
0.0 | 1 | 2003 | ModView, visualization of multiple protein sequences and structures · Bioinform. 2003 |
Bioinformatics and computational biology › structural bioinformatics
protein structure database |
0.0 | 1 | 2002 | LigBase: a database of families of aligned ligand binding sites in known protein sequences and structures · Bioinform. 2002 |
Bioinformatics and computational biology › structural biology
protein structure and function |
0.0 | 1 | 2010 | Prediction of protease substrates using sequence and structure features · Bioinform. 2010 |
Bioinformatics and computational biology
benchmark evaluation of prediction tools |
0.0 | 1 | 2001 | EVA: continuous automatic evaluation of protein structure prediction servers · Bioinform. 2001 |
Bioinformatics and computational biology › structural bioinformatics
protein structure classification |
0.0 | 1 | 2005 | PIBASE: a comprehensive database of structurally defined protein interfaces · Bioinform. 2005 |
Methods — techniques the papers use, named apart from their topics
file format agnostic processing · 0.9bayesian inference · 0.2profile-profile alignment · 0.1electron microscopy · 0.1dendrogram · 0.1cross-linking mass spectrometry · 0.1SAXS · 0.1NMR spectroscopy · 0.1support vector machine · 0.1HHsearch · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | tttrlib: modular software for integrating fluorescence spectroscopy, imaging, and molecular modelingabstractSUMMARY: We introduce software for reading, writing and processing fluorescence single-molecule and image spectroscopy data and developing analysis pipelines to unify various spectroscopic analysis tools. Our software can be used for processing multiple experiment types, e.g. for time-resolved single-molecule spectroscopy, laser scanning microscopy, fluorescence correlation spectroscopy and image correlation spectroscopy. The software is file format agnostic and processes multiple time-resolved data formats and outputs. Our software eliminates the need for data conversion and mitigates data archiving issues. AVAILABILITY AND IMPLEMENTATION: tttrlib is available via pip (https://pypi.org/project/tttrlib/) and bioconda while the open-source code is available via GitHub (https://github.com/fluorescence-tools/tttrlib). Presented examples and additional documentation demonstrating how to implement in vitro and live-cell image spectroscopy analysis are available at https://docs.peulen.xyz/tttrlib and https://zenodo.org/records/14002224. Thomas-Otavio Peulen, Katherina Hemmen, Annemarie Greife, Ben M. Webb, Suren Felekyan, Andrej Sali, Claus A. M. Seidel, Hugo Sanabria, Katrin G. Heinze |
Bioinform. | 6 |
| 2016 | Clustering of disulfide-rich peptides provides scaffolds for hit discovery by phage display: application to interleukin-23abstractAbstract Background Disulfide-rich peptides (DRPs) are found throughout nature. They are suitable scaffolds for drug development due to their small cores, whose disulfide bonds impart extraordinary chemical and biological stability. A challenge in developing a DRP therapeutic is to engineer binding to a specific target. This challenge can be overcome by (i) sampling the large sequence space of a given scaffold through a phage display library and by (ii) panning multiple libraries encoding structurally distinct scaffolds. Here, we implement a protocol for defining these diverse scaffolds, based on clustering structurally defined DRPs according to their conformational similarity. Results We developed and applied a hierarchical clustering protocol based on DRP structural similarity, followed by two post-processing steps, to classify 806 unique DRP structures into 81 clusters. The 20 most populated clusters comprised 85% of all DRPs. Representative scaffolds were selected from each of these clusters; the representatives were structurally distinct from one another, but similar to other DRPs in their respective clusters. To demonstrate the utility of the clusters, phage libraries were constructed for three of the representative scaffolds and panned against interleukin-23. One library produced a peptide that bound to this target with an IC50 of 3.3 μM. Conclusions Most DRP clusters contained members that were diverse in sequence, host organism, and interacting proteins, indicating that cluster members were functionally diverse despite having similar structure. Only 20 peptide scaffolds accounted for most of the natural DRP structural diversity, providing suitable starting points for seeding phage display experiments. Through selection of the scaffold surface to vary in phage display, libraries can be designed that present sequence diversity in architecturally distinct, biologically relevant combinations of secondary structures. We supported this hypothesis with a proof-of-concept experiment in which three phage libraries were constructed and panned against the IL-23 target, resulting in a single-digit μM hit and suggesting that a collection of libraries based on the full set of 20 scaffolds increases the potential to identify efficiently peptide binders to a protein target in a drug discovery program. David T. Barkan, Xiao-li Cheng, Herodion Celino, Tran Trung Tran, Ashok Bhandari, Charles S. Craik, Andrej Sali, Mark L. Smythe |
BMC Bioinform. | 7 |
| 2015 | Prediction of Functionally Important Phospho-Regulatory Events in Xenopus laevis OocytesabstractThe African clawed frog Xenopus laevis is an important model organism for studies in developmental and cell biology, including cell-signaling. However, our knowledge of X. laevis protein post-translational modifications remains scarce. Here, we used a mass spectrometry-based approach to survey the phosphoproteome of this species, compiling a list of 2636 phosphosites. We used structural information and phosphoproteomic data for 13 other species in order to predict functionally important phospho-regulatory events. We found that the degree of conservation of phosphosites across species is predictive of sites with known molecular function. In addition, we predicted kinase-protein interactions for a set of cell-cycle kinases across all species. The degree of conservation of kinase-protein interactions was found to be predictive of functionally relevant regulatory interactions. Finally, using comparative protein structure models, we find that phosphosites within structured domains tend to be located at positions with high conformational flexibility. Our analysis suggests that a small class of phosphosites occurs in positions that have the potential to regulate protein conformation. Jeffrey R. Johnson, Silvia D. Santos, Tasha Johnson, Ursula Pieper 0001, Marta Strumillo, Omar Wagih, Andrej Sali, Nevan J. Krogan, Pedro Beltrão |
PLoS Comput. Biol. | 7 |
| 2014 | A Systematic Computational Analysis of Biosynthetic Gene Cluster Evolution: Lessons for Engineering BiosynthesisabstractBacterial secondary metabolites are widely used as antibiotics, anticancer drugs, insecticides and food additives. Attempts to engineer their biosynthetic gene clusters (BGCs) to produce unnatural metabolites with improved properties are often frustrated by the unpredictability and complexity of the enzymes that synthesize these molecules, suggesting that genetic changes within BGCs are limited by specific constraints. Here, by performing a systematic computational analysis of BGC evolution, we derive evidence for three findings that shed light on the ways in which, despite these constraints, nature successfully invents new molecules: 1) BGCs for complex molecules often evolve through the successive merger of smaller sub-clusters, which function as independent evolutionary entities. 2) An important subset of polyketide synthases and nonribosomal peptide synthetases evolve by concerted evolution, which generates sets of sequence-homogenized domains that may hold promise for engineering efforts since they exhibit a high degree of functional interoperability, 3) Individual BGC families evolve in distinct ways, suggesting that design strategies should take into account family-specific functional constraints. These findings suggest novel strategies for using synthetic biology to rationally engineer biosynthetic pathways. Marnix H. Medema, Peter Cimermancic, Andrej Sali, Eriko Takano, Michael A. Fischbach |
PLoS Comput. Biol. | 3 |
| 2013 | Optimized atomic statistical potentials: assessment of protein interfaces and loopsabstractMOTIVATION: Statistical potentials have been widely used for modeling whole proteins and their parts (e.g. sidechains and loops) as well as interactions between proteins, nucleic acids and small molecules. Here, we formulate the statistical potentials entirely within a statistical framework, avoiding questionable statistical mechanical assumptions and approximations, including a definition of the reference state. RESULTS: We derive a general Bayesian framework for inferring statistically optimized atomic potentials (SOAP) in which the reference state is replaced with data-driven 'recovery' functions. Moreover, we restrain the relative orientation between two covalent bonds instead of a simple distance between two atoms, in an effort to capture orientation-dependent interactions such as hydrogen bonds. To demonstrate this general approach, we computed statistical potentials for protein-protein docking (SOAP-PP) and loop modeling (SOAP-Loop). For docking, a near-native model is within the top 10 scoring models in 40% of the PatchDock benchmark cases, compared with 23 and 27% for the state-of-the-art ZDOCK and FireDock scoring functions, respectively. Similarly, for modeling 12-residue loops in the PLOP benchmark, the average main-chain root mean square deviation of the best scored conformations by SOAP-Loop is 1.5 Å, close to the average root mean square deviation of the best sampled conformations (1.2 Å) and significantly better than that selected by Rosetta (2.1 Å), DFIRE (2.3 Å), DOPE (2.5 Å) and PLOP scoring functions (3.0 Å). Our Bayesian framework may also result in more accurate statistical potentials for additional modeling applications, thus affording better leverage of the experimentally determined protein structures. AVAILABILITY AND IMPLEMENTATION: SOAP-PP and SOAP-Loop are available as part of MODELLER (http://salilab.org/modeller). Guang Qiang Dong, Hao Fan 0002, Dina Schneidman, Ben M. Webb, Andrej Sali |
Bioinform. | 5 |
| 2013 | Target Prediction for an Open Access Set of Compounds Active against Mycobacterium tuberculosisabstractMycobacterium tuberculosis, the causative agent of tuberculosis (TB), infects an estimated two billion people worldwide and is the leading cause of mortality due to infectious disease. The development of new anti-TB therapeutics is required, because of the emergence of multi-drug resistance strains as well as co-infection with other pathogens, especially HIV. Recently, the pharmaceutical company GlaxoSmithKline published the results of a high-throughput screen (HTS) of their two million compound library for anti-mycobacterial phenotypes. The screen revealed 776 compounds with significant activity against the M. tuberculosis H37Rv strain, including a subset of 177 prioritized compounds with high potency and low in vitro cytotoxicity. The next major challenge is the identification of the target proteins. Here, we use a computational approach that integrates historical bioassay data, chemical properties and structural comparisons of selected compounds to propose their potential targets in M. tuberculosis. We predicted 139 target--compound links, providing a necessary basis for further studies to characterize the mode of action of these compounds. The results from our analysis, including the predicted structural models, are available to the wider scientific community in the open source mode, to encourage further development of novel TB therapeutics. Francisco Martínez-Jiménez, George Papadatos, Lun Yang, Iain M. Wallace, Ursula Pieper 0001, Andrej Sali, James R. Brown, John P. Overington, Marc A. Martí-Renom |
PLoS Comput. Biol. | 7 |
| 2012 | SALIGN: a web server for alignment of multiple protein sequences and structuresabstractSUMMARY: Accurate alignment of protein sequences and/or structures is crucial for many biological analyses, including functional annotation of proteins, classifying protein sequences into families, and comparative protein structure modeling. Described here is a web interface to SALIGN, the versatile protein multiple sequence/structure alignment module of MODELLER. The web server automatically determines the best alignment procedure based on the inputs, while allowing the user to override default parameter values. Multiple alignments are guided by a dendrogram computed from a matrix of all pairwise alignment scores. When aligning sequences to structures, SALIGN uses structural environment information to place gaps optimally. If two multiple sequence alignments of related proteins are input to the server, a profile-profile alignment is performed. All features of the server have been previously optimized for accuracy, especially in the contexts of comparative modeling and identification of interacting protein partners. AVAILABILITY: The SALIGN web server is freely accessible to the academic community at http://salilab.org/salign. SALIGN is a module of the MODELLER software, also freely available to academic users (http://salilab.org/modeller). CONTACT: [email protected]; [email protected]. Hannes Braberg, Ben M. Webb, Elina Tjioe, Ursula Pieper 0001, Andrej Sali, M. S. Madhusudhan 0001 |
Bioinform. | 5 |
| 2012 | A method for integrative structure determination of protein-protein complexesabstractMOTIVATION: Structural characterization of protein interactions is necessary for understanding and modulating biological processes. On one hand, X-ray crystallography or NMR spectroscopy provide atomic resolution structures but the data collection process is typically long and the success rate is low. On the other hand, computational methods for modeling assembly structures from individual components frequently suffer from high false-positive rate, rarely resulting in a unique solution. RESULTS: Here, we present a combined approach that computationally integrates data from a variety of fast and accessible experimental techniques for rapid and accurate structure determination of protein-protein complexes. The integrative method uses atomistic models of two interacting proteins and one or more datasets from five accessible experimental techniques: a small-angle X-ray scattering (SAXS) profile, 2D class average images from negative-stain electron microscopy micrographs (EM), a 3D density map from single-particle negative-stain EM, residue type content of the protein-protein interface from NMR spectroscopy and chemical cross-linking detected by mass spectrometry. The method is tested on a docking benchmark consisting of 176 known complex structures and simulated experimental data. The near-native model is the top scoring one for up to 61% of benchmark cases depending on the included experimental datasets; in comparison to 10% for standard computational docking. We also collected SAXS, 2D class average images and 3D density map from negative-stain EM to model the PCSK9 antigen-J16 Fab antibody complex, followed by validation of the model by a subsequently available X-ray crystallographic structure. Dina Schneidman, Andrea Rossi 0005, Agustin Avila-Sakar, Seung Joong Kim, Javier A. Velázquez-Muriel, Pavel Strop, Kristin A. Krukenberg, Maofu Liao, Ho Min Kim, Solmaz Sobhanifar, Volker Dötsch, Arvind Rajpal, Jaume Pons, David A. Agard, Andrej Sali |
Bioinform. | 17 |
| 2010 | Prediction of protease substrates using sequence and structure featuresabstractMOTIVATION: Granzyme B (GrB) and caspases cleave specific protein substrates to induce apoptosis in virally infected and neoplastic cells. While substrates for both types of proteases have been determined experimentally, there are many more yet to be discovered in humans and other metazoans. Here, we present a bioinformatics method based on support vector machine (SVM) learning that identifies sequence and structural features important for protease recognition of substrate peptides and then uses these features to predict novel substrates. Our approach can act as a convenient hypothesis generator, guiding future experiments by high-confidence identification of peptide-protein partners. RESULTS: The method is benchmarked on the known substrates of both protease types, including our literature-curated GrB substrate set (GrBah). On these benchmark sets, the method outperforms a number of other methods that consider sequence only, predicting at a 0.87 true positive rate (TPR) and a 0.13 false positive rate (FPR) for caspase substrates, and a 0.79 TPR and a 0.21 FPR for GrB substrates. The method is then applied to approximately 25 000 proteins in the human proteome to generate a ranked list of predicted substrates of each protease type. Two of these predictions, AIF-1 and SMN1, were selected for further experimental analysis, and each was validated as a GrB substrate. AVAILABILITY: All predictions for both protease types are publically available at http://salilab.org/peptide. A web server is at the same site that allows a user to train new SVM models to make predictions for any protein that recognizes specific oligopeptide ligands. David T. Barkan, Daniel R. Hostetter, Sami Mahrus, Ursula Pieper 0001, James A. Wells, Charles S. Craik, Andrej Sali |
Bioinform. | 7 |
| 2010 | The Overlap of Small Molecule and Protein Binding Sites within Families of Protein StructuresabstractProtein-protein interactions are challenging targets for modulation by small molecules. Here, we propose an approach that harnesses the increasing structural coverage of protein complexes to identify small molecules that may target protein interactions. Specifically, we identify ligand and protein binding sites that overlap upon alignment of homologous proteins. Of the 2,619 protein structure families observed to bind proteins, 1,028 also bind small molecules (250-1000 Da), and 197 exhibit a statistically significant (p<0.01) overlap between ligand and protein binding positions. These "bi-functional positions", which bind both ligands and proteins, are particularly enriched in tyrosine and tryptophan residues, similar to "energetic hotspots" described previously, and are significantly less conserved than mono-functional and solvent exposed positions. Homology transfer identifies ligands whose binding sites overlap at least 20% of the protein interface for 35% of domain-domain and 45% of domain-peptide mediated interactions. The analysis recovered known small-molecule modulators of protein interactions as well as predicted new interaction targets based on the sequence similarity of ligand binding sites. We illustrate the predictive utility of the method by suggesting structural mechanisms for the effects of sanglifehrin A on HIV virion production, bepridil on the cellular entry of anthrax edema factor, and fusicoccin on vertebrate developmental pathways. The results, available at http://pibase.janelia.org, represent a comprehensive collection of structurally characterized modulators of protein interactions, and suggest that homologous structures are a useful resource for the rational design of interaction modulators. Fred P. Davis, Andrej Sali |
PLoS Comput. Biol. | 2 |
| 2009 | ModLink+: improving fold recognition by using protein-protein interactionsabstractMOTIVATION: Several strategies have been developed to predict the fold of a target protein sequence, most of which are based on aligning the target sequence to other sequences of known structure. Previously, we demonstrated that the consideration of protein-protein interactions significantly increases the accuracy of fold assignment compared with PSI-BLAST sequence comparisons. A drawback of our method was the low number of proteins to which a fold could be assigned. Here, we present an improved version of the method that addresses this limitation. We also compare our method to other state-of-the-art fold assignment methodologies. RESULTS: Our approach (ModLink+) has been tested on 3716 proteins with domain folds classified in the Structural Classification Of Proteins (SCOP) as well as known interacting partners in the Database of Interacting Proteins (DIP). For this test set, the ratio of success [positive predictive value (PPV)] on fold assignment increases from 75% for PSI-BLAST, 83% for HHSearch and 81% for PRC to >90% for ModLink+at the e-value cutoff of 10(-3). Under this e-value, ModLink+can assign a fold to 30-45% of the proteins in the test set, while our previous method could cover <25%. When applied to 6384 proteins with unknown fold in the yeast proteome, ModLink+combined with PSI-BLAST assigns a fold for domains in 3738 proteins, while PSI-BLAST alone covers only 2122 proteins, HHSearch 2969 and PRC 2826 proteins, using a threshold e-value that would represent a PPV >82% for each method in the test set. AVAILABILITY: The ModLink+server is freely accessible in the World Wide Web at http://sbi.imim.es/modlink/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Oriol Fornes, Ramon Aragues, Jordi Espadaler, Marc A. Martí-Renom, Andrej Sali, Baldomero Oliva |
Bioinform. | 5 |
| 2009 | Global Motions of the Nuclear Pore Complex: Insights from Elastic Network ModelsabstractThe nuclear pore complex (NPC) is the gate to the nucleus. Recent determination of the configuration of proteins in the yeast NPC at approximately 5 nm resolution permits us to study the NPC global dynamics using coarse-grained structural models. We investigate these large-scale motions by using an extended elastic network model (ENM) formalism applied to several coarse-grained representations of the NPC. Two types of collective motions (global modes) are predicted by the ENMs to be intrinsically favored by the NPC architecture: global bending and extension/contraction from circular to elliptical shapes. These motions are shown to be robust against tested variations in the representation of the NPC, and are largely captured by a simple model of a toroid with axially varying mass density. We demonstrate that spoke multiplicity significantly affects the accessible number of symmetric low-energy modes of motion; the NPC-like toroidal structures composed of 8 spokes have access to highly cooperative symmetric motions that are inaccessible to toroids composed of 7 or 9 spokes. The analysis reveals modes of motion that may facilitate macromolecular transport through the NPC, consistent with previous experimental observations. Timothy R. Lezon, Andrej Sali, Ivet Bahar |
PLoS Comput. Biol. | 2 |
| 2008 | Prediction of enzyme function by combining sequence similarity and protein interactionsabstractBACKGROUND: A number of studies have used protein interaction data alone for protein function prediction. Here, we introduce a computational approach for annotation of enzymes, based on the observation that similar protein sequences are more likely to perform the same function if they share similar interacting partners. RESULTS: The method has been tested against the PSI-BLAST program using a set of 3,890 protein sequences from which interaction data was available. For protein sequences that align with at least 40% sequence identity to a known enzyme, the specificity of our method in predicting the first three EC digits increased from 80% to 90% at 80% coverage when compared to PSI-BLAST. CONCLUSION: Our method can also be used in proteins for which homologous sequences with known interacting partners can be detected. Thus, our method could increase 10% the specificity of genome-wide enzyme predictions based on sequence matching by PSI-BLAST alone. Jordi Espadaler, Narayanan Eswar, Enrique Querol, Francesc X. Avilés, Andrej Sali, Marc A. Martí-Renom, Baldomero Oliva |
BMC Bioinform. | 5 |
| 2008 | Evolutionarily Conserved Substrate Substructures for Automated Annotation of Enzyme SuperfamiliesabstractThe evolution of enzymes affects how well a species can adapt to new environmental conditions. During enzyme evolution, certain aspects of molecular function are conserved while other aspects can vary. Aspects of function that are more difficult to change or that need to be reused in multiple contexts are often conserved, while those that vary may indicate functions that are more easily changed or that are no longer required. In analogy to the study of conservation patterns in enzyme sequences and structures, we have examined the patterns of conservation and variation in enzyme function by analyzing graph isomorphisms among enzyme substrates of a large number of enzyme superfamilies. This systematic analysis of substrate substructures establishes the conservation patterns that typify individual superfamilies. Specifically, we determined the chemical substructures that are conserved among all known substrates of a superfamily and the substructures that are reacting in these substrates and then examined the relationship between the two. Across the 42 superfamilies that were analyzed, substantial variation was found in how much of the conserved substructure is reacting, suggesting that superfamilies may not be easily grouped into discrete and separable categories. Instead, our results suggest that many superfamilies may need to be treated individually for analyses of evolution, function prediction, and guiding enzyme engineering strategies. Annotating superfamilies with these conserved and reacting substructure patterns provides information that is orthogonal to information provided by studies of conservation in superfamily sequences and structures, thereby improving the precision with which we can predict the functions of enzymes of unknown function and direct studies in enzyme engineering. Because the method is automated, it is suitable for large-scale characterization and comparison of fundamental functional capabilities of both characterized and uncharacterized enzyme superfamilies. Ranyee A. Chiang, Andrej Sali, Patricia C. Babbitt |
PLoS Comput. Biol. | 2 |
| 2007 | The AnnoLite and AnnoLyze programs for comparative annotation of protein structuresabstractBACKGROUND: Advances in structural biology, including structural genomics, have resulted in a rapid increase in the number of experimentally determined protein structures. However, about half of the structures deposited by the structural genomics consortia have little or no information about their biological function. Therefore, there is a need for tools for automatically and comprehensively annotating the function of protein structures. We aim to provide such tools by applying comparative protein structure annotation that relies on detectable relationships between protein structures to transfer functional annotations. Here we introduce two programs, AnnoLite and AnnoLyze, which use the structural alignments deposited in the DBAli database. DESCRIPTION: AnnoLite predicts the SCOP, CATH, EC, InterPro, PfamA, and GO terms with an average sensitivity of ~90% and average precision of ~80%. AnnoLyze predicts ligand binding site and domain interaction patches with an average sensitivity of ~70% and average precision of ~30%, correctly localizing binding sites for small molecules in ~95% of its predictions. CONCLUSION: The AnnoLite and AnnoLyze programs for comparative annotation of protein structures can reliably and automatically annotate new protein structures. The programs are fully accessible via the Internet as part of the DBAli suite of tools at http://salilab.org/DBAli/. Marc A. Martí-Renom, Andrea Rossi 0005, Fátima Al-Shahrour, Fred P. Davis, Ursula Pieper 0001, Joaquín Dopazo, Andrej Sali |
BMC Bioinform. | 7 |
| 2007 | Characterization of Protein Hubs by Inferring Interacting Motifs from Protein InteractionsabstractThe characterization of protein interactions is essential for understanding biological systems. While genome-scale methods are available for identifying interacting proteins, they do not pinpoint the interacting motifs (e.g., a domain, sequence segments, a binding site, or a set of residues). Here, we develop and apply a method for delineating the interacting motifs of hub proteins (i.e., highly connected proteins). The method relies on the observation that proteins with common interaction partners tend to interact with these partners through a common interacting motif. The sole input for the method are binary protein interactions; neither sequence nor structure information is needed. The approach is evaluated by comparing the inferred interacting motifs with domain families defined for 368 proteins in the Structural Classification of Proteins (SCOP). The positive predictive value of the method for detecting proteins with common SCOP families is 75% at sensitivity of 10%. Most of the inferred interacting motifs were significantly associated with sequence patterns, which could be responsible for the common interactions. We find that yeast hubs with multiple interacting motifs are more likely to be essential than hubs with one or two interacting motifs, thus rationalizing the previously observed correlation between essentiality and the number of interacting partners of a protein. We also find that yeast hubs with multiple interacting motifs evolve slower than the average protein, contrary to the hubs with one or two interacting motifs. The proposed method will help us discover unknown interacting motifs and provide biological insights about protein hubs and their roles in interaction networks. Ramon Aragues, Andrej Sali, Jaume Bonet, Marc A. Martí-Renom, Baldomero Oliva |
PLoS Comput. Biol. | 2 |
| 2007 | Stereochemical Criteria for Prediction of the Effects of Proline Mutations on Protein StabilityabstractWhen incorporated into a polypeptide chain, proline (Pro) differs from all other naturally occurring amino acid residues in two important respects. The phi dihedral angle of Pro is constrained to values close to -65 degrees and Pro lacks an amide hydrogen. Consequently, mutations which result in introduction of Pro can significantly affect protein stability. In the present work, we describe a procedure to accurately predict the effect of Pro introduction on protein thermodynamic stability. Seventy-seven of the 97 non-Pro amino acid residues in the model protein, CcdB, were individually mutated to Pro, and the in vivo activity of each mutant was characterized. A decision tree to classify the mutation as perturbing or nonperturbing was created by correlating stereochemical properties of mutants to activity data. The stereochemical properties including main chain dihedral angle phi and main chain amide H-bonds (hydrogen bonds) were determined from 3D models of the mutant proteins built using MODELLER. We assessed the performance of the decision tree on a large dataset of 163 single-site Pro mutations of T4 lysozyme, 74 nsSNPs, and 52 other Pro substitutions from the literature. The overall accuracy of this algorithm was found to be 81% in the case of CcdB, 77% in the case of lysozyme, 76% in the case of nsSNPs, and 71% in the case of other Pro substitution data. The accuracy of Pro scanning mutagenesis for secondary structure assignment was also assessed and found to be at best 69%. Our prediction procedure will be useful in annotating uncharacterized nsSNPs of disease-associated proteins and for protein engineering and design. Kanika Bajaj, M. S. Madhusudhan 0001, Bharat V. Adkar, Purbani Chakrabarti, C. Ramakrishnan, Andrej Sali, Raghavan Varadarajan |
PLoS Comput. Biol. | 6 |
| 2007 | Functional Impact of Missense Variants in BRCA1 Predicted by Supervised LearningabstractMany individuals tested for inherited cancer susceptibility at the BRCA1 gene locus are discovered to have variants of unknown clinical significance (UCVs). Most UCVs cause a single amino acid residue (missense) change in the BRCA1 protein. They can be biochemically assayed, but such evaluations are time-consuming and labor-intensive. Computational methods that classify and suggest explanations for UCV impact on protein function can complement functional tests. Here we describe a supervised learning approach to classification of BRCA1 UCVs. Using a novel combination of 16 predictive features, the algorithms were applied to retrospectively classify the impact of 36 BRCA1 C-terminal (BRCT) domain UCVs biochemically assayed to measure transactivation function and to blindly classify 54 documented UCVs. Majority vote of three supervised learning algorithms is in agreement with the assay for more than 94% of the UCVs. Two UCVs found deleterious by both the assay and the classifiers reveal a previously uncharacterized putative binding site. Clinicians may soon be able to use computational classifiers such as those described here to better inform patients. These classifiers can be adapted to other cancer susceptibility genes and systematically applied to prioritize the growing number of potential causative loci and variants found by large-scale disease association studies. Rachel Karchin, Alvaro N. A. Monteiro, Sean V. Tavtigian, Marcelo A. Carvalho, Andrej Sali |
PLoS Comput. Biol. | 5 |
| 2006 | Structural Modeling of Protein Interactions by Analogy: Application to PSD-95abstractWe describe comparative patch analysis for modeling the structures of multidomain proteins and protein complexes, and apply it to the PSD-95 protein. Comparative patch analysis is a hybrid of comparative modeling based on a template complex and protein docking, with a greater applicability than comparative modeling and a higher accuracy than docking. It relies on structurally defined interactions of each of the complex components, or their homologs, with any other protein, irrespective of its fold. For each component, its known binding modes with other proteins of any fold are collected and expanded by the known binding modes of its homologs. These modes are then used to restrain conventional molecular docking, resulting in a set of binary domain complexes that are subsequently ranked by geometric complementarity and a statistical potential. The method is evaluated by predicting 20 binary complexes of known structure. It is able to correctly identify the binding mode in 70% of the benchmark complexes compared with 30% for protein docking. We applied comparative patch analysis to model the complex of the third PSD-95, DLG, and ZO-1 (PDZ) domain and the SH3-GK domains in the PSD-95 protein, whose structure is unknown. In the first predicted configuration of the domains, PDZ interacts with SH3, leaving both the GMP-binding site of guanylate kinase (GK) and the C-terminus binding cleft of PDZ accessible, while in the second configuration PDZ interacts with GK, burying both binding sites. We suggest that the two alternate configurations correspond to the different functional forms of PSD-95 and provide a possible structural description for the experimentally observed cooperative folding transitions in PSD-95 and its homologs. More generally, we expect that comparative patch analysis will provide useful spatial restraints for the structural characterization of an increasing number of binary and higher-order protein complexes. Dmitry Korkin, Fred P. Davis, Frank Alber, Tinh Luong, Min-Yi Shen, Vladan Lucic, Mary B. Kennedy, Andrej Sali |
PLoS Comput. Biol. | 8 |
| 2005 | PIBASE: a comprehensive database of structurally defined protein interfacesabstractMOTIVATION: In recent years, the Protein Data Bank (PDB) has experienced rapid growth. To maximize the utility of the high resolution protein-protein interaction data stored in the PDB, we have developed PIBASE, a comprehensive relational database of structurally defined interfaces between pairs of protein domains. It is composed of binary interfaces extracted from structures in the PDB and the Probable Quaternary Structure server using domain assignments from the Structural Classification of Proteins and CATH fold classification systems. RESULTS: PIBASE currently contains 158,915 interacting domain pairs between 105,061 domains from 2125 SCOP families. A diverse set of geometric, physiochemical and topologic properties are calculated for each complex, its domains, interfaces and binding sites. A subset of the interface properties are used to remove interface redundancy within PDB entries, resulting in 20,912 distinct domain-domain interfaces. The complexes are grouped into 989 topological classes based on their patterns of domain-domain contacts. The binary interfaces and their corresponding binding sites are categorized into 18,755 and 30,975 topological classes, respectively, based on the topology of secondary structure elements. The utility of the database is illustrated by outlining several current applications. AVAILABILITY: The database is accessible via the world wide web at http://salilab.org/pibase SUPPLEMENTARY INFORMATION: http://salilab.org/pibase/suppinfo.html. Fred P. Davis, Andrej Sali |
Bioinform. | 2 |
| 2005 | LS-SNP: large-scale annotation of coding non-synonymous SNPs based on multiple information sourcesabstractMOTIVATION: The NCBI dbSNP database lists over 9 million single nucleotide polymorphisms (SNPs) in the human genome, but currently contains limited annotation information. SNPs that result in amino acid residue changes (nsSNPs) are of critical importance in variation between individuals, including disease and drug sensitivity. RESULTS: We have developed LS-SNP, a genomic scale software pipeline to annotate nsSNPs. LS-SNP comprehensively maps nsSNPs onto protein sequences, functional pathways and comparative protein structure models, and predicts positions where nsSNPs destabilize proteins, interfere with the formation of domain-domain interfaces, have an effect on protein-ligand binding or severely impact human health. It currently annotates 28,043 validated SNPs that produce amino acid residue substitutions in human proteins from the SwissProt/TrEMBL database. Annotations can be viewed via a web interface either in the context of a genomic region or by selecting sets of SNPs, genes, proteins or pathways. These results are useful for identifying candidate functional SNPs within a gene, haplotype or pathway and in probing molecular mechanisms responsible for functional impacts of nsSNPs. AVAILABILITY: http://www.salilab.org/LS-SNP CONTACT: [email protected] SUPPLEMENTARY INFORMATION: http://salilab.org/LS-SNP/supp-info.pdf. Rachel Karchin, Mark Diekhans, Libusha Kelly, Daryl J. Thomas, Ursula Pieper 0001, Narayanan Eswar, David Haussler, Andrej Sali |
Bioinform. | 8 |
| 2003 | ModLoop: automated modeling of loops in protein structuresabstractSUMMARY: ModLoop is a web server for automated modeling of loops in protein structures. The input is the atomic coordinates of the protein structure in the Protein Data Bank format, and the specification of the starting and ending residues of one or more segments to be modeled, containing no more than 20 residues in total. The output is the coordinates of the non-hydrogen atoms in the modeled segments. A user provides the input to the server via a simple web interface, and receives the output by e-mail. The server relies on the loop modeling routine in MODELLER that predicts the loop conformations by satisfaction of spatial restraints, without relying on a database of known protein structures. For a rapid response, ModLoop runs on a cluster of Linux PC computers. AVAILABILITY: The server is freely accessible to academic users at http://salilab.org/modloop András Fiser, Andrej Sali |
Bioinform. | 2 |
| 2003 | ModView, visualization of multiple protein sequences and structuresabstractSUMMARY: We describe ModView, a web application for visualization of multiple protein sequences and structures. ModView integrates a multiple structure viewer, a multiple sequence alignment editor, and a database querying engine. It is possible to interactively manipulate hundreds of proteins, to visualize conservative and variable residues, active and binding sites, fragments, and domains in protein families, as well as to display large macromolecular complexes such as ribosomes or viruses. As a Netscape plug-in, ModView can be included in HTML pages along with text and figures, which makes it useful for teaching and presentations. ModView is also suitable as a graphical interface to various databases because it can be controlled through JavaScript commands and called from CGI scripts. AVAILABILITY: ModView is available at http://guitar.rockefeller.edu/modview. Valentin A. Ilyin, Ursula Pieper 0001, Ashley C. Stuart, Marc A. Martí-Renom, Linda McMahan, Andrej Sali |
Bioinform. | 6 |
| 2002 | LigBase: a database of families of aligned ligand binding sites in known protein sequences and structuresabstractA database comprising all ligand-binding sites of known structure aligned with all related protein sequences and structures is described. Currently, the database contains approximately 50000 ligand-binding sites for small molecules found in the Protein Data Bank (PDB). The structure-structure alignments are obtained by the Combinatorial Extension (CE) program (Shindyalov and Bourne, Protein Eng., 11, 739-747, 1998) and sequence-structure alignments are extracted from the ModBase database of comparative protein structure models for all known protein sequences (Sanchez et al., Nucleic Acids Res., 28, 250-253, 2000). It is possible to search for binding sites in LigBase by a variety of criteria. LigBase reports summarize ligand data including relevant structural information from the PDB file, such as ligand type and size, and contain links to all related protein sequences in the TrEMBL database. Residues in the binding sites are graphically depicted for comparison with other structurally defined family members. LigBase provides a resource for the analysis of families of related binding sites. Ashley C. Stuart, Valentin A. Ilyin, Andrej Sali |
Bioinform. | 3 |
| 2001 | The Third Georgia Tech-Emory International Conference on Bioinformatics: In Silico Biology; Bioinformatics After Human Genome (November 15-18, 2001, Atlanta, Georgia, USA)abstractMark Borodovsky, Eugene Koonin, Chris Burge, Jim Fickett, John Logsdon, Andrej Sali, Gary Stormo, Igor Zhulin; THE THIRD GEORGIA TECH–EMORY INTERNATIONAL C Mark Borodovsky, Eugene V. Koonin, Chris Burge, James W. Fickett, John Logsdon, Andrej Sali, Gary D. Stormo, Igor B. Zhulin |
Bioinform. | 6 |
| 2001 | EVA: continuous automatic evaluation of protein structure prediction serversabstractUNLABELLED: Evaluation of protein structure prediction methods is difficult and time-consuming. Here, we describe EVA, a web server for assessing protein structure prediction methods, in an automated, continuous and large-scale fashion. Currently, EVA evaluates the performance of a variety of prediction methods available through the internet. Every week, the sequences of the latest experimentally determined protein structures are sent to prediction servers, results are collected, performance is evaluated, and a summary is published on the web. EVA has so far collected data for more than 3000 protein chains. These results may provide valuable insight to both developers and users of prediction methods. AVAILABILITY: http://cubic.bioc.columbia.edu/eva. CONTACT: [email protected] Volker A. Eyrich, Marc A. Martí-Renom, Dariusz Przybylski, M. S. Madhusudhan 0001, András Fiser, Florencio Pazos, Alfonso Valencia, Andrej Sali, Burkhard Rost |
Bioinform. | 8 |
| 2001 | DBAli: a database of protein structure alignmentsabstractSUMMARY: The DBAli database includes approximately 35000 alignments of pairs of protein structures from SCOP (Lo Conte et al., Nucleic Acids Res., 28, 257-259, 2000) and CE (Shindyalov and Bourne, Protein Eng., 11, 739-747, 1998). DBAli is linked to several resources, including Compare3D (Shindyalov and Bourne, http://www.sdsc.edu/pb/software.htm, 1999) and ModView (Ilyin and Sali, http://guitar.rockefeller.edu/ModView/, 2001) for visualizing sequence alignments and structure superpositions. A flexible search of DBAli by protein sequence and structure properties allows construction of subsets of alignments suitable for a number of applications, such as benchmarking of sequence-sequence and sequence-structure alignment methods under a variety of conditions. AVAILABILITY: http://guitar.rockefeller.edu/DBAli/ Marc A. Martí-Renom, Valentin A. Ilyin, Andrej Sali |
Bioinform. | 3 |
| 1999 | The Second Georgia Tech International Conference on Bioinformatics: Sequence, Structure and Function (November 11-14, 1999, Atlanta, Georgia, USA)abstractThis issue of Bioinformatics contains reports on selected papers presented at the international conference in Atlanta.The conference was held at one of the midtown hotels offering a magnificent bird's-eye view to the cosmopolitan capital of the Southeast of the USA.The conference agenda included keynote lectures by Russell Doolittle, Walter Fitch and Walter Gilbert, 19 plenary and contributed talks and more than 50 poster presentations.The focus of the conference was on the most intriguing issues of comparative and structural genomics, functional interpretation of nucleic acid and protein sequences, inferring 3D structure from sequence data and deciphering the history of biomolecular evolution (http://exon.biology.gatech.edu/conference).Submission of manuscripts to the Bioinformatics special issue has been a parallel avenue of presentation for the conference participants.Eighteen papers were submitted to the peer review process and nine were accepted for publication.The issue starts with the papers devoted to DNA sequence analysis and gene finding, followed by the papers on protein sequence analysis and protein structure prediction.Single pass sequencing of the 5 and 3 ends of cloned cDNA sequences represents an efficient approach to large scale characterization of the set of genes expressed by an organism.In the past few years huge numbers of such sequences, termed 'expressed sequence tags' or ESTs, have been determined from human and a variety of other organisms, and deposited in public databases.Such sequences represent a powerful resource for identifying genes, but suffer from the relatively high rates of sequencing errors typical of single pass data.Frameshift errors, in particular, can obscure the results of similarity searches based on translations of putative EST ORFs.The paper entitled 'FramePlus: Aligning DNA to protein sequences' by E. Halperin, S. Faigler and R. Gill-More describes a new algorithm specifically tailored to allow for moderate to high rates of frameshift errors.The new techniques resulted in improvements in the ability to detect homologies between ESTs and known proteins in peptide database searches. Chris Burge, Andrej Sali, Mark Borodovsky |
Bioinform. | 2 |