Nick V. Grishin

dblp:18/6563 · DBLP profile ↗
← Back
39ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0003-4108-1153ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 39 · 4 since 2021
YearPublicationVenuePosition
2026 ECOD: Classification of domains in AFDB Swiss-Prot structure predictions
abstract
The development of highly accurate protein structure prediction algorithms has led to an explosion of structural data, transforming our understanding of protein structure-function relationships across diverse organisms. Domain classifications such as the Evolutionary Classification of Protein Domains (ECOD) have incorporated these computational predictions alongside experimental structures to create comprehensive resources for the research community. The AlphaFold Protein Structure Database (AFDB) plays a unique role, providing millions of predicted structures that ECOD has systematically classified for human proteins, small pathogens, and reference proteomes. Here, we extend this classification framework to the UniProtKB/Swiss-Prot dataset, applying the Domain Parser for AlphaFold Models (DPAM) pipeline to classify domains from over 542,000 Swiss-Prot protein structure predictions, resulting in more than 1,032,000 classified domains. These domains span 3,493 ECOD topologies and display high assignment confidence (mean DPAM probability: 0.992), with extensive taxonomic and functional diversity. Notably, over 100,000 domains lack existing Pfam mappings, reflecting the extended sensitivity of structure-based classification and identifying domain groups not yet captured by sequence-based profiles. These results significantly expand ECOD's coverage into a functionally and taxonomically diverse protein space, anchoring high-confidence structure predictions in an evolutionary framework. By integrating Swiss-Prot predictions, we enhance the utility and interpretability of AlphaFold models and establish a foundation for future large-scale, functionally informed domain classifications.
R. Dustin Schaeffer, Jing Zhang 0115, Qian Cong, Nick V. Grishin
PLoS Comput. Biol.4
2025 Using evolutionary context to classify difficult protein folds
abstract
Recent advances in protein structure prediction, such as AlphaFold2, have enabled identification of vast numbers of putative novel protein domains across the sequence space, many of which adopt structures dissimilar to known folds. Based on structural segmentation and classification, the Encyclopedia of Domains (TED) project recently cataloged more than 7400 low-symmetry, structure-based domains as candidate novel-fold (CNF) domains. To place these domains in their broader evolutionary and structural context, we applied DPAM (Domain Parser for AlphaFold Models), a complementary method that combines AlphaFold-derived confidence metrics with sensitive sequence and structure similarity searches, to parse domains for the AlphaFold models of proteins containing TED CNF domains. We identified 8044 DPAM domains with significant overlap with TED CNF domains, among which 2490 were confidently assigned to entries in the ECOD (Evolutionary Classification of protein Domains) structural classification hierarchy. Our results suggest that a substantial subset of TED candidate novel-fold domains are distant homologs of existing ECOD domains. Comparison of domain boundaries between TED and DPAM showed varied patterns: more than one-third of cases featured TED CNF domains largely embedded within DPAM domains-often representing insertions or extensions into enzymatic or repeat folds. A smaller fraction (17%) exhibited consistent domain boundaries between TED and DPAM. These consistently defined domains are often characterized by significant structural diversity, including long insertions and duplications. An even smaller subset showed the reverse relationship, with DPAM domains largely embedded within TED CNF domains. In these cases, DPAM effectively separated multiple structural units that TED grouped as single domains. Together, these findings highlight the complementarity of structural and evolutionary approaches for domain annotation and demonstrate the power of integrative methods, such as DPAM, in refining the classification of challenging protein folds and uncovering distant evolutionary relationships.
Jimin Pei, R. Dustin Schaeffer, Qian Cong, Nick V. Grishin
PLoS Comput. Biol.4
2024 ECOD domain classification of 48 whole proteomes from AlphaFold Structure Database using DPAM2
abstract
Protein structure prediction has now been deployed widely across several different large protein sets. Large-scale domain annotation of these predictions can aid in the development of biological insights. Using our Evolutionary Classification of Protein Domains (ECOD) from experimental structures as a basis for classification, we describe the detection and cataloging of domains from 48 whole proteomes deposited in the AlphaFold Database. On average, we can provide positive classification (either of domains or other identifiable non-domain regions) for 90% of residues in all proteomes. We classified 746,349 domains from 536,808 proteins comprised of over 226,424,000 amino acid residues. We examine the varying populations of homologous groups in both eukaryotes and bacteria. In addition to containing a higher fraction of disordered regions and unassigned domains, eukaryotes show a higher proportion of repeated proteins, both globular and small repeats. We enumerate those highly populated domains that are shared in both eukaryotes and bacteria, such as the Rossmann domains, TIM barrels, and P-loop domains. Additionally, we compare the sampling of homologous groups from this whole proteome set against our stable ECOD reference and discuss groups that have been enriched by structure predictions. Finally, we discuss the implication of these results for protein target selection for future classification strategies for very large protein sets.
R. Dustin Schaeffer, Jing Zhang 0115, Kirill E. Medvedev, Lisa N. Kinch, Qian Cong, Nick V. Grishin
PLoS Comput. Biol.6
2022 DisEnrich: database of enriched regions in human dark proteome
abstract
MOTIVATION: Intrinsically disordered proteins (IDPs) are involved in numerous processes crucial for living organisms. Bias in amino acid composition of these proteins determines their unique biophysical and functional features. Distinct intrinsically disordered regions (IDRs) with compositional bias play different important roles in various biological processes. IDRs enriched in particular amino acids in human proteome have not been described consistently. RESULTS: We developed DisEnrich-the database of human proteome IDRs that are significantly enriched in particular amino acids. Each human protein is described using Gene Ontology (GO) function terms, disorder prediction for the full-length sequence using three methods, enriched IDR composition and ranks of human proteins with similar enriched IDRs. Distribution analysis of enriched IDRs among broad functional categories revealed significant overrepresentation of R- and Y-enriched IDRs in metabolic and enzymatic activities and F-enriched IDRs in transport. About 75% of functional categories contain IDPs with IDRs significantly enriched in hydrophobic residues that are important for protein-protein interactions. AVAILABILITY AND IMPLEMENTATION: The database is available at http://prodata.swmed.edu/DisEnrichDB/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics Advances online.
Kirill E. Medvedev, Jimin Pei, Nick V. Grishin
Bioinform.3
2020 pCRM1exportome: database of predicted CRM1-dependent Nuclear Export Signal (NES) motifs in cancer-related genes
abstract
MOTIVATION: The consensus pattern of Nuclear Export Signal (NES) is a short sequence motif that is commonly identified in protein sequences, whether the motif acts as an NES (true positive) or not (false positive). Finding more plausible NES functioning regions among the vast array of consensus-matching segments would provide an interesting resource for further experimental validation. Better defined NES should also allow meaningful mapping of cancer-related mutation positions, leading to plausible explanations for the relationship between nuclear export and disease. RESULTS: Possible NES candidate regions are extracted from the cancer-related human reference proteome. Extracted NES are scored for reliability by combining sequence-based and structure-based approaches. The confidently identified NES candidate motifs were checked for overlap with cancer-related mutation positions annotated in the COSMIC database. Among the ∼700 cancer-related sequences in the COSMIC Cancer Gene Census, 178 sequences are predicted to have possible NES motifs containing cancer-related mutations at their key positions. These lists are organized into our database (pCRM1exportome), and other protein sequences in the human reference proteome can also be retrieved by their UniProt IDs. AVAILABILITY AND IMPLEMENTATION: The database is freely available at http://prodata.swmed.edu/pCRM1exportome. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jordan M. Baumhardt, Jimin Pei, Yuh Min Chook, Nick V. Grishin
Bioinform.5
2020 Mutation severity spectrum of rare alleles in the human genome is predictive of disease type
abstract
The human genome harbors a variety of genetic variations. Single-nucleotide changes that alter amino acids in protein-coding regions are one of the major causes of human phenotypic variation and diseases. These single-amino acid variations (SAVs) are routinely found in whole genome and exome sequencing. Evaluating the functional impact of such genomic alterations is crucial for diagnosis of genetic disorders. We developed DeepSAV, a deep-learning convolutional neural network to differentiate disease-causing and benign SAVs based on a variety of protein sequence, structural and functional properties. Our method outperforms most stand-alone programs, and the version incorporating population and gene-level information (DeepSAV+PG) has similar predictive power as some of the best available. We transformed DeepSAV scores of rare SAVs in the human population into a quantity termed "mutation severity measure" for each human protein-coding gene. It reflects a gene's tolerance to deleterious missense mutations and serves as a useful tool to study gene-disease associations. Genes implicated in cancer, autism, and viral interaction are found by this measure as intolerant to mutations, while genes associated with a number of other diseases are scored as tolerant. Among known disease-associated genes, those that are mutation-intolerant are likely to function in development and signal transduction pathways, while those that are mutation-tolerant tend to encode metabolic and mitochondrial proteins.
Jimin Pei, Lisa N. Kinch, Zbyszek Otwinowski, Nick V. Grishin
PLoS Comput. Biol.4
2019 Functional analysis of Rossmann-like domains reveals convergent evolution of topology and reaction pathways
abstract
Rossmann folds are ancient, frequently diverged domains found in many biological reaction pathways where they have adapted for different functions. Consequently, discernment and classification of their homologous relations and function can be complicated. We define a minimal Rossmann-like structure motif (RLM) that corresponds for the common core of known Rossmann domains and use this motif to identify all RLM domains in the Protein Data Bank (PDB), thus finding they constitute about 20% of all known 3D structures. The Evolutionary Classification of protein structure Domains (ECOD) classifies RLM domains in a number of groups that lack evidence for homology (X-groups), which suggests that they could have evolved independently multiple times. Closely related, homologous RLM enzyme families can diverge to bind different ligands using similar binding sites and to catalyze different reactions. Conversely, non-homologous RLM domains can converge to catalyze the same reactions or to bind the same ligand with alternate binding modes. We discuss a special case of such convergent evolution that is relevant to the polypharmacology paradigm, wherein the same drug (methotrexate) binds to multiple non-homologous RLM drug targets with different topologies. Finally, assigning proteins with RLM domain to the Enzyme Commission classification suggest that RLM enzymes function mainly in metabolism (and comprise 38% of reference metabolic pathways) and are overrepresented in extant pathways that represent ancient biosynthetic routes such as nucleotide metabolism, energy metabolism, and metabolism of amino acids. In fact, RLM enzymes take part in five out of eight enzymatic reactions of the Wood-Ljungdahl metabolic pathway thought to be used by the last universal common ancestor (LUCA). The prevalence of RLM domains in this ancient metabolism might explain their wide distribution among enzymes.
Kirill E. Medvedev, Lisa N. Kinch, R. Dustin Schaeffer, Nick V. Grishin
PLoS Comput. Biol.4
2018 A sequence family database built on ECOD structural domains
abstract
Motivation: The ECOD database classifies protein domains based on their evolutionary relationships, considering both remote and close homology. The family group in ECOD provides classification of domains that are closely related to each other based on sequence similarity. Due to different perspectives on domain definition, direct application of existing sequence domain databases, such as Pfam, to ECOD struggles with several shortcomings. Results: We created multiple sequence alignments and profiles from ECOD domains with the help of structural information in alignment building and boundary delineation. We validated the alignment quality by scoring structure superposition to demonstrate that they are comparable to curated seed alignments in Pfam. Comparison to Pfam and CDD reveals that 27 and 16% of ECOD families are new, but they are also dominated by small families, likely because of the sampling bias from the PDB database. There are 35 and 48% of families whose boundaries are modified comparing to counterparts in Pfam and CDD, respectively. Availability and implementation: The new families are now integrated in the ECOD website. The aggregate HMMER profile library and alignment are available for download on ECOD website (http://prodata.swmed.edu/ecod). Supplementary information: Supplementary data are available at Bioinformatics online.
Yuxing Liao, R. Dustin Schaeffer, Jimin Pei, Nick V. Grishin
Bioinform.4
2016 Phylogeny Reconstruction with Alignment-Free Method That Corrects for Horizontal Gene Transfer
abstract
Advances in sequencing have generated a large number of complete genomes. Traditionally, phylogenetic analysis relies on alignments of orthologs, but defining orthologs and separating them from paralogs is a complex task that may not always be suited to the large datasets of the future. An alternative to traditional, alignment-based approaches are whole-genome, alignment-free methods. These methods are scalable and require minimal manual intervention. We developed SlopeTree, a new alignment-free method that estimates evolutionary distances by measuring the decay of exact substring matches as a function of match length. SlopeTree corrects for horizontal gene transfer, for composition variation and low complexity sequences, and for branch-length nonlinearity caused by multiple mutations at the same site. We tested SlopeTree on 495 bacteria, 73 archaea, and 72 strains of Escherichia coli and Shigella. We compared our trees to the NCBI taxonomy, to trees based on concatenated alignments, and to trees produced by other alignment-free methods. The results were consistent with current knowledge about prokaryotic evolution. We assessed differences in tree topology over different methods and settings and found that the majority of bacteria and archaea have a core set of proteins that evolves by descent. In trees built from complete genomes rather than sets of core genes, we observed some grouping by phenotype rather than phylogeny, for instance with a cluster of sulfur-reducing thermophilic bacteria coming together irrespective of their phyla. The source-code for SlopeTree is available at: http://prodata.swmed.edu/download/pub/slopetree_v1/slopetree.tar.gz.
Raquel Bromberg, Nick V. Grishin, Zbyszek Otwinowski
PLoS Comput. Biol.2
2015 LocNES: a computational tool for locating classical NESs in CRM1 cargo proteins
abstract
MOTIVATION: Classical nuclear export signals (NESs) are short cognate peptides that direct proteins out of the nucleus via the CRM1-mediated export pathway. CRM1 regulates the localization of hundreds of macromolecules involved in various cellular functions and diseases. Due to the diverse and complex nature of NESs, reliable prediction of the signal remains a challenge despite several attempts made in the last decade. RESULTS: We present a new NES predictor, LocNES. LocNES scans query proteins for NES consensus-fitting peptides and assigns these peptides probability scores using Support Vector Machine model, whose feature set includes amino acid sequence, disorder propensity, and the rank of position-specific scoring matrix score. LocNES demonstrates both higher sensitivity and precision over existing NES prediction tools upon comparative analysis using experimentally identified NESs. AVAILABILITY AND IMPLEMENTATION: LocNES is freely available at http://prodata.swmed.edu/LocNES CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Darui Xu, Kara Marquis, Jimin Pei, Szu-Chin Fu, Tolga Cagatay, Nick V. Grishin, Yuh Min Chook
Bioinform.6
2015 SFESA: a web server for pairwise alignment refinement by secondary structure shifts
abstract
BACKGROUND: Protein sequence alignment is essential for a variety of tasks such as homology modeling and active site prediction. Alignment errors remain the main cause of low-quality structure models. A bioinformatics tool to refine alignments is needed to make protein alignments more accurate. RESULTS: We developed the SFESA web server to refine pairwise protein sequence alignments. Compared to the previous version of SFESA, which required a set of 3D coordinates for a protein, the new server will search a sequence database for the closest homolog with an available 3D structure to be used as a template. For each alignment block defined by secondary structure elements in the template, SFESA evaluates alignment variants generated by local shifts and selects the best-scoring alignment variant. A scoring function that combines the sequence score of profile-profile comparison and the structure score of template-derived contact energy is used for evaluation of alignments. PROMALS pairwise alignments refined by SFESA are more accurate than those produced by current advanced alignment methods such as HHpred and CNFpred. In addition, SFESA also improves alignments generated by other software. CONCLUSIONS: SFESA is a web-based tool for alignment refinement, designed for researchers to compute, refine, and evaluate pairwise alignments with a combined sequence and structure scoring of alignment blocks. To our knowledge, the SFESA web server is the only tool that refines alignments by evaluating local shifts of secondary structure elements. The SFESA web server is available at http://prodata.swmed.edu/sfesa.
Jing Tong, Jimin Pei, Nick V. Grishin
BMC Bioinform.3
2014 ECOD: An Evolutionary Classification of Protein Domains
abstract
Understanding the evolution of a protein, including both close and distant relationships, often reveals insight into its structure and function. Fast and easy access to such up-to-date information facilitates research. We have developed a hierarchical evolutionary classification of all proteins with experimentally determined spatial structures, and presented it as an interactive and updatable online database. ECOD (Evolutionary Classification of protein Domains) is distinct from other structural classifications in that it groups domains primarily by evolutionary relationships (homology), rather than topology (or "fold"). This distinction highlights cases of homology between domains of differing topology to aid in understanding of protein structure evolution. ECOD uniquely emphasizes distantly related homologs that are difficult to detect, and thus catalogs the largest number of evolutionary links among structural domain classifications. Placing distant homologs together underscores the ancestral similarities of these proteins and draws attention to the most important regions of sequence and structure, as well as conserved functional sites. ECOD also recognizes closer sequence-based relationships between protein domains. Currently, approximately 100,000 protein structures are classified in ECOD into 9,000 sequence families clustered into close to 2,000 evolutionary groups. The classification is assisted by an automated pipeline that quickly and consistently classifies weekly releases of PDB structures and allows for continual updates. This synchronization with PDB uniquely distinguishes ECOD among all protein classifications. Finally, we present several case studies of homologous proteins not recorded in other classifications, illustrating the potential of how ECOD can be used to further biological and evolutionary studies.
R. Dustin Schaeffer, Yuxing Liao, Lisa N. Kinch, Jimin Pei, Shuoyong Shi, Bong-Hyun Kim, Nick V. Grishin
PLoS Comput. Biol.8
2013 Defining and predicting structurally conserved regions in protein superfamilies
abstract
MOTIVATION: The structures of homologous proteins are generally better conserved than their sequences. This phenomenon is demonstrated by the prevalence of structurally conserved regions (SCRs) even in highly divergent protein families. Defining SCRs requires the comparison of two or more homologous structures and is affected by their availability and divergence, and our ability to deduce structurally equivalent positions among them. In the absence of multiple homologous structures, it is necessary to predict SCRs of a protein using information from only a set of homologous sequences and (if available) a single structure. Accurate SCR predictions can benefit homology modelling and sequence alignment. RESULTS: Using pairwise DaliLite alignments among a set of homologous structures, we devised a simple measure of structural conservation, termed structural conservation index (SCI). SCI was used to distinguish SCRs from non-SCRs. A database of SCRs was compiled from 386 SCOP superfamilies containing 6489 protein domains. Artificial neural networks were then trained to predict SCRs with various features deduced from a single structure and homologous sequences. Assessment of the predictions via a 5-fold cross-validation method revealed that predictions based on features derived from a single structure perform similarly to ones based on homologous sequences, while combining sequence and structural features was optimal in terms of accuracy (0.755) and Matthews correlation coefficient (0.476). These results suggest that even without information from multiple structures, it is still possible to effectively predict SCRs for a protein. Finally, inspection of the structures with the worst predictions pinpoints difficulties in SCR definitions. AVAILABILITY: The SCR database and the prediction server can be found at http://prodata.swmed.edu/SCR. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics Online.
Ivan K. Huang, Jimin Pei, Nick V. Grishin
Bioinform.3
2013 M2SG: mapping human disease-related genetic variants to protein sequences and genomic loci
abstract
SUMMARY: Online Mendelian Inheritance in Man (OMIM) is a manually curated compendium of human genetic variants and the corresponding phenotypes, mostly human diseases. Instead of directly documenting the native sequences for gene entries, OMIM links its entries to protein and DNA sequences in other databases. However, because of the existence of gene isoforms and errors in OMIM records, mapping a specific OMIM mutation to its corresponding protein sequence is not trivial. Combining computer programs and extensive manual curation of OMIM full-text descriptions and original literature, we mapped 98% of OMIM amino acid substitutions (AASs) and all SwissProt Variant (SwissVar) disease-related AASs to reference sequences and confidently mapped 99.96% of all AASs to the genomic loci. Based on the results, we developed an online database and interactive web server (M2SG) to (i) retrieve the mapped OMIM and SwissVar variants for a given protein sequence; and (ii) obtain related proteins and mutations for an input disease phenotype. This database will be useful for analyzing sequences, understanding the effect of mutations, identifying important genetic variations and designing experiments on a protein of interest. AVAILABILITY AND IMPLEMENTATION: The database and web server are freely available at http://prodata.swmed.edu/M2S/mut2seq.cgi.
Renkai Ji, Qian Cong, Nick V. Grishin
Bioinform.4
2013 Pclust: protein network visualization highlighting experimental data
abstract
SUMMARY: One approach to infer functions of new proteins from their homologs utilizes visualization of an all-against-all pairwise similarity network (A2ApsN) that exploits the speed of BLAST and avoids the complexity of multiple sequence alignment. However, identifying functions of the protein clusters in A2ApsN is never trivial, due to a lack of linking characterized proteins to their relevant information in current software packages. Given the database errors introduced by automatic annotation transfer, functional deduction should be made from proteins with experimental studies, i.e. 'reference proteins'. Here, we present a web server, termed Pclust, which provides a user-friendly interface to visualize the A2ApsN, placing emphasis on such 'reference proteins' and providing access to their full information in source databases, e.g. articles in PubMed. The identification of 'reference proteins' and the ease of cross-database linkage will facilitate understanding the functions of protein clusters in the network, thus promoting interpretation of proteins of interest. AVAILABILITY: The Pclust server is freely available at http://prodata.swmed.edu/pclust
Lisa N. Kinch, Nick V. Grishin
Bioinform.3
2013 Seq2Ref: a web server to facilitate functional interpretation
abstract
BACKGROUND: The size of the protein sequence database has been exponentially increasing due to advances in genome sequencing. However, experimentally characterized proteins only constitute a small portion of the database, such that the majority of sequences have been annotated by computational approaches. Current automatic annotation pipelines inevitably introduce errors, making the annotations unreliable. Instead of such error-prone automatic annotations, functional interpretation should rely on annotations of 'reference proteins' that have been experimentally characterized or manually curated. RESULTS: The Seq2Ref server uses BLAST to detect proteins homologous to a query sequence and identifies the reference proteins among them. Seq2Ref then reports publications with experimental characterizations of the identified reference proteins that might be relevant to the query. Furthermore, a plurality-based rating system is developed to evaluate the homologous relationships and rank the reference proteins by their relevance to the query. CONCLUSIONS: The reference proteins detected by our server will lend insight into proteins of unknown function and provide extensive information to develop in-depth understanding of uncharacterized proteins. Seq2Ref is available at: http://prodata.swmed.edu/seq2ref.
Qian Cong, Lisa N. Kinch, Nick V. Grishin
BMC Bioinform.4
2012 Self consistency grouping: a stringent clustering method
abstract
BACKGROUND: Numerous types of clustering like single linkage and K-means have been widely studied and applied to a variety of scientific problems. However, the existing methods are not readily applicable for the problems that demand high stringency. METHODS: Our method, self consistency grouping, i.e. SCG, yields clusters whose members are closer in rank to each other than to any member outside the cluster. We do not define a distance metric; we use the best known distance metric and presume that it measures the correct distance. SCG does not impose any restriction on the size or the number of the clusters that it finds. The boundaries of clusters are determined by the inconsistencies in the ranks. In addition to the direct implementation that finds the complete structure of the (sub)clusters we implemented two faster versions. The fastest version is guaranteed to find only the clusters that are not subclusters of any other clusters and the other version yields the same output as the direct implementation but does so more efficiently. RESULTS: Our tests have demonstrated that SCG yields very few false positives. This was accomplished by introducing errors in the distance measurement. Clustering of protein domain representatives by structural similarity showed that SCG could recover homologous groups with high precision. CONCLUSIONS: SCG has potential for finding biological relationships under stringent conditions.
Bong-Hyun Kim, Bhadrachalam Chitturi, Nick V. Grishin
BMC Bioinform.3
2011 An automatic method for CASP9 free modeling structure prediction assessment
abstract
MOTIVATION: Manual inspection has been applied to and is well accepted for assessing critical assessment of protein structure prediction (CASP) free modeling (FM) category predictions over the years. Such manual assessment requires expertise and significant time investment, yet has the problems of being subjective and unable to differentiate models of similar quality. It is beneficial to incorporate the ideas behind manual inspection to an automatic score system, which could provide objective and reproducible assessment of structure models. RESULTS: Inspired by our experience in CASP9 FM category assessment, we developed an automatic superimposition independent method named Quality Control Score (QCS) for structure prediction assessment. QCS captures both global and local structural features, with emphasis on global topology. We applied this method to all FM targets from CASP9, and overall the results showed the best agreement with Manual Inspection Scores among automatic prediction assessment methods previously applied in CASPs, such as Global Distance Test Total Score (GDT_TS) and Contact Score (CS). As one of the important components to guide our assessment of CASP9 FM category predictions, this method correlates well with other scoring methods and yet is able to reveal good-quality models that are missed by GDT_TS. AVAILABILITY: The script for QCS calculation is available at http://prodata.swmed.edu/QCS/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Qian Cong, Lisa N. Kinch, Jimin Pei, Shuoyong Shi, Vyacheslav N. Grishin, Nick V. Grishin
Bioinform.7
2010 Structural Differences between Proteins with Similar Sequences
abstract
Similarity between protein sequences is usually predictive of similarity in structures. However, in some rare cases protein domains with significant sequence similarity adopt different structures. Here, we carry out a survey of protein domain pairs with high sequence similarity (measured by HHsearch probability) and low structural similarity (measured by Dali Z-score), aiming to identify the reasons for this discordance. Besides methodological problems with either sequences or structures of domains, we find and describe novel examples of homologs with structural changes.
Qian Cong, Bong-Hyun Kim, Lisa N. Kinch, Nick V. Grishin
BIBE4
2010 HangOut: generating clean PSI-BLAST profiles for domains with long insertions
abstract
UNLABELLED: Profile-based similarity search is an essential step in structure-function studies of proteins. However, inclusion of non-homologous sequence segments into a profile causes its corruption and results in false positives. Profile corruption is common in multidomain proteins, and single domains with long insertions are a significant source of errors. We developed a procedure (HangOut) that, for a single domain with specified insertion position, cleans erroneously extended PSI-BLAST alignments to generate better profiles. AVAILABILITY: HangOut is implemented in Python 2.3 and runs on all Unix-compatible platforms. The source code is available under the GNU GPL license at http://prodata.swmed.edu/HangOut/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Bong-Hyun Kim, Qian Cong, Nick V. Grishin
Bioinform.3
2009 Phenotypic categorization of genetic skin diseases reveals new relations between phenotypes, genes and pathways
abstract
MOTIVATION: Systematic analysis of connection between proteins, their cellular function and phenotypic manifestations in disease is a central problem of biological and clinical research. The solution to this problem requires the development of new approaches to link the rapidly growing dataset of gene-disease associations with the many complex and overlapping phenotypes of human disease. RESULTS: We analyze genetic skin disorders and suggest a manually designed set of elementary phenotypes whose combinations define diseases as points in a multidimensional space, providing a basis for phenotypic disease clustering. Placing the known gene-disease associations in the context of this space reveals new patterns that suggest previously unknown functional links between proteins, signaling pathways and disease phenotypes. For example, analysis of telangiectasias (spider vein diseases) reveals a previously unrecognized interplay between the TGF-beta signaling pathway and pentose phosphate pathway. This interaction may mediate glucose-dependent regulation of TGF-beta signaling, providing a clue to the known association between angiopathies and diabetes and implying new gene candidates for mutational analysis and drug targeting.
Ruslan Sadreyev, Jamison D. Feramisco, Hensin Tsao, Nick V. Grishin
Bioinform.4
2009 Structure similarity measure with penalty for close non-equivalent residues
abstract
MOTIVATION: Recent improvement in homology-based structure modeling emphasizes the importance of sensitive evaluation measures that help identify and correct modest distortions in models compared with the target structures. Global Distance Test Total Score (GDT_TS), otherwise a very powerful and effective measure for model evaluation, is still insensitive to and can even reward such distortions, as observed for remote homology modeling in the latest CASP8 (Comparative Assessment of Structure Prediction). RESULTS: We develop a new measure that balances GDT_TS reward for the closeness of equivalent model and target residues ('attraction' term) with the penalty for the closeness of non-equivalent residues ('repulsion' term). Compared with GDT_TS, the resulting score, TR (total score with repulsion), is much more sensitive to structure compression both in real remote homologs and in CASP models. TR is correlated yet different from other measures of structure similarity. The largest difference from GDT_TS is observed in models of mid-range quality based on remote homology modeling. AVAILABILITY: The script for TR calculation is included in Supplementary Material. TR scores for all server models in CASP8 are available at http://prodata.swmed.edu/CASP8.
Ruslan Sadreyev, Shuoyong Shi, David Baker 0001, Nick V. Grishin
Bioinform.4
2009 PROCAIN server for remote protein sequence similarity search
abstract
UNLABELLED: Sensitive and accurate detection of distant protein homology is essential for the studies of protein structure, function and evolution. We recently developed PROCAIN, a method that is based on sequence profile comparison and involves the analysis of four signals--similarities of residue content at the profile positions combined with three types of assisting information: sequence motifs, residue conservation and predicted secondary structure. Here we present the PROCAIN web server that allows the user to submit a query sequence or multiple sequence alignment and perform the search in a profile database of choice. The output is structured similar to that of BLAST, with the list of detected homologs sorted by E-value and followed by profile-profile alignments. The front page allows the user to adjust multiple options of input processing and output formatting, as well as search settings, including the relative weights assigned to the three types of assisting information. AVAILABILITY: http://prodata.swmed.edu/procain/.
Yong Wang 0025, Ruslan Sadreyev, Nick V. Grishin
Bioinform.3
2009 Considering scores between unrelated proteins in the search database improves profile comparison
abstract
BACKGROUND: Profile-based comparison of multiple sequence alignments is a powerful methodology for the detection remote protein sequence similarity, which is essential for the inference and analysis of protein structure, function, and evolution. Accurate estimation of statistical significance of detected profile similarities is essential for further development of this methodology. Here we analyze a novel approach to estimate the statistical significance of profile similarity: the explicit consideration of background score distributions for each database template (subject). RESULTS: Using a simple scheme to combine and analytically approximate query- and subject-based distributions, we show that (i) inclusion of background distributions for the subjects increases the quality of homology detection; (ii) this increase is higher when the distributions are based on the scores to all known non-homologs of the subject rather than a small calibration subset of the database representatives; and (iii) these all known non-homolog distributions of scores for the subject make the dominant contribution to the improved performance: adding the calibration distribution of the query has a negligible additional effect. CONCLUSION: The construction of distributions based on the complete sets of non-homologs for each subject is particularly relevant in the setting of structure prediction where the database consists of proteins with solved 3D structure (PDB, SCOP, CATH, etc.) and therefore structural relationships between proteins are known. These results point to a potential new direction in the development of more powerful methods for remote homology detection.
Ruslan Sadreyev, Yong Wang 0025, Nick V. Grishin
BMC Bioinform.3
2007 PROMALS: towards accurate multiple sequence alignments of distantly related proteins
abstract
MOTIVATION: Accurate multiple sequence alignments are essential in protein structure modeling, functional prediction and efficient planning of experiments. Although the alignment problem has attracted considerable attention, preparation of high-quality alignments for distantly related sequences remains a difficult task. RESULTS: We developed PROMALS, a multiple alignment method that shows promising results for protein homologs with sequence identity below 10%, aligning close to half of the amino acid residues correctly on average. This is about three times more accurate than traditional pairwise sequence alignment methods. PROMALS algorithm derives its strength from several sources: (i) sequence database searches to retrieve additional homologs; (ii) accurate secondary structure prediction; (iii) a hidden Markov model that uses a novel combined scoring of amino acids and secondary structures; (iv) probabilistic consistency-based scoring applied to progressive alignment of profiles. Compared to the best alignment methods that do not use secondary structure prediction and database searches (e.g. MUMMALS, ProbCons and MAFFT), PROMALS is up to 30% more accurate, with improvement being most prominent for highly divergent homologs. Compared to SPEM and HHalign, which also employ database searches and secondary structure prediction, PROMALS shows an accuracy improvement of several percent. AVAILABILITY: The PROMALS web server is available at: http://prodata.swmed.edu/promals/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jimin Pei, Nick V. Grishin
Bioinform.2
2007 Searching for three-dimensional secondary structural patterns in proteins with ProSMoS
abstract
MOTIVATION: Many evolutionarily distant, but functionally meaningful links between proteins come to light through comparison of spatial structures. Most programs that assess structural similarity compare two proteins to each other and find regions in common between them. Structural classification experts look for a particular structural motif instead. Programs base similarity scores on superposition or closeness of either Cartesian coordinates or inter-residue contacts. Experts pay more attention to the general orientation of the main chain and mutual spatial arrangement of secondary structural elements. There is a need for a computational tool to find proteins with the same secondary structures, topological connections and spatial architecture, regardless of subtle differences in 3D coordinates. RESULTS: We developed ProSMoS--a Protein Structure Motif Search program that emulates an expert. Starting from a spatial structure, the program uses previously delineated secondary structural elements. A meta-matrix of interactions between the elements (parallel or antiparallel) minding handedness of connections (left or right) and other features (e.g. element lengths and hydrogen bonds) is constructed prior to or during the searches. All structures are reduced to such meta-matrices that contain just enough information to define a protein fold, but this definition remains very general and deviations in 3D coordinates are tolerated. User supplies a meta-matrix for a structural motif of interest, and ProSMoS finds all proteins in the protein data bank (PDB) that match the meta-matrix. ProSMoS performance is compared to other programs and is illustrated on a beta-Grasp motif. A brief analysis of all beta-Grasp-containing proteins is presented. Program availability: ProSMoS is freely available for non-commercial use from ftp://iole.swmed.edu/pub/ProSMoS.
Shuoyong Shi, Indraneel Majumdar, S. Sri Krishna, Nick V. Grishin
Bioinform.5
2007 A comprehensive system for evaluation of remote sequence similarity detection
abstract
BACKGROUND: Accurate and sensitive performance evaluation is crucial for both effective development of better structure prediction methods based on sequence similarity, and for the comparative analysis of existing methods. Up to date, there has been no satisfactory comprehensive evaluation method that (i) is based on a large and statistically unbiased set of proteins with clearly defined relationships; and (ii) covers all performance aspects of sequence-based structure predictors, such as sensitivity and specificity, alignment accuracy and coverage, and structure template quality. RESULTS: With the aim of designing such a method, we (i) select a statistically balanced set of divergent protein domains from SCOP, and define similarity relationships for the majority of these domains by complementing the best of information available in SCOP with a rigorous SVM-based algorithm; and (ii) develop protocols for the assessment of similarity detection and alignment quality from several complementary perspectives. The evaluation of similarity detection is based on ROC-like curves and includes several complementary approaches to the definition of true/false positives. Reference-dependent approaches use the 'gold standard' of pre-defined domain relationships and structure-based alignments. Reference-independent approaches assess the quality of structural match predicted by the sequence alignment, with respect to the whole domain length (global mode) or to the aligned region only (local mode). Similarly, the evaluation of alignment quality includes several reference-dependent and -independent measures, in global and local modes. As an illustration, we use our benchmark to compare the performance of several methods for the detection of remote sequence similarities, and show that different aspects of evaluation reveal different properties of the evaluated methods, highlighting their advantages, weaknesses, and potential for further development. CONCLUSION: The presented benchmark provides a new tool for a statistically unbiased assessment of methods for remote sequence similarity detection, from various complementary perspectives. This tool should be useful both for users choosing the best method for a given purpose, and for developers designing new, more powerful methods. The benchmark set, reference alignments, and evaluation codes can be downloaded from ftp://iole.swmed.edu/pub/evaluation/.
Ruslan Sadreyev, Yong Wang 0025, Bong-Hyun Kim, Nick V. Grishin
BMC Bioinform.5
2006 The HicAB cassette, a putative novel, RNA-targeting toxin-antitoxin system in archaea and bacteria
abstract
Toxin-antitoxin systems (TAS) are abundant, diverse, horizontally mobile gene modules that encode powerful resistance mechanisms in prokaryotes. We use the comparative-genomic approach to predict a new TAS that consists of a two-gene cassette encoding uncharacterized HicA and HicB proteins. Numerous bacterial and archaeal genomes encode from one to eight HicAB modules which appear to be highly prone to horizontal gene transfer. The HicB protein (COG1598/COG4226) has a partially degraded RNAse H fold, whereas HicA (COG1724) contains a double-stranded RNA-binding domain. The stable combination of these two domains suggests a link to RNA metabolism, possibly, via an RNA interference-type mechanism. In most HicB proteins, the RNAse H-like domain is fused to a DNA-binding domain, either of the ribbon-helix-helix or of the helix-turn-helix class; in other TAS, proteins containing these DNA-binding domains function as antitoxins. Thus, the HicAB module is predicted to be a novel TAS whose mechanism involves RNA-binding and, possibly, cleavage.
Kira S. Makarova, Nick V. Grishin, Eugene V. Koonin
Bioinform.2
2006 Prediction of functional specificity determinants from protein sequences using log-likelihood ratios
abstract
MOTIVATION: A number of methods have been developed to predict functional specificity determinants in protein families based on sequence information. Most of these methods rely on pre-defined functional subgroups. Manual subgroup definition is difficult because of the limited number of experimentally characterized subfamilies with differing specificity, while automatic subgroup partitioning using computational tools is a non-trivial task and does not always yield ideal results. RESULTS: We propose a new approach SPEL (specificity positions by evolutionary likelihood) to detect positions that are likely to be functional specificity determinants. SPEL, which does not require subgroup definition, takes a multiple sequence alignment of a protein family as the only input, and assigns a P-value to every position in the alignment. Positions with low P-values are likely to be important for functional specificity. An evolutionary tree is reconstructed during the calculation, and P-value estimation is based on a random model that involves evolutionary simulations. Evolutionary log-likelihood is chosen as a measure of amino acid distribution at a position. To illustrate the performance of the method, we carried out a detailed analysis of two protein families (LacI/PurR and G protein alpha subunit), and compared our method with two existing methods (evolutionary trace and mutual information based). All three methods were also compared on a set of protein families with known ligand-bound structures. AVAILABILITY: SPEL is freely available for non-commercial use. Its pre-compiled versions for several platforms and alignments used in this work are available at ftp://iole.swmed.edu/pub/SPEL/
Jimin Pei, Lisa N. Kinch, Nick V. Grishin
Bioinform.4
2005 Structural drift: a possible path to protein fold change
abstract
SUMMARY: Along with their mutating sequences, protein structures change in time. Analyzing a formate dehydrogenase domain that is evolutionarily related to ferredoxin, but simultaneously contains all the structural elements of a beta-Grasp fold, we illustrate here a mechanism termed as structural drift, by which changes to a protein fold can occur. CONTACT: [email protected].
S. Sri Krishna, Nick V. Grishin
Bioinform.2
2005 PALSSE: A program to delineate linear secondary structural elements from protein structures
abstract
BACKGROUND: The majority of residues in protein structures are involved in the formation of alpha-helices and beta-strands. These distinctive secondary structure patterns can be used to represent a protein for visual inspection and in vector-based protein structure comparison. Success of such structural comparison methods depends crucially on the accurate identification and delineation of secondary structure elements. RESULTS: We have developed a method PALSSE (Predictive Assignment of Linear Secondary Structure Elements) that delineates secondary structure elements (SSEs) from protein Calpha coordinates and specifically addresses the requirements of vector-based protein similarity searches. Our program identifies two types of secondary structures: helix and beta-strand, typically those that can be well approximated by vectors. In contrast to traditional secondary structure algorithms, which identify a secondary structure state for every residue in a protein chain, our program attributes residues to linear SSEs. Consecutive elements may overlap, thus allowing residues located at the overlapping region to have more than one secondary structure type. CONCLUSION: PALSSE is predictive in nature and can assign about 80% of the protein chain to SSEs as compared to 53% by DSSP and 57% by P-SEA. Such a generous assignment ensures almost every residue is part of an element and is used in structural comparisons. Our results are in agreement with human judgment and DSSP. The method is robust to coordinate errors and can be used to define SSEs even in poorly refined and low-resolution structures. The program and results are available at http://prodata.swmed.edu/palsse/.
Indraneel Majumdar, S. Sri Krishna, Nick V. Grishin
BMC Bioinform.3
2004 PCOAT: positional correlation analysis using multiple methods
abstract
UNLABELLED: PCOAT (Positional COrrelation Analysis Tool) is a program to perform positional correlation analysis for protein multiple sequence alignment in order to identify structurally or functionally important interactions between positions in a protein family. We implement different statistical methods to detect highly correlated position pairs, amino acid pairs, individual positions and networks of correlated positions, and utilize multiple sequence weighting and sampling methods to eliminate background correlations caused by phylogeny and stochastic events. Our program runs relatively fast and is suitable for analyzing alignments containing large number of sequences. AVAILABILITY: ftp://iole.swmed.edu/pub/PCOAT/. SUPPLEMENTARY INFORMATION: The PCOAT ftp site contains a detailed description of the program, and the results of PCOAT analysis on C2H2 alignment and ACT domain alignment.
Nick V. Grishin
Bioinform.2
2004 Quality of alignment comparison by COMPASS improves with inclusion of diverse confident homologs
abstract
MOTIVATION: Adding more distant homologs to a multiple alignment and thus increasing its diversity may eventually deteriorate the numerical profile constructed from this alignment. Here, we addressed the question whether such a diversity limit can be reached in the alignments of confident homologs found by PSI-BLAST, and we analyzed the dependence of the quality of the profile-profile comparison made by COMPASS on the sequence diversity within these alignments. RESULTS: Protein families that have a greater number of diverse confident homologs in the current sequence databases provide an increased quality of similarity detection in profile databases, but produce on average less accurate profile-profile alignments with their remote relatives. This lower alignment accuracy cannot be improved when the most distant members of these families are excluded from their profiles. On the contrary, the presence of more diverse members results in more accurate alignments. For families with a high diversity of confident homologs, the lower quality of profile alignments with their remote relatives seems to be an attribute of these families or their alignments, rather than to be caused by the large number of diverse sequences itself. Our results suggest that at any level of profile diversity, one should include in the multiple alignment as many confident sequence homologs as possible in order to produce the most accurate results.
Ruslan Sadreyev, Nick V. Grishin
Bioinform.2
2004 SCOPmap: Automated assignment of protein structures to evolutionary superfamilies
abstract
BACKGROUND: Inference of remote homology between proteins is very challenging and remains a prerogative of an expert. Thus a significant drawback to the use of evolutionary-based protein structure classifications is the difficulty in assigning new proteins to unique positions in the classification scheme with automatic methods. To address this issue, we have developed an algorithm to map protein domains to an existing structural classification scheme and have applied it to the SCOP database. RESULTS: The general strategy employed by this algorithm is to combine the results of several existing sequence and structure comparison tools applied to a query protein of known structure in order to find the homologs already classified in SCOP database and thus determine classification assignments. The algorithm is able to map domains within newly solved structures to the appropriate SCOP superfamily level with approximately 95% accuracy. Examples of correctly mapped remote homologs are discussed. The algorithm is also capable of identifying potential evolutionary relationships not specified in the SCOP database, thus helping to make it better. The strategy of the mapping algorithm is not limited to SCOP and can be applied to any other evolutionary-based classification scheme as well. SCOPmap is available for download. CONCLUSION: The SCOPmap program is useful for assigning domains in newly solved structures to appropriate superfamilies and for identifying evolutionary links between different superfamilies.
Sara Cheek, S. Sri Krishna, Lisa N. Kinch, Nick V. Grishin
BMC Bioinform.5
2004 Estimates of statistical significance for comparison of individual positions in multiple sequence alignments
abstract
BACKGROUND: Profile-based analysis of multiple sequence alignments (MSA) allows for accurate comparison of protein families. Here, we address the problems of detecting statistically confident dissimilarities between (1) MSA position and a set of predicted residue frequencies, and (2) between two MSA positions. These problems are important for (i) evaluation and optimization of methods predicting residue occurrence at protein positions; (ii) detection of potentially misaligned regions in automatically produced alignments and their further refinement; and (iii) detection of sites that determine functional or structural specificity in two related families. RESULTS: For problems (1) and (2), we propose analytical estimates of P-value and apply them to the detection of significant positional dissimilarities in various experimental situations. (a) We compare structure-based predictions of residue propensities at a protein position to the actual residue frequencies in the MSA of homologs. (b) We evaluate our method by the ability to detect erroneous position matches produced by an automatic sequence aligner. (c) We compare MSA positions that correspond to residues aligned by automatic structure aligners. (d) We compare MSA positions that are aligned by high-quality manual superposition of structures. Detected dissimilarities reveal shortcomings of the automatic methods for residue frequency prediction and alignment construction. For the high-quality structural alignments, the dissimilarities suggest sites of potential functional or structural importance. CONCLUSION: The proposed computational method is of significant potential value for the analysis of protein families.
Ruslan Sadreyev, Nick V. Grishin
BMC Bioinform.2
2003 Probabilistic scoring measures for profile-profile comparison yield more accurate short seed alignments
abstract
MOTIVATION: The development of powerful automatic methods for the comparison of protein sequences has become increasingly important. Profile-to-profile comparisons allow for the use of broader information about protein families, resulting in more sensitive and accurate comparisons of distantly related sequences. A key part in the comparison of two profiles is the method for the calculation of scores for the position matches. A number of methods based on various theoretical considerations have been proposed. We implemented several previously reported scoring functions as well as our own functions, and compared them on the basis of their ability to produce accurate short ungapped alignments of a given length. RESULTS: Our results suggest that the family of the probabilistic methods (log-odds based methods and prof_sim) may be the more appropriate choice for the generation of initial 'seeds' as the first step to produce local profile-profile alignments. The most effective scoring systems were the closely related modifications of functions previously implemented in the COMPASS and Picasso methods.
David Mittelman, Ruslan Sadreyev, Nick V. Grishin
Bioinform.3
2003 PCMA: fast and accurate multiple sequence alignment based on profile consistency
abstract
UNLABELLED: PCMA (profile consistency multiple sequence alignment) is a progressive multiple sequence alignment program that combines two different alignment strategies. Highly similar sequences are aligned in a fast way as in ClustalW, forming pre-aligned groups. The T-Coffee strategy is applied to align the relatively divergent groups based on profile-profile comparison and consistency. The scoring function for local alignments of pre-aligned groups is based on a novel profile-profile comparison method that is a generalization of the PSI-BLAST approach to profile-sequence comparison. PCMA balances speed and accuracy in a flexible way and is suitable for aligning large numbers of sequences. AVAILABILITY: PCMA is freely available for non-commercial use. Pre-compiled versions for several platforms can be downloaded from ftp://iole.swmed.edu/pub/PCMA/.
Jimin Pei, Ruslan Sadreyev, Nick V. Grishin
Bioinform.3
2002 Euclidian space and grouping of biological objects
abstract
MOTIVATION: Biological objects tend to cluster into discrete groups. Objects within a group typically possess similar properties. It is important to have fast and efficient tools for grouping objects that result in biologically meaningful clusters. Protein sequences reflect biological diversity and offer an extraordinary variety of objects for polishing clustering strategies. Grouping of sequences should reflect their evolutionary history and their functional properties. Visualization of relationships between sequences is of no less importance. Tree-building methods are typically used for such visualization. An alternative concept to visualization is a multidimensional sequence space. In this space, proteins are defined as points and distances between the points reflect the relationships between the proteins. Such a space can also be a basis for model-based clustering strategies that typically produce results correlating better with biological properties of proteins. RESULTS: We developed an approach to classification of biological objects that combines evolutionary measures of their similarity with a model-based clustering procedure. We apply the methodology to amino acid sequences. On the first step, given a multiple sequence alignment, we estimate evolutionary distances between proteins measured in expected numbers of amino acid substitutions per site. These distances are additive and are suitable for evolutionary tree reconstruction. On the second step, we find the best fit approximation of the evolutionary distances by Euclidian distances and thus represent each protein by a point in a multidimensional space. The Euclidian space may be projected in two or three dimensions and the projections can be used to visualize relationships between proteins. On the third step, we find a non-parametric estimate of the probability density of the points and cluster the points that belong to the same local maximum of this density in a group. The number of groups is controlled by a sigma-parameter that determines the shape of the density estimate and the number of maxima in it. The grouping procedure outperforms commonly used methods such as UPGMA and single linkage clustering.
Vyacheslav N. Grishin, Nick V. Grishin
Bioinform.2
2001 AL2CO: calculation of positional conservation in a protein sequence alignment
abstract
Abstract Motivation: Amino acid sequence alignments are widely used in the analysis of protein structure, function and evolutionary relationships. Proteins within a superfamily usually share the same fold and possess related functions. These structural and functional constraints are reflected in the alignment conservation patterns. Positions of functional and/or structural importance tend to be more conserved. Conserved positions are usually clustered in distinct motifs surrounded by sequence segments of low conservation. Poorly conserved regions might also arise from the imperfections in multiple alignment algorithms and thus indicate possible alignment errors. Quantification of conservation by attributing a conservation index to each aligned position makes motif detection more convenient. Mapping these conservation indices onto a protein spatial structure helps to visualize spatial conservation features of the molecule and to predict functionally and/or structurally important sites. Analysis of conservation indices could be a useful tool in detection of potentially misaligned regions and will aid in improvement of multiple alignments. Results: We developed a program to calculate a conservation index at each position in a multiple sequence alignment using several methods. Namely, amino acid frequencies at each position are estimated and the conservation index is calculated from these frequencies. We utilize both unweighted frequencies and frequencies weighted using two different strategies. Three conceptually different approaches (entropy-based, variance-based and matrix score-based) are implemented in the algorithm to define the conservation index. Calculating conservation indices for 35522 positions in 284 alignments from SMART database we demonstrate that different methods result in highly correlated (correlation coefficient more than 0.85) conservation indices. Conservation indices show statistically significant correlation between sequentially adjacent positions \batchmode \documentclass[fleqn,10pt,legalpaper]{article} \usepackage{amssymb} \usepackage{amsfonts} \usepackage{amsmath} \pagestyle{empty} \begin{document} \(i\) \end{document}and \batchmode \documentclass[fleqn,10pt,legalpaper]{article} \usepackage{amssymb} \usepackage{amsfonts} \usepackage{amsmath} \pagestyle{empty} \begin{document} \(i\ +\ j\) \end{document}, where \batchmode \documentclass[fleqn,10pt,legalpaper]{article} \usepackage{amssymb} \usepackage{amsfonts} \usepackage{amsmath} \pagestyle{empty} \begin{document} \(j\ {<}\ 13\) \end{document}, and averaging of the indices over the window of three positions is optimal for motif detection. Positions with gaps display substantially lower conservation properties. We compare conservation properties of the SMART alignments or FSSP structural alignments to those of the ClustalW alignments. The results suggest that conservation indices should be a valuable tool of alignment quality assessment and might be used as an objective function for refinement of multiple alignments. Availability: The C code of the AL2CO program and its pre-compiled versions for several platforms as well as the details of the analysis are freely available at ftp://iole.swmed.edu/pub/al2co/. Contact: [email protected] * To whom correspondence should be addressed.
Jimin Pei, Nick V. Grishin
Bioinform.2