VLDB 2026 Research / reviewers in the wild / expert
Jimin Pei
dblp:06/566
· DBLP profile ↗
15ranked-venue papers
7as first author
3since 2021 · last 2025
0000-0002-3505-9665ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 7 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
11 papers |
Bioinformatics and computational biology · 100% |
Topics — the 21 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
protein structure prediction |
0.7 | 2 | 2022 | Human mitochondrial protein complexes revealed by large-scale coevolution analysis and deep learning-based structure modeling · Bioinform. 2022 An automatic method for CASP9 free modeling structure prediction assessment · Bioinform. 2011 |
Bioinformatics and computational biology › protein function prediction
nuclear export signal prediction |
0.7 | 2 | 2020 | pCRM1exportome: database of predicted CRM1-dependent Nuclear Export Signal (NES) motifs in cancer-related genes · Bioinform. 2020 LocNES: a computational tool for locating classical NESs in CRM1 cargo proteins · Bioinform. 2015 |
Bioinformatics and computational biology
protein sequence analysis |
0.6 | 4 | 2020 | pCRM1exportome: database of predicted CRM1-dependent Nuclear Export Signal (NES) motifs in cancer-related genes · Bioinform. 2020 PROMALS: towards accurate multiple sequence alignments of distantly related proteins · Bioinform. 2007 Prediction of functional specificity determinants from protein sequences using log-likelihood ratios · Bioinform. 2006 |
Bioinformatics and computational biology › protein-protein interaction prediction
coevolution-based interaction prediction |
0.6 | 1 | 2022 | Human mitochondrial protein complexes revealed by large-scale coevolution analysis and deep learning-based structure modeling · Bioinform. 2022 |
Bioinformatics and computational biology › protein structure analysis › intrinsically disordered protein
intrinsically disordered protein database |
0.6 | 1 | 2022 | DisEnrich: database of enriched regions in human dark proteome · Bioinform. 2022 |
Bioinformatics and computational biology › protein structure analysis › quaternary structure
protein complex modeling |
0.6 | 1 | 2022 | Human mitochondrial protein complexes revealed by large-scale coevolution analysis and deep learning-based structure modeling · Bioinform. 2022 |
Bioinformatics and computational biology
protein-protein interaction prediction |
0.6 | 1 | 2022 | Human mitochondrial protein complexes revealed by large-scale coevolution analysis and deep learning-based structure modeling · Bioinform. 2022 |
Bioinformatics and computational biology
multiple sequence alignment |
0.5 | 4 | 2018 | A sequence family database built on ECOD structural domains · Bioinform. 2018 PROMALS: towards accurate multiple sequence alignments of distantly related proteins · Bioinform. 2007 PCMA: fast and accurate multiple sequence alignment based on profile consistency · Bioinform. 2003 |
Bioinformatics and computational biology
protein function prediction |
0.4 | 2 | 2022 | LocNES: a computational tool for locating classical NESs in CRM1 cargo proteins · Bioinform. 2015 DisEnrich: database of enriched regions in human dark proteome · Bioinform. 2022 |
Bioinformatics and computational biology
sequence analysis |
0.4 | 2 | 2018 | A sequence family database built on ECOD structural domains · Bioinform. 2018 PCMA: fast and accurate multiple sequence alignment based on profile consistency · Bioinform. 2003 |
Bioinformatics and computational biology › sequence analysis
profile hidden markov model |
0.3 | 1 | 2018 | A sequence family database built on ECOD structural domains · Bioinform. 2018 |
Bioinformatics and computational biology › protein structure analysis › protein domain identification
protein domain classification |
0.3 | 1 | 2018 | A sequence family database built on ECOD structural domains · Bioinform. 2018 |
Bioinformatics and computational biology › structural biology
protein structure and function |
0.3 | 1 | 2018 | A sequence family database built on ECOD structural domains · Bioinform. 2018 |
Bioinformatics and computational biology
protein structure analysis |
0.2 | 1 | 2013 | Defining and predicting structurally conserved regions in protein superfamilies · Bioinform. 2013 |
Bioinformatics and computational biology
cancer genomics |
0.1 | 1 | 2020 | pCRM1exportome: database of predicted CRM1-dependent Nuclear Export Signal (NES) motifs in cancer-related genes · Bioinform. 2020 |
Bioinformatics and computational biology › protein structure prediction
model quality assessment |
0.1 | 1 | 2011 | An automatic method for CASP9 free modeling structure prediction assessment · Bioinform. 2011 |
Bioinformatics and computational biology
phylogenetics |
0.1 | 1 | 2006 | Prediction of functional specificity determinants from protein sequences using log-likelihood ratios · Bioinform. 2006 |
Bioinformatics and computational biology › protein structure prediction › template-based modeling
homology modeling |
0.0 | 1 | 2013 | Defining and predicting structurally conserved regions in protein superfamilies · Bioinform. 2013 |
Bioinformatics and computational biology › protein sequence analysis
protein superfamily analysis |
0.0 | 1 | 2013 | Defining and predicting structurally conserved regions in protein superfamilies · Bioinform. 2013 |
Bioinformatics and computational biology › multiple sequence alignment
progressive alignment |
0.0 | 1 | 2003 | PCMA: fast and accurate multiple sequence alignment based on profile consistency · Bioinform. 2003 |
Bioinformatics and computational biology › sequence analysis › sequence comparison
sequence conservation analysis |
0.0 | 1 | 2001 | AL2CO: calculation of positional conservation in a protein sequence alignment · Bioinform. 2001 |
Methods — techniques the papers use, named apart from their topics
gene ontology annotation · 0.6disorder prediction · 0.6deep learning-based structure prediction · 0.6co-evolution analysis · 0.6structure-based scoring · 0.4sequence-based scoring · 0.4structure superposition · 0.3HMMER · 0.3position-specific scoring matrix · 0.2disorder propensity · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Using evolutionary context to classify difficult protein foldsabstractRecent advances in protein structure prediction, such as AlphaFold2, have enabled identification of vast numbers of putative novel protein domains across the sequence space, many of which adopt structures dissimilar to known folds. Based on structural segmentation and classification, the Encyclopedia of Domains (TED) project recently cataloged more than 7400 low-symmetry, structure-based domains as candidate novel-fold (CNF) domains. To place these domains in their broader evolutionary and structural context, we applied DPAM (Domain Parser for AlphaFold Models), a complementary method that combines AlphaFold-derived confidence metrics with sensitive sequence and structure similarity searches, to parse domains for the AlphaFold models of proteins containing TED CNF domains. We identified 8044 DPAM domains with significant overlap with TED CNF domains, among which 2490 were confidently assigned to entries in the ECOD (Evolutionary Classification of protein Domains) structural classification hierarchy. Our results suggest that a substantial subset of TED candidate novel-fold domains are distant homologs of existing ECOD domains. Comparison of domain boundaries between TED and DPAM showed varied patterns: more than one-third of cases featured TED CNF domains largely embedded within DPAM domains-often representing insertions or extensions into enzymatic or repeat folds. A smaller fraction (17%) exhibited consistent domain boundaries between TED and DPAM. These consistently defined domains are often characterized by significant structural diversity, including long insertions and duplications. An even smaller subset showed the reverse relationship, with DPAM domains largely embedded within TED CNF domains. In these cases, DPAM effectively separated multiple structural units that TED grouped as single domains. Together, these findings highlight the complementarity of structural and evolutionary approaches for domain annotation and demonstrate the power of integrative methods, such as DPAM, in refining the classification of challenging protein folds and uncovering distant evolutionary relationships. Jimin Pei, R. Dustin Schaeffer, Qian Cong, Nick V. Grishin |
PLoS Comput. Biol. | 1 |
| 2022 | DisEnrich: database of enriched regions in human dark proteomeabstractMOTIVATION: Intrinsically disordered proteins (IDPs) are involved in numerous processes crucial for living organisms. Bias in amino acid composition of these proteins determines their unique biophysical and functional features. Distinct intrinsically disordered regions (IDRs) with compositional bias play different important roles in various biological processes. IDRs enriched in particular amino acids in human proteome have not been described consistently. RESULTS: We developed DisEnrich-the database of human proteome IDRs that are significantly enriched in particular amino acids. Each human protein is described using Gene Ontology (GO) function terms, disorder prediction for the full-length sequence using three methods, enriched IDR composition and ranks of human proteins with similar enriched IDRs. Distribution analysis of enriched IDRs among broad functional categories revealed significant overrepresentation of R- and Y-enriched IDRs in metabolic and enzymatic activities and F-enriched IDRs in transport. About 75% of functional categories contain IDPs with IDRs significantly enriched in hydrophobic residues that are important for protein-protein interactions. AVAILABILITY AND IMPLEMENTATION: The database is available at http://prodata.swmed.edu/DisEnrichDB/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics Advances online. Kirill E. Medvedev, Jimin Pei, Nick V. Grishin |
Bioinform. | 2 |
| 2022 | Human mitochondrial protein complexes revealed by large-scale coevolution analysis and deep learning-based structure modelingabstractMOTIVATION: Recent development of deep-learning methods has led to a breakthrough in the prediction accuracy of 3D protein structures. Extending these methods to protein pairs is expected to allow large-scale detection of protein-protein interactions (PPIs) and modeling protein complexes at the proteome level. RESULTS: We applied RoseTTAFold and AlphaFold, two of the latest deep-learning methods for structure predictions, to analyze coevolution of human proteins residing in mitochondria, an organelle of vital importance in many cellular processes including energy production, metabolism, cell death and antiviral response. Variations in mitochondrial proteins have been linked to a plethora of human diseases and genetic conditions. RoseTTAFold, with high computational speed, was used to predict the coevolution of about 95% of mitochondrial protein pairs. Top-ranked pairs were further subject to modeling of the complex structures by AlphaFold, which also produced contact probability with high precision and in many cases consistent with RoseTTAFold. Most top-ranked pairs with high contact probability were supported by known PPIs and/or similarities to experimental structural complexes. For high-scoring pairs without experimental complex structures, our coevolution analyses and structural models shed light on the details of their interfaces, including CHCHD4-AIFM1, MTERF3-TRUB2, FMC1-ATPAF2 and ECSIT-NDUFAF1. We also identified novel PPIs (PYURF-NDUFAF5, LYRM1-MTRF1L and COA8-COX10) for several proteins without experimentally characterized interaction partners, leading to predictions of their molecular functions and the biological processes they are involved in. AVAILABILITY AND IMPLEMENTATION: Data of mitochondrial proteins and their interactions are available at: http://conglab.swmed.edu/mitochondria. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jimin Pei, Jing Zhang 0115, Qian Cong |
Bioinform. | 1 |
| 2020 | pCRM1exportome: database of predicted CRM1-dependent Nuclear Export Signal (NES) motifs in cancer-related genesabstractMOTIVATION: The consensus pattern of Nuclear Export Signal (NES) is a short sequence motif that is commonly identified in protein sequences, whether the motif acts as an NES (true positive) or not (false positive). Finding more plausible NES functioning regions among the vast array of consensus-matching segments would provide an interesting resource for further experimental validation. Better defined NES should also allow meaningful mapping of cancer-related mutation positions, leading to plausible explanations for the relationship between nuclear export and disease. RESULTS: Possible NES candidate regions are extracted from the cancer-related human reference proteome. Extracted NES are scored for reliability by combining sequence-based and structure-based approaches. The confidently identified NES candidate motifs were checked for overlap with cancer-related mutation positions annotated in the COSMIC database. Among the ∼700 cancer-related sequences in the COSMIC Cancer Gene Census, 178 sequences are predicted to have possible NES motifs containing cancer-related mutations at their key positions. These lists are organized into our database (pCRM1exportome), and other protein sequences in the human reference proteome can also be retrieved by their UniProt IDs. AVAILABILITY AND IMPLEMENTATION: The database is freely available at http://prodata.swmed.edu/pCRM1exportome. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jordan M. Baumhardt, Jimin Pei, Yuh Min Chook, Nick V. Grishin |
Bioinform. | 3 |
| 2020 | Mutation severity spectrum of rare alleles in the human genome is predictive of disease typeabstractThe human genome harbors a variety of genetic variations. Single-nucleotide changes that alter amino acids in protein-coding regions are one of the major causes of human phenotypic variation and diseases. These single-amino acid variations (SAVs) are routinely found in whole genome and exome sequencing. Evaluating the functional impact of such genomic alterations is crucial for diagnosis of genetic disorders. We developed DeepSAV, a deep-learning convolutional neural network to differentiate disease-causing and benign SAVs based on a variety of protein sequence, structural and functional properties. Our method outperforms most stand-alone programs, and the version incorporating population and gene-level information (DeepSAV+PG) has similar predictive power as some of the best available. We transformed DeepSAV scores of rare SAVs in the human population into a quantity termed "mutation severity measure" for each human protein-coding gene. It reflects a gene's tolerance to deleterious missense mutations and serves as a useful tool to study gene-disease associations. Genes implicated in cancer, autism, and viral interaction are found by this measure as intolerant to mutations, while genes associated with a number of other diseases are scored as tolerant. Among known disease-associated genes, those that are mutation-intolerant are likely to function in development and signal transduction pathways, while those that are mutation-tolerant tend to encode metabolic and mitochondrial proteins. Jimin Pei, Lisa N. Kinch, Zbyszek Otwinowski, Nick V. Grishin |
PLoS Comput. Biol. | 1 |
| 2018 | A sequence family database built on ECOD structural domainsabstractMotivation: The ECOD database classifies protein domains based on their evolutionary relationships, considering both remote and close homology. The family group in ECOD provides classification of domains that are closely related to each other based on sequence similarity. Due to different perspectives on domain definition, direct application of existing sequence domain databases, such as Pfam, to ECOD struggles with several shortcomings. Results: We created multiple sequence alignments and profiles from ECOD domains with the help of structural information in alignment building and boundary delineation. We validated the alignment quality by scoring structure superposition to demonstrate that they are comparable to curated seed alignments in Pfam. Comparison to Pfam and CDD reveals that 27 and 16% of ECOD families are new, but they are also dominated by small families, likely because of the sampling bias from the PDB database. There are 35 and 48% of families whose boundaries are modified comparing to counterparts in Pfam and CDD, respectively. Availability and implementation: The new families are now integrated in the ECOD website. The aggregate HMMER profile library and alignment are available for download on ECOD website (http://prodata.swmed.edu/ecod). Supplementary information: Supplementary data are available at Bioinformatics online. Yuxing Liao, R. Dustin Schaeffer, Jimin Pei, Nick V. Grishin |
Bioinform. | 3 |
| 2015 | LocNES: a computational tool for locating classical NESs in CRM1 cargo proteinsabstractMOTIVATION: Classical nuclear export signals (NESs) are short cognate peptides that direct proteins out of the nucleus via the CRM1-mediated export pathway. CRM1 regulates the localization of hundreds of macromolecules involved in various cellular functions and diseases. Due to the diverse and complex nature of NESs, reliable prediction of the signal remains a challenge despite several attempts made in the last decade. RESULTS: We present a new NES predictor, LocNES. LocNES scans query proteins for NES consensus-fitting peptides and assigns these peptides probability scores using Support Vector Machine model, whose feature set includes amino acid sequence, disorder propensity, and the rank of position-specific scoring matrix score. LocNES demonstrates both higher sensitivity and precision over existing NES prediction tools upon comparative analysis using experimentally identified NESs. AVAILABILITY AND IMPLEMENTATION: LocNES is freely available at http://prodata.swmed.edu/LocNES CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Darui Xu, Kara Marquis, Jimin Pei, Szu-Chin Fu, Tolga Cagatay, Nick V. Grishin, Yuh Min Chook |
Bioinform. | 3 |
| 2015 | SFESA: a web server for pairwise alignment refinement by secondary structure shiftsabstractBACKGROUND: Protein sequence alignment is essential for a variety of tasks such as homology modeling and active site prediction. Alignment errors remain the main cause of low-quality structure models. A bioinformatics tool to refine alignments is needed to make protein alignments more accurate. RESULTS: We developed the SFESA web server to refine pairwise protein sequence alignments. Compared to the previous version of SFESA, which required a set of 3D coordinates for a protein, the new server will search a sequence database for the closest homolog with an available 3D structure to be used as a template. For each alignment block defined by secondary structure elements in the template, SFESA evaluates alignment variants generated by local shifts and selects the best-scoring alignment variant. A scoring function that combines the sequence score of profile-profile comparison and the structure score of template-derived contact energy is used for evaluation of alignments. PROMALS pairwise alignments refined by SFESA are more accurate than those produced by current advanced alignment methods such as HHpred and CNFpred. In addition, SFESA also improves alignments generated by other software. CONCLUSIONS: SFESA is a web-based tool for alignment refinement, designed for researchers to compute, refine, and evaluate pairwise alignments with a combined sequence and structure scoring of alignment blocks. To our knowledge, the SFESA web server is the only tool that refines alignments by evaluating local shifts of secondary structure elements. The SFESA web server is available at http://prodata.swmed.edu/sfesa. Jing Tong, Jimin Pei, Nick V. Grishin |
BMC Bioinform. | 2 |
| 2014 | ECOD: An Evolutionary Classification of Protein DomainsabstractUnderstanding the evolution of a protein, including both close and distant relationships, often reveals insight into its structure and function. Fast and easy access to such up-to-date information facilitates research. We have developed a hierarchical evolutionary classification of all proteins with experimentally determined spatial structures, and presented it as an interactive and updatable online database. ECOD (Evolutionary Classification of protein Domains) is distinct from other structural classifications in that it groups domains primarily by evolutionary relationships (homology), rather than topology (or "fold"). This distinction highlights cases of homology between domains of differing topology to aid in understanding of protein structure evolution. ECOD uniquely emphasizes distantly related homologs that are difficult to detect, and thus catalogs the largest number of evolutionary links among structural domain classifications. Placing distant homologs together underscores the ancestral similarities of these proteins and draws attention to the most important regions of sequence and structure, as well as conserved functional sites. ECOD also recognizes closer sequence-based relationships between protein domains. Currently, approximately 100,000 protein structures are classified in ECOD into 9,000 sequence families clustered into close to 2,000 evolutionary groups. The classification is assisted by an automated pipeline that quickly and consistently classifies weekly releases of PDB structures and allows for continual updates. This synchronization with PDB uniquely distinguishes ECOD among all protein classifications. Finally, we present several case studies of homologous proteins not recorded in other classifications, illustrating the potential of how ECOD can be used to further biological and evolutionary studies. R. Dustin Schaeffer, Yuxing Liao, Lisa N. Kinch, Jimin Pei, Shuoyong Shi, Bong-Hyun Kim, Nick V. Grishin |
PLoS Comput. Biol. | 5 |
| 2013 | Defining and predicting structurally conserved regions in protein superfamiliesabstractMOTIVATION: The structures of homologous proteins are generally better conserved than their sequences. This phenomenon is demonstrated by the prevalence of structurally conserved regions (SCRs) even in highly divergent protein families. Defining SCRs requires the comparison of two or more homologous structures and is affected by their availability and divergence, and our ability to deduce structurally equivalent positions among them. In the absence of multiple homologous structures, it is necessary to predict SCRs of a protein using information from only a set of homologous sequences and (if available) a single structure. Accurate SCR predictions can benefit homology modelling and sequence alignment. RESULTS: Using pairwise DaliLite alignments among a set of homologous structures, we devised a simple measure of structural conservation, termed structural conservation index (SCI). SCI was used to distinguish SCRs from non-SCRs. A database of SCRs was compiled from 386 SCOP superfamilies containing 6489 protein domains. Artificial neural networks were then trained to predict SCRs with various features deduced from a single structure and homologous sequences. Assessment of the predictions via a 5-fold cross-validation method revealed that predictions based on features derived from a single structure perform similarly to ones based on homologous sequences, while combining sequence and structural features was optimal in terms of accuracy (0.755) and Matthews correlation coefficient (0.476). These results suggest that even without information from multiple structures, it is still possible to effectively predict SCRs for a protein. Finally, inspection of the structures with the worst predictions pinpoints difficulties in SCR definitions. AVAILABILITY: The SCR database and the prediction server can be found at http://prodata.swmed.edu/SCR. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics Online. Ivan K. Huang, Jimin Pei, Nick V. Grishin |
Bioinform. | 2 |
| 2011 | An automatic method for CASP9 free modeling structure prediction assessmentabstractMOTIVATION: Manual inspection has been applied to and is well accepted for assessing critical assessment of protein structure prediction (CASP) free modeling (FM) category predictions over the years. Such manual assessment requires expertise and significant time investment, yet has the problems of being subjective and unable to differentiate models of similar quality. It is beneficial to incorporate the ideas behind manual inspection to an automatic score system, which could provide objective and reproducible assessment of structure models. RESULTS: Inspired by our experience in CASP9 FM category assessment, we developed an automatic superimposition independent method named Quality Control Score (QCS) for structure prediction assessment. QCS captures both global and local structural features, with emphasis on global topology. We applied this method to all FM targets from CASP9, and overall the results showed the best agreement with Manual Inspection Scores among automatic prediction assessment methods previously applied in CASPs, such as Global Distance Test Total Score (GDT_TS) and Contact Score (CS). As one of the important components to guide our assessment of CASP9 FM category predictions, this method correlates well with other scoring methods and yet is able to reveal good-quality models that are missed by GDT_TS. AVAILABILITY: The script for QCS calculation is available at http://prodata.swmed.edu/QCS/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qian Cong, Lisa N. Kinch, Jimin Pei, Shuoyong Shi, Vyacheslav N. Grishin, Nick V. Grishin |
Bioinform. | 3 |
| 2007 | PROMALS: towards accurate multiple sequence alignments of distantly related proteinsabstractMOTIVATION: Accurate multiple sequence alignments are essential in protein structure modeling, functional prediction and efficient planning of experiments. Although the alignment problem has attracted considerable attention, preparation of high-quality alignments for distantly related sequences remains a difficult task. RESULTS: We developed PROMALS, a multiple alignment method that shows promising results for protein homologs with sequence identity below 10%, aligning close to half of the amino acid residues correctly on average. This is about three times more accurate than traditional pairwise sequence alignment methods. PROMALS algorithm derives its strength from several sources: (i) sequence database searches to retrieve additional homologs; (ii) accurate secondary structure prediction; (iii) a hidden Markov model that uses a novel combined scoring of amino acids and secondary structures; (iv) probabilistic consistency-based scoring applied to progressive alignment of profiles. Compared to the best alignment methods that do not use secondary structure prediction and database searches (e.g. MUMMALS, ProbCons and MAFFT), PROMALS is up to 30% more accurate, with improvement being most prominent for highly divergent homologs. Compared to SPEM and HHalign, which also employ database searches and secondary structure prediction, PROMALS shows an accuracy improvement of several percent. AVAILABILITY: The PROMALS web server is available at: http://prodata.swmed.edu/promals/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jimin Pei, Nick V. Grishin |
Bioinform. | 1 |
| 2006 | Prediction of functional specificity determinants from protein sequences using log-likelihood ratiosabstractMOTIVATION: A number of methods have been developed to predict functional specificity determinants in protein families based on sequence information. Most of these methods rely on pre-defined functional subgroups. Manual subgroup definition is difficult because of the limited number of experimentally characterized subfamilies with differing specificity, while automatic subgroup partitioning using computational tools is a non-trivial task and does not always yield ideal results. RESULTS: We propose a new approach SPEL (specificity positions by evolutionary likelihood) to detect positions that are likely to be functional specificity determinants. SPEL, which does not require subgroup definition, takes a multiple sequence alignment of a protein family as the only input, and assigns a P-value to every position in the alignment. Positions with low P-values are likely to be important for functional specificity. An evolutionary tree is reconstructed during the calculation, and P-value estimation is based on a random model that involves evolutionary simulations. Evolutionary log-likelihood is chosen as a measure of amino acid distribution at a position. To illustrate the performance of the method, we carried out a detailed analysis of two protein families (LacI/PurR and G protein alpha subunit), and compared our method with two existing methods (evolutionary trace and mutual information based). All three methods were also compared on a set of protein families with known ligand-bound structures. AVAILABILITY: SPEL is freely available for non-commercial use. Its pre-compiled versions for several platforms and alignments used in this work are available at ftp://iole.swmed.edu/pub/SPEL/ Jimin Pei, Lisa N. Kinch, Nick V. Grishin |
Bioinform. | 1 |
| 2003 | PCMA: fast and accurate multiple sequence alignment based on profile consistencyabstractUNLABELLED: PCMA (profile consistency multiple sequence alignment) is a progressive multiple sequence alignment program that combines two different alignment strategies. Highly similar sequences are aligned in a fast way as in ClustalW, forming pre-aligned groups. The T-Coffee strategy is applied to align the relatively divergent groups based on profile-profile comparison and consistency. The scoring function for local alignments of pre-aligned groups is based on a novel profile-profile comparison method that is a generalization of the PSI-BLAST approach to profile-sequence comparison. PCMA balances speed and accuracy in a flexible way and is suitable for aligning large numbers of sequences. AVAILABILITY: PCMA is freely available for non-commercial use. Pre-compiled versions for several platforms can be downloaded from ftp://iole.swmed.edu/pub/PCMA/. Jimin Pei, Ruslan Sadreyev, Nick V. Grishin |
Bioinform. | 1 |
| 2001 | AL2CO: calculation of positional conservation in a protein sequence alignmentabstractAbstract Motivation: Amino acid sequence alignments are widely used in the analysis of protein structure, function and evolutionary relationships. Proteins within a superfamily usually share the same fold and possess related functions. These structural and functional constraints are reflected in the alignment conservation patterns. Positions of functional and/or structural importance tend to be more conserved. Conserved positions are usually clustered in distinct motifs surrounded by sequence segments of low conservation. Poorly conserved regions might also arise from the imperfections in multiple alignment algorithms and thus indicate possible alignment errors. Quantification of conservation by attributing a conservation index to each aligned position makes motif detection more convenient. Mapping these conservation indices onto a protein spatial structure helps to visualize spatial conservation features of the molecule and to predict functionally and/or structurally important sites. Analysis of conservation indices could be a useful tool in detection of potentially misaligned regions and will aid in improvement of multiple alignments. Results: We developed a program to calculate a conservation index at each position in a multiple sequence alignment using several methods. Namely, amino acid frequencies at each position are estimated and the conservation index is calculated from these frequencies. We utilize both unweighted frequencies and frequencies weighted using two different strategies. Three conceptually different approaches (entropy-based, variance-based and matrix score-based) are implemented in the algorithm to define the conservation index. Calculating conservation indices for 35522 positions in 284 alignments from SMART database we demonstrate that different methods result in highly correlated (correlation coefficient more than 0.85) conservation indices. Conservation indices show statistically significant correlation between sequentially adjacent positions \batchmode \documentclass[fleqn,10pt,legalpaper]{article} \usepackage{amssymb} \usepackage{amsfonts} \usepackage{amsmath} \pagestyle{empty} \begin{document} \(i\) \end{document}and \batchmode \documentclass[fleqn,10pt,legalpaper]{article} \usepackage{amssymb} \usepackage{amsfonts} \usepackage{amsmath} \pagestyle{empty} \begin{document} \(i\ +\ j\) \end{document}, where \batchmode \documentclass[fleqn,10pt,legalpaper]{article} \usepackage{amssymb} \usepackage{amsfonts} \usepackage{amsmath} \pagestyle{empty} \begin{document} \(j\ {<}\ 13\) \end{document}, and averaging of the indices over the window of three positions is optimal for motif detection. Positions with gaps display substantially lower conservation properties. We compare conservation properties of the SMART alignments or FSSP structural alignments to those of the ClustalW alignments. The results suggest that conservation indices should be a valuable tool of alignment quality assessment and might be used as an objective function for refinement of multiple alignments. Availability: The C code of the AL2CO program and its pre-compiled versions for several platforms as well as the details of the analysis are freely available at ftp://iole.swmed.edu/pub/al2co/. Contact: [email protected] * To whom correspondence should be addressed. Jimin Pei, Nick V. Grishin |
Bioinform. | 1 |