EDBT 2026 Demo / reviewers in the wild / expert
Jeffrey Skolnick
dblp:85/4219
· DBLP profile ↗
30ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0002-1877-4958ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 28 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scalable discovery and validation of order-specific electronic health record event trajectories for interpretable adverse-outcome risk estimationabstractOBJECTIVE: Clinicians must estimate patients' risk of adverse outcomes, yet many electronic health record tools represent health history as an unordered "bag of codes," discarding temporal order. Machine learning tools that leverage sequence order may offer better predictions but often lack interpretability for clinical audit. We introduce an interpretable, order-aware framework that mines frequent event pairs (A → B) and tests whether their ordering provides prognostic information for estimating risk of adverse outcomes (C) beyond the same codes without order. METHODS: Using the NIH All of Us Controlled Tier, 432,617 eligible participants were split 70/30 into discovery and confirmation. We combined 340,687 A → B pairs with nine outcomes (C) to create 3,066,183 candidate trajectories and tested 85,535 trajectories in this pilot using observation-aware indexing, a 90-day latency and 5-year A → B gap, inverse-probability weighting, weighted competing-risk cumulative incidence, four prespecified comparators (primary: B with no prior A), and Benjamini-Hochberg false discovery rate (FDR) in a discovery/confirmation design. RESULTS: After discovery FDR, 663 unique trajectories advanced and 171 validated for ≥ 1 horizon in confirmation. At 5 years, the median risk ratio (RR) was 2.02 versus the primary comparator and 3.89 versus a calendar-time baseline (median absolute risk difference 2.49 percentage-points). Reverse-order checks were feasible for ∼ 87% of hypotheses (median A → B → C vs B → A → C RR 1.18) at 5 years. CONCLUSION: Ordered trajectories can provide interpretable, audit-ready signal beyond unordered features and can augment white-box risk estimation for adverse clinical outcomes. Brice Edelman, Hannah Kim 0013, Jeffrey Skolnick |
J. Biomed. Informatics | 3 |
| 2026 | MENDELSEEK: An algorithm that predicts mendelian genes and elucidates what makes them specialabstractAlthough individual Mendelian diseases-those caused by a single gene-are rare, their collective disease burden is substantial. Identifying the causal gene for each condition is essential for accurate diagnosis and effective treatment. Yet, despite decades of research, the genetic basis of more than half of all known Mendelian diseases remains unresolved. To address this gap, we introduce MENDELSEEK, a machine learning framework that predicts Mendelian genes by integrating residue variation scores with pathway participation, Gene Ontology processes, and protein language model features. In benchmarking across 16,946 human genes with 10-fold cross-validation, MENDELSEEK achieved an AUC of 0.869 and an AUPR of 0.737-substantially outperforming the next best methods, ENTPRISE+ENTPRISE-X (AUC 0.781; AUPR 0.626), and REVEL (AUC 0.585; AUPR 0.401). When applied to the full set of 17,858 human genes, MENDELSEEK predicted 1,277 novel Mendelian gene candidates with precision greater than 0.7. Analysis further revealed that Mendelian genes engage in significantly more protein-protein interactions than non-Mendelian genes and are evolutionarily ancient. Together, these results highlight MENDELSEEK as a major advance over existing methods, offering new insights into the biochemical features that distinguish Mendelian from non-Mendelian genes. Hongyi Zhou, Brice Edelman, Jeffrey Skolnick |
PLoS Comput. Biol. | 3 |
| 2025 | AF3Complex yields improved structural predictions of protein complexesabstractMOTIVATION: Accurate structures of protein complexes are essential for understanding biological pathway function. A previous study showed how downstream modifications to AlphaFold 2 could yield AF2Complex, a model better suited for protein complexes. Here, we introduce AF3Complex, a model equipped with both similar and novel improvements, built on AlphaFold 3. RESULTS: Benchmarking AF3Complex and AlphaFold 3 on a large dataset of protein complexes, it was shown that AF3Complex outperforms AlphaFold 3. Moreover, by evaluating the structures generated by AF3Complex on datasets of protein-peptide complexes and antibody-antigen complexes, it was established that AF3Complex could create high-fidelity structures for these challenging complex types. Additionally, when deployed to generate structural predictions for protein complexes used in the recent CASP16 competition, AF3Complex yielded structures that would have placed it among the top models in the competition. AVAILABILITY AND IMPLEMENTATION: The AF3Complex code is freely available at https://github.com/Jfeldman34/AF3Complex.git. Jonathan Feldman 0001, Jeffrey Skolnick |
Bioinform. | 2 |
| 2025 | Valsci: an open-source, self-hostable literature review utility for automated large-batch scientific claim verification using large language modelsabstractBACKGROUND: The exponential growth of scientific publications poses a formidable challenge for researchers seeking to validate emerging hypotheses or synthesize existing evidence. In this paper, we introduce Valsci, an open-source, self-hostable utility that automates large-batch scientific claim verification using any OpenAI-compatible large language model. Valsci unites retrieval-augmented generation with structured bibliometric scoring and chain-of-thought prompting, enabling users to efficiently search, evaluate, and summarize evidence from the Semantic Scholar database and other academic sources. Unlike conventional standalone LLMs, which often suffer from hallucinations and unreliable citations, Valsci grounds its analyses in verifiable published findings. A guided prompt-flow approach is employed to generate query expansions, retrieve relevant excerpts, and synthesize coherent, evidence-based reports. RESULTS: Preliminary evaluations across claims from the SciFact benchmark dataset reveal that Valsci significantly outperforms base GPT-4o outputs in citation hallucination rate while maintaining a low misclassification rate. The system is highly scalable, processing hundreds of claims per hour through asynchronous parallelization. CONCLUSIONS: By providing an open and transparent platform for large-batch literature verification, Valsci substantially lowers the barrier to comprehensive evidence-based reviews and fosters a more reproducible research ecosystem. Brice Edelman, Jeffrey Skolnick |
BMC Bioinform. | 2 |
| 2024 | HiGen: Hierarchy-Aware Sequence Generation for Hierarchical Text ClassificationabstractHierarchical text classification (HTC) is a complex subtask under multi-label text classification, characterized by a hierarchical label taxonomy and data imbalance. The best-performing models aim to learn a static representation by combining document and hierarchical label information. However, the relevance of document sections can vary based on the hierarchy level, necessitating a dynamic document representation. To address this, we propose HiGen, a text-generation-based framework utilizing language models to encode dynamic text representations. We introduce a level-guided loss function to capture the relationship between text and label name semantics. Our approach incorporates a task-specific pretraining strategy, adapting the language model to in-domain knowledge and significantly enhancing performance for classes with limited examples. Furthermore, we present a new and valuable dataset called ENZYME, designed for HTC, which comprises articles from PubMed with the goal of predicting Enzyme Commission (EC) numbers. Through extensive experiments on the ENZYME dataset and the widely recognized WOS and NYT datasets, our methodology demonstrates superior performance, surpassing existing approaches while efficiently handling data and mitigating class imbalance. We release our code and dataset here: https://github.com/viditjain99/HiGen. Vidit Jain, Mukund Rungta, Yuchen Zhuang, Yue Yu 0001, Mu Gao, Jeffrey Skolnick, Chao Zhang 0014 |
EACL (1) | 7 |
| 2023 | Predicted structural proteome of Sphagnum divinum and proteome-scale annotationabstractMOTIVATION: Sphagnum-dominated peatlands store a substantial amount of terrestrial carbon. The genus is undersampled and under-studied. No experimental crystal structure from any Sphagnum species exists in the Protein Data Bank and fewer than 200 Sphagnum-related genes have structural models available in the AlphaFold Protein Structure Database. Tools and resources are needed to help bridge these gaps, and to enable the analysis of other structural proteomes now made possible by accurate structure prediction. RESULTS: We present the predicted structural proteome (25 134 primary transcripts) of Sphagnum divinum computed using AlphaFold, structural alignment results of all high-confidence models against an annotated nonredundant crystallographic database of over 90,000 structures, a structure-based classification of putative Enzyme Commission (EC) numbers across this proteome, and the computational method to perform this proteome-scale structure-based annotation. AVAILABILITY AND IMPLEMENTATION: All data and code are available in public repositories, detailed at https://github.com/BSDExabio/SAFA. The structural models of the S. divinum proteome have been deposited in the ModelArchive repository at https://modelarchive.org/doi/10.5452/ma-ornl-sphdiv. Russell B. Davidson, Mark Coletti, Mu Gao, Bryan Piatkowski, Avinash Sreedasyam, Farhan Quadir, David J. Weston, Jeremy Schmutz, Jianlin Cheng, Jeffrey Skolnick, Jerry M. Parks, Ada Sedova |
Bioinform. | 10 |
| 2021 | A novel sequence alignment algorithm based on deep learning of the protein folding codeabstractMOTIVATION: From evolutionary interference, function annotation to structural prediction, protein sequence comparison has provided crucial biological insights. While many sequence alignment algorithms have been developed, existing approaches often cannot detect hidden structural relationships in the 'twilight zone' of low sequence identity. To address this critical problem, we introduce a computational algorithm that performs protein Sequence Alignments from deep-Learning of Structural Alignments (SAdLSA, silent 'd'). The key idea is to implicitly learn the protein folding code from many thousands of structural alignments using experimentally determined protein structures. RESULTS: To demonstrate that the folding code was learned, we first show that SAdLSA trained on pure α-helical proteins successfully recognizes pairs of structurally related pure β-sheet protein domains. Subsequent training and benchmarking on larger, highly challenging datasets show significant improvement over established approaches. For challenging cases, SAdLSA is ∼150% better than HHsearch for generating pairwise alignments and ∼50% better for identifying the proteins with the best alignments in a sequence library. The time complexity of SAdLSA is O(N) thanks to GPU acceleration. AVAILABILITY AND IMPLEMENTATION: Datasets and source codes of SAdLSA are available free of charge for academic users at http://sites.gatech.edu/cssb/sadlsa/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mu Gao, Jeffrey Skolnick |
Bioinform. | 2 |
| 2016 | A knowledge-based approach for predicting gene-disease associationsabstractMOTIVATION: Recent advances of next-generation sequence technologies have made it possible to rapidly and inexpensively identify gene variations. Knowing the disease association of these gene variations is important for early intervention to treat deadly diseases and provide possible targets to cure these diseases. Genome-wide association studies (GWAS) have identified many individual genes associated with common diseases. To exploit the large amount of data obtained from GWAS studies and leverage our understanding of common as well as rare diseases, we have developed a knowledge-based approach to predict gene-disease associations. We first derive gene-gene mutual information by utilizing the cooccurrence of genes in known gene-disease association data. Subsequently, the mutual information is combined with known protein-protein interaction networks by a boosted tree regression method. RESULTS: The method called Know-GENE is compared with the method of random walking on the heterogeneous network using the same input data. For a set of 960 diseases, using the same training data in testing in 3-fold cross-validation, the average recall rate within the top ranked 100 genes by Know-GENE is 65.0% compared with 37.9% by the state of the art random walking on heterogeneous network. This significant improvement is mostly due to the inclusion of knowledge-based mutual information. AVAILABILITY AND IMPLEMENTATION: Predictions for genes associated with the 960 diseases are available at http://cssb2.biology.gatech.edu/knowgene CONTACT: : [email protected]. Hongyi Zhou, Jeffrey Skolnick |
Bioinform. | 2 |
| 2015 | GS-align for glycan structure alignment and similarity measurementabstractMOTIVATION: Glycans play critical roles in many biological processes, and their structural diversity is key for specific protein-glycan recognition. Comparative structural studies of biological molecules provide useful insight into their biological relationships. However, most computational tools are designed for protein structure, and despite their importance, there is no currently available tool for comparing glycan structures in a sequence order- and size-independent manner. RESULTS: A novel method, GS-align, is developed for glycan structure alignment and similarity measurement. GS-align generates possible alignments between two glycan structures through iterative maximum clique search and fragment superposition. The optimal alignment is then determined by the maximum structural similarity score, GS-score, which is size-independent. Benchmark tests against the Protein Data Bank (PDB) N-linked glycan library and PDB homologous/non-homologous N-glycoprotein sets indicate that GS-align is a robust computational tool to align glycan structures and quantify their structural similarity. GS-align is also applied to template-based glycan structure prediction and monosaccharide substitution matrix generation to illustrate its utility. AVAILABILITY AND IMPLEMENTATION: http://www.glycanstructure.org/gsalign. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hui Sun Lee, Sunhwan Jo, Srayanta Mukherjee, Jeffrey Skolnick, Wonpil Im |
Bioinform. | 5 |
| 2015 | LIGSIFT: an open-source tool for ligand structural alignment and virtual screeningabstractMOTIVATION: Shape-based alignment of small molecules is a widely used approach in computer-aided drug discovery. Most shape-based ligand structure alignment applications, both commercial and freely available ones, use the Tanimoto coefficient or similar functions for evaluating molecular similarity. Major drawbacks of using such functions are the size dependence of the score and the fact that the statistical significance of the molecular match using such metrics is not reported. RESULTS: We describe a new open-source ligand structure alignment and virtual screening (VS) algorithm, LIGSIFT, that uses Gaussian molecular shape overlay for fast small molecule alignment and a size-independent scoring function for efficient VS based on the statistical significance of the score. LIGSIFT was tested against the compounds for 40 protein targets available in the Directory of Useful Decoys and the performance was evaluated using the area under the ROC curve (AUC), the Enrichment Factor (EF) and Hit Rate (HR). LIGSIFT-based VS shows an average AUC of 0.79, average EF values of 20.8 and a HR of 59% in the top 1% of the screened library. AVAILABILITY AND IMPLEMENTATION: LIGSIFT software, including the source code, is freely available to academic users at http://cssb.biology.gatech.edu/LIGSIFT. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. CONTACT: [email protected]. Ambrish Roy, Jeffrey Skolnick |
Bioinform. | 2 |
| 2014 | Sliding of Proteins Non-specifically Bound to DNA: Brownian Dynamics Studies with Coarse-Grained Protein and DNA ModelsabstractDNA binding proteins efficiently search for their cognitive sites on long genomic DNA by combining 3D diffusion and 1D diffusion (sliding) along the DNA. Recent experimental results and theoretical analyses revealed that the proteins show a rotation-coupled sliding along DNA helical pitch. Here, we performed Brownian dynamics simulations using newly developed coarse-grained protein and DNA models for evaluating how hydrodynamic interactions between the protein and DNA molecules, binding affinity of the protein to DNA, and DNA fluctuations affect the one dimensional diffusion of the protein on the DNA. Our results indicate that intermolecular hydrodynamic interactions reduce 1D diffusivity by 30%. On the other hand, structural fluctuations of DNA give rise to steric collisions between the CG-proteins and DNA, resulting in faster 1D sliding of the protein. Proteins with low binding affinities consistent with experimental estimates of non-specific DNA binding show hopping along the CG-DNA. This hopping significantly increases sliding speed. These simulation studies provide additional insights into the mechanism of how DNA binding proteins find their target sites on the genome. Tadashi Ando, Jeffrey Skolnick |
PLoS Comput. Biol. | 2 |
| 2013 | APoc: large-scale identification of similar protein pocketsabstractMOTIVATION: Most proteins interact with small-molecule ligands such as metabolites or drug compounds. Over the past several decades, many of these interactions have been captured in high-resolution atomic structures. From a geometric point of view, most interaction sites for grasping these small-molecule ligands, as revealed in these structures, form concave shapes, or 'pockets', on the protein's surface. An efficient method for comparing these pockets could greatly assist the classification of ligand-binding sites, prediction of protein molecular function and design of novel drug compounds. RESULTS: We introduce a computational method, APoc (Alignment of Pockets), for the large-scale, sequence order-independent, structural comparison of protein pockets. A scoring function, the Pocket Similarity Score (PS-score), is derived to measure the level of similarity between pockets. Statistical models are used to estimate the significance of the PS-score based on millions of comparisons of randomly related pockets. APoc is a general robust method that may be applied to pockets identified by various approaches, such as ligand-binding sites as observed in experimental complex structures, or predicted pockets identified by a pocket-detection method. Finally, we curate large benchmark datasets to evaluate the performance of APoc and present interesting examples to demonstrate the usefulness of the method. We also demonstrate that APoc has better performance than the geometric hashing-based method SiteEngine. AVAILABILITY AND IMPLEMENTATION: The APoc software package including the source code is freely available at http://cssb.biology.gatech.edu/APoc. Mu Gao, Jeffrey Skolnick |
Bioinform. | 2 |
| 2013 | A Comprehensive Survey of Small-Molecule Binding Pockets in ProteinsabstractMany biological activities originate from interactions between small-molecule ligands and their protein targets. A detailed structural and physico-chemical characterization of these interactions could significantly deepen our understanding of protein function and facilitate drug design. Here, we present a large-scale study on a non-redundant set of about 20,000 known ligand-binding sites, or pockets, of proteins. We find that the structural space of protein pockets is crowded, likely complete, and may be represented by about 1,000 pocket shapes. Correspondingly, the growth rate of novel pockets deposited in the Protein Data Bank has been decreasing steadily over the recent years. Moreover, many protein pockets are promiscuous and interact with ligands of diverse scaffolds. Conversely, many ligands are promiscuous and interact with structurally different pockets. Through a physico-chemical and structural analysis, we provide insights into understanding both pocket promiscuity and ligand promiscuity. Finally, we discuss the implications of our study for the prediction of protein-ligand interactions based on pocket comparison. Mu Gao, Jeffrey Skolnick |
PLoS Comput. Biol. | 2 |
| 2013 | Restricted N-glycan Conformational Space in the PDB and Its Implication in Glycan Structure ModelingabstractUnderstanding glycan structure and dynamics is central to understanding protein-carbohydrate recognition and its role in protein-protein interactions. Given the difficulties in obtaining the glycan's crystal structure in glycoconjugates due to its flexibility and heterogeneity, computational modeling could play an important role in providing glycosylated protein structure models. To address if glycan structures available in the PDB can be used as templates or fragments for glycan modeling, we present a survey of the N-glycan structures of 35 different sequences in the PDB. Our statistical analysis shows that the N-glycan structures found on homologous glycoproteins are significantly conserved compared to the random background, suggesting that N-glycan chains can be confidently modeled with template glycan structures whose parent glycoproteins share sequence similarity. On the other hand, N-glycan structures found on non-homologous glycoproteins do not show significant global structural similarity. Nonetheless, the internal substructures of these N-glycans, particularly, the substructures that are closer to the protein, show significantly similar structures, suggesting that such substructures can be used as fragments in glycan modeling. Increased interactions with protein might be responsible for the restricted conformational space of N-glycan chains. Our results suggest that structure prediction/modeling of N-glycans of glycoconjugates using structure database could be effective and different modeling approaches would be needed depending on the availability of template structures. Sunhwan Jo, Hui Sun Lee, Jeffrey Skolnick, Wonpil Im |
PLoS Comput. Biol. | 3 |
| 2012 | EFICAz2.5: application of a high-precision enzyme function predictor to 396 proteomesabstractUNLABELLED: High-quality enzyme function annotation is essential for understanding the biochemistry, metabolism and disease processes of organisms. Previously, we developed a multi-component high-precision enzyme function predictor, EFICAz(2) (enzyme function inference by a combined approach). Here, we present an updated improved version, EFICAz(2.5), that is trained on a significantly larger data set of enzyme sequences and PROSITE patterns. We also present the results of the application of EFICAz(2.5) to the enzyme reannotation of 396 genomes cataloged in the ENSEMBL database. AVAILABILITY: The EFICAz(2.5) server and database is freely available with a use-friendly interface at http://cssb.biology.gatech.edu/EFICAz2.5. Jeffrey Skolnick |
Bioinform. | 2 |
| 2011 | Learning Protein Folding Energy Functionsabstractprotein folding is protein energy function design, which pertains to defining the energy of protein conformations in a way that makes folding most efficient and reliable. In this paper, we address this issue as a weight optimization problem and utilize a machine learning approach, learning-to-rank, to solve this problem. We investigate the ranking-via-classification approach, especially the RankingSVM method and compare it with the state-of-the-art approach to the problem using the MINUIT optimization package. To maintain the physicality of the results, we impose non-negativity constraints on the weights. For this we develop two efficient non-negative support vector machine (NNSVM) methods, derived from L2-norm SVM and L1-norm SVMs, respectively. We demonstrate an energy function which maintains the correct ordering with respect to structure dissimilarity to the native state more often, is more efficient and reliable for learning on large protein sets, and is qualitatively superior to the current state-of-the-art energy function. Wei Guan 0002, Arkadas Ozakin, Alexander G. Gray, Jose Borreguero, Shashi Bhushan Pandit, Anna Jagielska, Liliana Wroblewska, Jeffrey Skolnick |
ICDM | 8 |
| 2010 | iAlign: a method for the structural comparison of protein-protein interfacesabstractMOTIVATION: Protein-protein interactions play an essential role in many cellular processes. The rapid accumulation of protein-protein complex structures provides an unprecedented opportunity for comparative studies of protein-protein interactions. To facilitate such studies, it is necessary to develop an accurate and efficient computational algorithm for the comparison of protein-protein interaction modes. While there are many structural comparison approaches developed for individual proteins, very few methods are available for protein-protein complexes. RESULTS: We present a novel interface alignment method, iAlign, for the structural alignment of protein-protein interfaces. New scoring schemes for measuring interface similarity are introduced, and an iterative dynamic programming algorithm is implemented. We find that the similarity scores follow extreme value distributions. Using statistical models, we empirically estimate their statistical significance, which is in good agreement with manual classifications by human experts. Large-scale tests of iAlign were conducted on both artificial docking models and experimental structures. In a benchmark test on 1517 dimers, iAlign successfully detects biologically related, structurally similar protein-protein interfaces at a coverage percentage of 90% and an error per query of 0.05. When compared against previously published methods, iAlign is substantially more accurate and efficient. AVAILABILITY: The iAlign software package is freely available at http://cssb.biology.gatech.edu/iAlign. Mu Gao, Jeffrey Skolnick |
Bioinform. | 2 |
| 2010 | PSiFR: an integrated resource for prediction of protein structure and functionabstractUNLABELLED: In the post-genomic era, the annotation of protein function facilitates the understanding of various biological processes. To extend the range of function annotation methods to the twilight zone of sequence identity, we have developed approaches that exploit both protein tertiary structure and/or protein sequence evolutionary relationships. To serve the scientific community, we have integrated the structure prediction tools, TASSER, TASSER-Lite and METATASSER, and the functional inference tools, FINDSITE, a structure-based algorithm for binding site prediction, Gene Ontology molecular function inference and ligand screening, EFICAz(2), a sequence-based approach to enzyme function inference and DBD-hunter, an algorithm for predicting DNA-binding proteins and associated DNA-binding residues, into a unified web resource, Protein Structure and Function prediction Resource (PSiFR). AVAILABILITY AND IMPLEMENTATION: PSiFR is freely available for use on the web at http://psifr.cssb.biology.gatech.edu/ Shashi Bhushan Pandit, Michal Brylinski, Hongyi Zhou, Mu Gao, Adrian K. Arakaki, Jeffrey Skolnick |
Bioinform. | 6 |
| 2009 | FINDSITE: a combined evolution/structure-based approach to protein function predictionabstractA key challenge of the post-genomic era is the identification of the function(s) of all the molecules in a given organism. Here, we review the status of sequence and structure-based approaches to protein function inference and ligand screening that can provide functional insights for a significant fraction of the approximately 50% of ORFs of unassigned function in an average proteome. We then describe FINDSITE, a recently developed algorithm for ligand binding site prediction, ligand screening and molecular function prediction, which is based on binding site conservation across evolutionary distant proteins identified by threading. Importantly, FINDSITE gives comparable results when high-resolution experimental structures as well as predicted protein models are used. Jeffrey Skolnick, Michal Brylinski |
Briefings Bioinform. | 1 |
| 2009 | EFICAz2: enzyme function inference by a combined approach enhanced by machine learningabstractBACKGROUND: We previously developed EFICAz, an enzyme function inference approach that combines predictions from non-completely overlapping component methods. Two of the four components in the original EFICAz are based on the detection of functionally discriminating residues (FDRs). FDRs distinguish between member of an enzyme family that are homofunctional (classified under the EC number of interest) or heterofunctional (annotated with another EC number or lacking enzymatic activity). Each of the two FDR-based components is associated to one of two specific kinds of enzyme families. EFICAz exhibits high precision performance, except when the maximal test to training sequence identity (MTTSI) is lower than 30%. To improve EFICAz's performance in this regime, we: i) increased the number of predictive components and ii) took advantage of consensual information from the different components to make the final EC number assignment. RESULTS: We have developed two new EFICAz components, analogs to the two FDR-based components, where the discrimination between homo and heterofunctional members is based on the evaluation, via Support Vector Machine models, of all the aligned positions between the query sequence and the multiple sequence alignments associated to the enzyme families. Benchmark results indicate that: i) the new SVM-based components outperform their FDR-based counterparts, and ii) both SVM-based and FDR-based components generate unique predictions. We developed classification tree models to optimally combine the results from the six EFICAz components into a final EC number prediction. The new implementation of our approach, EFICAz2, exhibits a highly improved prediction precision at MTTSI < 30% compared to the original EFICAz, with only a slight decrease in prediction recall. A comparative analysis of enzyme function annotation of the human proteome by EFICAz2 and KEGG shows that: i) when both sources make EC number assignments for the same protein sequence, the assignments tend to be consistent and ii) EFICAz2 generates considerably more unique assignments than KEGG. CONCLUSION: Performance benchmarks and the comparison with KEGG demonstrate that EFICAz2 is a powerful and precise tool for enzyme function annotation, with multiple applications in genome analysis and metabolic pathway reconstruction. The EFICAz2 web service is available at: http://cssb.biology.gatech.edu/skolnick/webservice/EFICAz2/index.html. Adrian K. Arakaki, Jeffrey Skolnick |
BMC Bioinform. | 3 |
| 2009 | FINDSITELHM: A Threading-Based Approach to Ligand Homology ModelingabstractLigand virtual screening is a widely used tool to assist in new pharmaceutical discovery. In practice, virtual screening approaches have a number of limitations, and the development of new methodologies is required. Previously, we showed that remotely related proteins identified by threading often share a common binding site occupied by chemically similar ligands. Here, we demonstrate that across an evolutionarily related, but distant family of proteins, the ligands that bind to the common binding site contain a set of strongly conserved anchor functional groups as well as a variable region that accounts for their binding specificity. Furthermore, the sequence and structure conservation of residues contacting the anchor functional groups is significantly higher than those contacting ligand variable regions. Exploiting these insights, we developed FINDSITE(LHM) that employs structural information extracted from weakly related proteins to perform rapid ligand docking by homology modeling. In large scale benchmarking, using the predicted anchor-binding mode and the crystal structure of the receptor, FINDSITE(LHM) outperforms classical docking approaches with an average ligand RMSD from native of approximately 2.5 A. For weakly homologous receptor protein models, using FINDSITE(LHM), the fraction of recovered binding residues and specific contacts is 0.66 (0.55) and 0.49 (0.38) for highly confident (all) targets, respectively. Finally, in virtual screening for HIV-1 protease inhibitors, using similarity to the ligand anchor region yields significantly improved enrichment factors. Thus, the rather accurate, computationally inexpensive FINDSITE(LHM) algorithm should be a useful approach to assist in the discovery of novel biopharmaceuticals. Michal Brylinski, Jeffrey Skolnick |
PLoS Comput. Biol. | 2 |
| 2009 | From Nonspecific DNA-Protein Encounter Complexes to the Prediction of DNA-Protein InteractionsabstractDNA-protein interactions are involved in many essential biological activities. Because there is no simple mapping code between DNA base pairs and protein amino acids, the prediction of DNA-protein interactions is a challenging problem. Here, we present a novel computational approach for predicting DNA-binding protein residues and DNA-protein interaction modes without knowing its specific DNA target sequence. Given the structure of a DNA-binding protein, the method first generates an ensemble of complex structures obtained by rigid-body docking with a nonspecific canonical B-DNA. Representative models are subsequently selected through clustering and ranking by their DNA-protein interfacial energy. Analysis of these encounter complex models suggests that the recognition sites for specific DNA binding are usually favorable interaction sites for the nonspecific DNA probe and that nonspecific DNA-protein interaction modes exhibit some similarity to specific DNA-protein binding modes. Although the method requires as input the knowledge that the protein binds DNA, in benchmark tests, it achieves better performance in identifying DNA-binding sites than three previously established methods, which are based on sophisticated machine-learning techniques. We further apply our method to protein structures predicted through modeling and demonstrate that our method performs satisfactorily on protein models whose root-mean-square Calpha deviation from native is up to 5 A from their native structures. This study provides valuable structural insights into how a specific DNA-binding protein interacts with a nonspecific DNA sequence. The similarity between the specific DNA-protein interaction mode and nonspecific interaction modes may reflect an important sampling step in search of its specific DNA targets by a DNA-binding protein. Mu Gao, Jeffrey Skolnick |
PLoS Comput. Biol. | 2 |
| 2009 | A Threading-Based Method for the Prediction of DNA-Binding Proteins with Application to the Human GenomeabstractDiverse mechanisms for DNA-protein recognition have been elucidated in numerous atomic complex structures from various protein families. These structural data provide an invaluable knowledge base not only for understanding DNA-protein interactions, but also for developing specialized methods that predict the DNA-binding function from protein structure. While such methods are useful, a major limitation is that they require an experimental structure of the target as input. To overcome this obstacle, we develop a threading-based method, DNA-Binding-Domain-Threader (DBD-Threader), for the prediction of DNA-binding domains and associated DNA-binding protein residues. Our method, which uses a template library composed of DNA-protein complex structures, requires only the target protein's sequence. In our approach, fold similarity and DNA-binding propensity are employed as two functional discriminating properties. In benchmark tests on 179 DNA-binding and 3,797 non-DNA-binding proteins, using templates whose sequence identity is less than 30% to the target, DBD-Threader achieves a sensitivity/precision of 56%/86%. This performance is considerably better than the standard sequence comparison method PSI-BLAST and is comparable to DBD-Hunter, which requires an experimental structure as input. Moreover, for over 70% of predicted DNA-binding domains, the backbone Root Mean Square Deviations (RMSDs) of the top-ranked structural models are within 6.5 A of their experimental structures, with their associated DNA-binding sites identified at satisfactory accuracy. Additionally, DBD-Threader correctly assigned the SCOP superfamily for most predicted domains. To demonstrate that DBD-Threader is useful for automatic function annotation on a large-scale, DBD-Threader was applied to 18,631 protein sequences from the human genome; 1,654 proteins are predicted to have DNA-binding function. Comparison with existing Gene Ontology (GO) annotations suggests that approximately 30% of our predictions are new. Finally, we present some interesting predictions in detail. In particular, it is estimated that approximately 20% of classic zinc finger domains play a functional role not related to direct DNA-binding. Mu Gao, Jeffrey Skolnick |
PLoS Comput. Biol. | 2 |
| 2008 | Fr-TM-align: a new protein structural alignment method based on fragment alignments and the TM-scoreabstractBACKGROUND: Protein tertiary structure comparisons are employed in various fields of contemporary structural biology. Most structure comparison methods involve generation of an initial seed alignment, which is extended and/or refined to provide the best structural superposition between a pair of protein structures as assessed by a structure comparison metric. One such metric, the TM-score, was recently introduced to provide a combined structure quality measure of the coordinate root mean square deviation between a pair of structures and coverage. Using the TM-score, the TM-align structure alignment algorithm was developed that was often found to have better accuracy and coverage than the most commonly used structural alignment programs; however, there were a number of situations when this was not true. RESULTS: To further improve structure alignment quality, the Fr-TM-align algorithm has been developed where aligned fragment pairs are used to generate the initial seed alignments that are then refined using dynamic programming to maximize the TM-score. For the assessment of the structural alignment quality from Fr-TM-align in comparison to other programs such as CE and TM-align, we examined various alignment quality assessment scores such as PSI and TM-score. The assessment showed that the structural alignment quality from Fr-TM-align is better in comparison to both CE and TM-align. On average, the structural alignments generated using Fr-TM-align have a higher TM-score (~9%) and coverage (~7%) in comparison to those generated by TM-align. Fr-TM-align uses an exhaustive procedure to generate initial seed alignments. Hence, the algorithm is computationally more expensive than TM-align. CONCLUSION: Fr-TM-align, a new algorithm that employs fragment alignment and assembly provides better structural alignments in comparison to TM-align. The source code and executables of Fr-TM-align are freely downloadable at: http://cssb.biology.gatech.edu/skolnick/files/FrTMalign/. Shashi Bhushan Pandit, Jeffrey Skolnick |
BMC Bioinform. | 2 |
| 2006 | Structure Modeling of All Identified G Protein-Coupled Receptors in the Human GenomeabstractG protein-coupled receptors (GPCRs), encoded by about 5% of human genes, comprise the largest family of integral membrane proteins and act as cell surface receptors responsible for the transduction of endogenous signal into a cellular response. Although tertiary structural information is crucial for function annotation and drug design, there are few experimentally determined GPCR structures. To address this issue, we employ the recently developed threading assembly refinement (TASSER) method to generate structure predictions for all 907 putative GPCRs in the human genome. Unlike traditional homology modeling approaches, TASSER modeling does not require solved homologous template structures; moreover, it often refines the structures closer to native. These features are essential for the comprehensive modeling of all human GPCRs when close homologous templates are absent. Based on a benchmarked confidence score, approximately 820 predicted models should have the correct folds. The majority of GPCR models share the characteristic seven-transmembrane helix topology, but 45 ORFs are predicted to have different structures. This is due to GPCR fragments that are predominantly from extracellular or intracellular domains as well as database annotation errors. Our preliminary validation includes the automated modeling of bovine rhodopsin, the only solved GPCR in the Protein Data Bank. With homologous templates excluded, the final model built by TASSER has a global C(alpha) root-mean-squared deviation from native of 4.6 angstroms, with a root-mean-squared deviation in the transmembrane helix region of 2.1 angstroms. Models of several representative GPCRs are compared with mutagenesis and affinity labeling data, and consistent agreement is demonstrated. Structure clustering of the predicted models shows that GPCRs with similar structures tend to belong to a similar functional class even when their sequences are diverse. These results demonstrate the usefulness and robustness of the in silico models for GPCR functional analysis. All predicted GPCR models are freely available for noncommercial users on our Web site (http://www.bioinformatics.buffalo.edu/GPCR). Yang Zhang 0040, Mark E. DeVries, Jeffrey Skolnick |
PLoS Comput. Biol. | 3 |
| 2006 | Correction: Structure Modeling of All Identified G Protein-Coupled Receptors in the Human Genome
Yang Zhang 0040, Mark E. DeVries, Jeffrey Skolnick |
PLoS Comput. Biol. | 3 |
| 2004 | Large-scale assessment of the utility of low-resolution protein structures for biochemical function assignmentabstractMOTIVATION: Several protein function prediction methods employ structural features captured in three-dimensional (3D) descriptors of biologically relevant sites. These methods are successful when applied to high-resolution structures, but their detection ability in lower resolution predicted structures has only been tested for a few cases. RESULTS: A method that automatically generates a library of 3D functional descriptors for the structure-based prediction of enzyme active sites (automated functional templates, 593 in total for 162 different enzymes), based on functional and structural information automatically extracted from public databases, has been developed and evaluated using decoy structures. The applicability to predicted structures was investigated by analyzing decoys of varying quality, derived from enzyme native structures. For 35% of decoy structures, our method identifies the active site in models having 3-4 A coordinate root mean square deviation from the native structure, a quality that is reachable using state of the art protein structure prediction algorithms. AVAILABILITY: See http://www.bioinformatics.buffalo.edu/resources/aft/ Adrian K. Arakaki, Yang Zhang 0040, Jeffrey Skolnick |
Bioinform. | 3 |
| 2001 | BioMolQuest: integrated database-based retrieval of protein structural and functional informationabstractAbstract Motivation: Information about a particular protein or protein family is usually distributed among multiple databases and often in more than one entry in each database. Retrieval and organization of this information can be a laborious task. This task is complicated even further by the existence of alternative terms for the same concept. Results: The PDB, SWISS-PROT, ENZYME, and CATH databases have been imported into a combined relational database, BioMolQuest. A powerful search engine has been built using this database as a back end. The search engine achieves significant improvements in query performance by automatically utilizing cross-references between the legacy databases. The results of the queries are presented in an organized, hierarchical way. Availability: http://bioinformatics.danforthcenter.org/yury/public/home.html. Unavailable for commercial users. Contact: [email protected] * To whom correspondence should be addressed. Yury V. Bukhman, Jeffrey Skolnick |
Bioinform. | 2 |
| 1998 | A self-consistent field optimization approach to build energetically and geometrically correct lattice models of proteinsabstractLattice modeling of proteins is commonly used in studying the protein folding problem.A finite number of possible conformations of lattice models enormously facilitates exploration of the conformational space.In this work, we suggest a method to search for the optimal lattice models that reproduced the off-lattice structures with minimal errors in geometry and energetics.The method is based on the self-consistent field optimization of a combined pseudoenergy function that includes two force fields: an "interaction field", which drives the residues to optimize the chain energy, and a "geometrical field", which attracts the residues towards their native positions. Boris A. Reva, Alexei V. Finkelstein, Jeffrey Skolnick |
RECOMB | 3 |
| 1994 | Flexible algorithm for direct multiple alignment of protein structures and sequencesabstractThe recently described equivalence between the alignment of two proteins and a conformation of a lattice chain on a two-dimensional square lattice is extended to multiple alignments. The search for the optimal multiple alignment between several proteins, which is equivalent to finding the energy minimum in the conformational space of a multi-dimensional lattice chain, is studied by the Monte Carlo approach. This method, while not deterministic, and for two-dimensional problems slower than dynamic programming, can accept arbitrary scoring functions, including non-local ones, and its speed decreases slowly with increasing number of dimensions. For the local scoring functions, the MC algorithm can also reproduce known exact solutions for the direct multiple alignments. As illustrated by examples, both for structure- and sequence-based alignments, direct multi-dimensional alignments are able to capture weak similarities between divergent families much better than ones built from pairwise alignments by a hierarchical approach. Adam Godzik, Jeffrey Skolnick |
Comput. Appl. Biosci. | 2 |