EDBT 2026 Demo / reviewers in the wild / expert
Qian Cong
dblp:62/8321
· DBLP profile ↗
12ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0002-8909-0414ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reciprocal best matching: a new pipeline for scoring models with unknown stoichiometry in CASP experimentsabstractBACKGROUND: Accurate prediction of protein complex structures remains a significant challenge, particularly when stoichiometry information is unavailable. In the recent Critical Assessment of Structure Prediction Round XVI (CASP16), the “Phase 0” challenge was introduced to stimulate progress in this area. However, existing evaluation tools, such as OpenStructure, might introduce systematic biases when evaluating models with stoichiometries different from the target, sometimes favoring those with excess subunits and inflating scores for models with incorrect stoichiometries. RESULTS: To address this issue, we developed the Reciprocal Best Matching (RBM) pipeline. RBM compares predicted and target structures by bidirectionally matching interfaces and assigning penalizations to unmatched interfaces. This approach penalizes incorrect stoichiometries in a consistent and unbiased manner while preserving strong correlation with established CASP metrics. Application of RBM in CASP16 assessments revealed improved discrimination between correctly and incorrectly stoichiometric models. CONCLUSIONS: Our method, RBM, could correct the systematic bias in the existing assessment protocol for protein complex structure prediction without stoichiometry information. We provide a standalone software implementation of our RBM pipeline to stimulate further method development in protein complex structure prediction and to support future CASP experiments. Rongqing Yuan, Qian Cong |
BMC Bioinform. | 3 |
| 2026 | ECOD: Classification of domains in AFDB Swiss-Prot structure predictionsabstractThe development of highly accurate protein structure prediction algorithms has led to an explosion of structural data, transforming our understanding of protein structure-function relationships across diverse organisms. Domain classifications such as the Evolutionary Classification of Protein Domains (ECOD) have incorporated these computational predictions alongside experimental structures to create comprehensive resources for the research community. The AlphaFold Protein Structure Database (AFDB) plays a unique role, providing millions of predicted structures that ECOD has systematically classified for human proteins, small pathogens, and reference proteomes. Here, we extend this classification framework to the UniProtKB/Swiss-Prot dataset, applying the Domain Parser for AlphaFold Models (DPAM) pipeline to classify domains from over 542,000 Swiss-Prot protein structure predictions, resulting in more than 1,032,000 classified domains. These domains span 3,493 ECOD topologies and display high assignment confidence (mean DPAM probability: 0.992), with extensive taxonomic and functional diversity. Notably, over 100,000 domains lack existing Pfam mappings, reflecting the extended sensitivity of structure-based classification and identifying domain groups not yet captured by sequence-based profiles. These results significantly expand ECOD's coverage into a functionally and taxonomically diverse protein space, anchoring high-confidence structure predictions in an evolutionary framework. By integrating Swiss-Prot predictions, we enhance the utility and interpretability of AlphaFold models and establish a foundation for future large-scale, functionally informed domain classifications. R. Dustin Schaeffer, Jing Zhang 0115, Qian Cong, Nick V. Grishin |
PLoS Comput. Biol. | 3 |
| 2025 | DPAM-AI: a domain parser for AlphaFold models powered by artificial intelligenceabstractMOTIVATION: Due to the breakthrough in protein structure prediction by AlphaFold, the scientific community has access to 200 million predicted protein structures with near-atomic accuracy from the AlphaFold protein structure DataBase (AFDB), covering nearly the entire protein universe. Segmenting these models into domains and classifying them into an evolutionary hierarchy hold tremendous potential for unraveling essential insights into protein function. RESULTS: We introduce DPAM-AI, a Domain Parser for AlphaFold Models based on Artificial Intelligence. DPAM-AI utilizes a convolutional neural network trained with previously classified domains in the Evolutionary Classification Of protein Domains (ECOD) database. DPAM-AI integrates inter-residue distances, predicted aligned errors, and sequence and structural alignments to previously classified domains detected via sequence (HHsuite) and structural (Dali) similarity searches. DPAM-AI has demonstrated its power through rigorous tests, excelling in several benchmark sets compared to its predecessor, DPAM, and other recently published domain parsers, Merizo and Chainsaw. We applied DPAM-AI to representative AFDB models for proteins classified in Pfam. We obtained representative 3D structures for 18 487 (89%) of the 20 795 Pfam families. The remaining families either (i) belong to viral proteins that were excluded from AFDB or (ii) do not adopt globular 3D structures. Our structure-aware domain delineation uncovered a considerable fraction (15%) of Pfam domains containing multiple structural and evolutionary units and refined the boundaries for over half. AVAILABILITY AND IMPLEMENTATION: Pfam and corresponding DPAM-AI domains are at http://prodata.swmed.edu/DPAM-pfam/. Our code is deposited at https://github.com/Jsauce5p/DPAM/tree/dpam_ai, and updates will be released through https://github.com/CongLabCode/DPAM. Jesse Durham, Jing Zhang 0115, Richard D. Schaeffer, Qian Cong |
Bioinform. | 4 |
| 2025 | Using evolutionary context to classify difficult protein foldsabstractRecent advances in protein structure prediction, such as AlphaFold2, have enabled identification of vast numbers of putative novel protein domains across the sequence space, many of which adopt structures dissimilar to known folds. Based on structural segmentation and classification, the Encyclopedia of Domains (TED) project recently cataloged more than 7400 low-symmetry, structure-based domains as candidate novel-fold (CNF) domains. To place these domains in their broader evolutionary and structural context, we applied DPAM (Domain Parser for AlphaFold Models), a complementary method that combines AlphaFold-derived confidence metrics with sensitive sequence and structure similarity searches, to parse domains for the AlphaFold models of proteins containing TED CNF domains. We identified 8044 DPAM domains with significant overlap with TED CNF domains, among which 2490 were confidently assigned to entries in the ECOD (Evolutionary Classification of protein Domains) structural classification hierarchy. Our results suggest that a substantial subset of TED candidate novel-fold domains are distant homologs of existing ECOD domains. Comparison of domain boundaries between TED and DPAM showed varied patterns: more than one-third of cases featured TED CNF domains largely embedded within DPAM domains-often representing insertions or extensions into enzymatic or repeat folds. A smaller fraction (17%) exhibited consistent domain boundaries between TED and DPAM. These consistently defined domains are often characterized by significant structural diversity, including long insertions and duplications. An even smaller subset showed the reverse relationship, with DPAM domains largely embedded within TED CNF domains. In these cases, DPAM effectively separated multiple structural units that TED grouped as single domains. Together, these findings highlight the complementarity of structural and evolutionary approaches for domain annotation and demonstrate the power of integrative methods, such as DPAM, in refining the classification of challenging protein folds and uncovering distant evolutionary relationships. Jimin Pei, R. Dustin Schaeffer, Qian Cong, Nick V. Grishin |
PLoS Comput. Biol. | 3 |
| 2024 | ECOD domain classification of 48 whole proteomes from AlphaFold Structure Database using DPAM2abstractProtein structure prediction has now been deployed widely across several different large protein sets. Large-scale domain annotation of these predictions can aid in the development of biological insights. Using our Evolutionary Classification of Protein Domains (ECOD) from experimental structures as a basis for classification, we describe the detection and cataloging of domains from 48 whole proteomes deposited in the AlphaFold Database. On average, we can provide positive classification (either of domains or other identifiable non-domain regions) for 90% of residues in all proteomes. We classified 746,349 domains from 536,808 proteins comprised of over 226,424,000 amino acid residues. We examine the varying populations of homologous groups in both eukaryotes and bacteria. In addition to containing a higher fraction of disordered regions and unassigned domains, eukaryotes show a higher proportion of repeated proteins, both globular and small repeats. We enumerate those highly populated domains that are shared in both eukaryotes and bacteria, such as the Rossmann domains, TIM barrels, and P-loop domains. Additionally, we compare the sampling of homologous groups from this whole proteome set against our stable ECOD reference and discuss groups that have been enriched by structure predictions. Finally, we discuss the implication of these results for protein target selection for future classification strategies for very large protein sets. R. Dustin Schaeffer, Jing Zhang 0115, Kirill E. Medvedev, Lisa N. Kinch, Qian Cong, Nick V. Grishin |
PLoS Comput. Biol. | 5 |
| 2022 | Human mitochondrial protein complexes revealed by large-scale coevolution analysis and deep learning-based structure modelingabstractMOTIVATION: Recent development of deep-learning methods has led to a breakthrough in the prediction accuracy of 3D protein structures. Extending these methods to protein pairs is expected to allow large-scale detection of protein-protein interactions (PPIs) and modeling protein complexes at the proteome level. RESULTS: We applied RoseTTAFold and AlphaFold, two of the latest deep-learning methods for structure predictions, to analyze coevolution of human proteins residing in mitochondria, an organelle of vital importance in many cellular processes including energy production, metabolism, cell death and antiviral response. Variations in mitochondrial proteins have been linked to a plethora of human diseases and genetic conditions. RoseTTAFold, with high computational speed, was used to predict the coevolution of about 95% of mitochondrial protein pairs. Top-ranked pairs were further subject to modeling of the complex structures by AlphaFold, which also produced contact probability with high precision and in many cases consistent with RoseTTAFold. Most top-ranked pairs with high contact probability were supported by known PPIs and/or similarities to experimental structural complexes. For high-scoring pairs without experimental complex structures, our coevolution analyses and structural models shed light on the details of their interfaces, including CHCHD4-AIFM1, MTERF3-TRUB2, FMC1-ATPAF2 and ECSIT-NDUFAF1. We also identified novel PPIs (PYURF-NDUFAF5, LYRM1-MTRF1L and COA8-COX10) for several proteins without experimentally characterized interaction partners, leading to predictions of their molecular functions and the biological processes they are involved in. AVAILABILITY AND IMPLEMENTATION: Data of mitochondrial proteins and their interactions are available at: http://conglab.swmed.edu/mitochondria. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jimin Pei, Jing Zhang 0115, Qian Cong |
Bioinform. | 3 |
| 2020 | Protein contact prediction using metagenome sequence data and residual neural networksabstractMOTIVATION: Almost all protein residue contact prediction methods rely on the availability of deep multiple sequence alignments (MSAs). However, many proteins from the poorly populated families do not have sufficient number of homologs in the conventional UniProt database. Here we aim to solve this issue by exploring the rich sequence data from the metagenome sequencing projects. RESULTS: Based on the improved MSA constructed from the metagenome sequence data, we developed MapPred, a new deep learning-based contact prediction method. MapPred consists of two component methods, DeepMSA and DeepMeta, both trained with the residual neural networks. DeepMSA was inspired by the recent method DeepCov, which was trained on 441 matrices of covariance features. By considering the symmetry of contact map, we reduced the number of matrices to 231, which makes the training more efficient in DeepMSA. Experiments show that DeepMSA outperforms DeepCov by 10-13% in precision. DeepMeta works by combining predicted contacts and other sequence profile features. Experiments on three benchmark datasets suggest that the contribution from the metagenome sequence data is significant with P-values less than 4.04E-17. MapPred is shown to be complementary and comparable the state-of-the-art methods. The success of MapPred is attributed to three factors: the deeper MSA from the metagenome sequence data, improved feature design in DeepMSA and optimized training by the residual neural networks. AVAILABILITY AND IMPLEMENTATION: http://yanglab.nankai.edu.cn/mappred/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qi Wu 0016, Zhen-Ling Peng, Ivan Anishchenko, Qian Cong, David Baker 0001, Jianyi Yang 0002 |
Bioinform. | 4 |
| 2013 | M2SG: mapping human disease-related genetic variants to protein sequences and genomic lociabstractSUMMARY: Online Mendelian Inheritance in Man (OMIM) is a manually curated compendium of human genetic variants and the corresponding phenotypes, mostly human diseases. Instead of directly documenting the native sequences for gene entries, OMIM links its entries to protein and DNA sequences in other databases. However, because of the existence of gene isoforms and errors in OMIM records, mapping a specific OMIM mutation to its corresponding protein sequence is not trivial. Combining computer programs and extensive manual curation of OMIM full-text descriptions and original literature, we mapped 98% of OMIM amino acid substitutions (AASs) and all SwissProt Variant (SwissVar) disease-related AASs to reference sequences and confidently mapped 99.96% of all AASs to the genomic loci. Based on the results, we developed an online database and interactive web server (M2SG) to (i) retrieve the mapped OMIM and SwissVar variants for a given protein sequence; and (ii) obtain related proteins and mutations for an input disease phenotype. This database will be useful for analyzing sequences, understanding the effect of mutations, identifying important genetic variations and designing experiments on a protein of interest. AVAILABILITY AND IMPLEMENTATION: The database and web server are freely available at http://prodata.swmed.edu/M2S/mut2seq.cgi. Renkai Ji, Qian Cong, Nick V. Grishin |
Bioinform. | 2 |
| 2013 | Seq2Ref: a web server to facilitate functional interpretationabstractBACKGROUND: The size of the protein sequence database has been exponentially increasing due to advances in genome sequencing. However, experimentally characterized proteins only constitute a small portion of the database, such that the majority of sequences have been annotated by computational approaches. Current automatic annotation pipelines inevitably introduce errors, making the annotations unreliable. Instead of such error-prone automatic annotations, functional interpretation should rely on annotations of 'reference proteins' that have been experimentally characterized or manually curated. RESULTS: The Seq2Ref server uses BLAST to detect proteins homologous to a query sequence and identifies the reference proteins among them. Seq2Ref then reports publications with experimental characterizations of the identified reference proteins that might be relevant to the query. Furthermore, a plurality-based rating system is developed to evaluate the homologous relationships and rank the reference proteins by their relevance to the query. CONCLUSIONS: The reference proteins detected by our server will lend insight into proteins of unknown function and provide extensive information to develop in-depth understanding of uncharacterized proteins. Seq2Ref is available at: http://prodata.swmed.edu/seq2ref. Qian Cong, Lisa N. Kinch, Nick V. Grishin |
BMC Bioinform. | 2 |
| 2011 | An automatic method for CASP9 free modeling structure prediction assessmentabstractMOTIVATION: Manual inspection has been applied to and is well accepted for assessing critical assessment of protein structure prediction (CASP) free modeling (FM) category predictions over the years. Such manual assessment requires expertise and significant time investment, yet has the problems of being subjective and unable to differentiate models of similar quality. It is beneficial to incorporate the ideas behind manual inspection to an automatic score system, which could provide objective and reproducible assessment of structure models. RESULTS: Inspired by our experience in CASP9 FM category assessment, we developed an automatic superimposition independent method named Quality Control Score (QCS) for structure prediction assessment. QCS captures both global and local structural features, with emphasis on global topology. We applied this method to all FM targets from CASP9, and overall the results showed the best agreement with Manual Inspection Scores among automatic prediction assessment methods previously applied in CASPs, such as Global Distance Test Total Score (GDT_TS) and Contact Score (CS). As one of the important components to guide our assessment of CASP9 FM category predictions, this method correlates well with other scoring methods and yet is able to reveal good-quality models that are missed by GDT_TS. AVAILABILITY: The script for QCS calculation is available at http://prodata.swmed.edu/QCS/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qian Cong, Lisa N. Kinch, Jimin Pei, Shuoyong Shi, Vyacheslav N. Grishin, Nick V. Grishin |
Bioinform. | 1 |
| 2010 | Structural Differences between Proteins with Similar SequencesabstractSimilarity between protein sequences is usually predictive of similarity in structures. However, in some rare cases protein domains with significant sequence similarity adopt different structures. Here, we carry out a survey of protein domain pairs with high sequence similarity (measured by HHsearch probability) and low structural similarity (measured by Dali Z-score), aiming to identify the reasons for this discordance. Besides methodological problems with either sequences or structures of domains, we find and describe novel examples of homologs with structural changes. Qian Cong, Bong-Hyun Kim, Lisa N. Kinch, Nick V. Grishin |
BIBE | 1 |
| 2010 | HangOut: generating clean PSI-BLAST profiles for domains with long insertionsabstractUNLABELLED: Profile-based similarity search is an essential step in structure-function studies of proteins. However, inclusion of non-homologous sequence segments into a profile causes its corruption and results in false positives. Profile corruption is common in multidomain proteins, and single domains with long insertions are a significant source of errors. We developed a procedure (HangOut) that, for a single domain with specified insertion position, cleans erroneously extended PSI-BLAST alignments to generate better profiles. AVAILABILITY: HangOut is implemented in Python 2.3 and runs on all Unix-compatible platforms. The source code is available under the GNU GPL license at http://prodata.swmed.edu/HangOut/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bong-Hyun Kim, Qian Cong, Nick V. Grishin |
Bioinform. | 2 |