VLDB 2026 Research / reviewers in the wild / expert
Haicang Zhang
dblp:138/0439
· DBLP profile ↗
15ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0001-6268-4258ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
7 papers |
Bioinformatics and computational biology · 100% | |
| Artificial intelligence
4 papers |
Generative modeling · 88% 3D vision · 12% |
Topics — the 19 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
protein design |
2.2 | 3 | 2024 | Antibody Design Using a Score-based Diffusion Model Guided by Evolutionary, Physical and Geometric Constraints · ICML 2024 CarbonNovo: Joint Design of Protein Structure and Sequence Using a Unified Energy-based Model · ICML 2024 Accurate and efficient protein sequence design through learning concise local environment of residues · Bioinform. 2023 |
Bioinformatics and computational biology › protein design
antibody design |
1.6 | 2 | 2025 | Multi-objective antibody design with constrained preference optimization · ICLR 2025 Antibody Design Using a Score-based Diffusion Model Guided by Evolutionary, Physical and Geometric Constraints · ICML 2024 |
Machine learning › Generative modeling › diffusion model
score-based generative model |
1.5 | 2 | 2024 | Antibody Design Using a Score-based Diffusion Model Guided by Evolutionary, Physical and Geometric Constraints · ICML 2024 CarbonNovo: Joint Design of Protein Structure and Sequence Using a Unified Energy-based Model · ICML 2024 |
Bioinformatics and computational biology › protein design › antibody design
sequence-structure co-design |
1.5 | 2 | 2024 | Antibody Design Using a Score-based Diffusion Model Guided by Evolutionary, Physical and Geometric Constraints · ICML 2024 CarbonNovo: Joint Design of Protein Structure and Sequence Using a Unified Energy-based Model · ICML 2024 |
Machine learning › Generative modeling
diffusion model |
1.4 | 2 | 2024 | Antibody Design Using a Score-based Diffusion Model Guided by Evolutionary, Physical and Geometric Constraints · ICML 2024 Predicting mutational effects on protein-protein binding via a side-chain diffusion probabilistic model · NeurIPS 2023 |
Bioinformatics and computational biology › protein design
computational protein design |
0.9 | 1 | 2025 | Multi-objective antibody design with constrained preference optimization · ICLR 2025 |
Machine learning › Generative modeling
energy-based model |
0.8 | 1 | 2024 | CarbonNovo: Joint Design of Protein Structure and Sequence Using a Unified Energy-based Model · ICML 2024 |
Bioinformatics and computational biology › protein design
de novo protein design |
0.8 | 1 | 2024 | CarbonNovo: Joint Design of Protein Structure and Sequence Using a Unified Energy-based Model · ICML 2024 |
Computer vision › 3D vision
molecular structure |
0.7 | 1 | 2023 | Predicting mutational effects on protein-protein binding via a side-chain diffusion probabilistic model · NeurIPS 2023 |
Machine learning › Generative modeling › diffusion model › geometric diffusion model
riemannian diffusion model |
0.7 | 1 | 2023 | Predicting mutational effects on protein-protein binding via a side-chain diffusion probabilistic model · NeurIPS 2023 |
Bioinformatics and computational biology
protein engineering |
0.7 | 1 | 2023 | Predicting mutational effects on protein-protein binding via a side-chain diffusion probabilistic model · NeurIPS 2023 |
Bioinformatics and computational biology › protein design
protein sequence design |
0.7 | 1 | 2023 | Accurate and efficient protein sequence design through learning concise local environment of residues · Bioinform. 2023 |
Bioinformatics and computational biology › statistical genetics
variant effect prediction |
0.7 | 1 | 2023 | Predicting mutational effects on protein-protein binding via a side-chain diffusion probabilistic model · NeurIPS 2023 |
Bioinformatics and computational biology
protein structure prediction |
0.5 | 2 | 2017 | Improving protein fold recognition by extracting fold-specific features from predicted residue-residue contacts · Bioinform. 2017 FALCON@home: a high-throughput protein structure prediction server based on remote homologue recognition · Bioinform. 2016 |
Bioinformatics and computational biology › protein structure prediction › template-based modeling
fold recognition |
0.3 | 1 | 2017 | Improving protein fold recognition by extracting fold-specific features from predicted residue-residue contacts · Bioinform. 2017 |
Bioinformatics and computational biology › sequence analysis › homology detection
remote homology detection |
0.2 | 1 | 2016 | FALCON@home: a high-throughput protein structure prediction server based on remote homologue recognition · Bioinform. 2016 |
Bioinformatics and computational biology › protein structure prediction
template-based modeling |
0.2 | 1 | 2016 | FALCON@home: a high-throughput protein structure prediction server based on remote homologue recognition · Bioinform. 2016 |
Bioinformatics and computational biology › protein structure prediction
residue contact prediction |
0.1 | 1 | 2017 | Improving protein fold recognition by extracting fold-specific features from predicted residue-residue contacts · Bioinform. 2017 |
Distributed systems › grid computing
volunteer computing |
0.1 | 1 | 2016 | FALCON@home: a high-throughput protein structure prediction server based on remote homologue recognition · Bioinform. 2016 |
Methods — techniques the papers use, named apart from their topics
protein language model · 4.8primal-dual optimization · 1.7constrained preference optimization · 1.7score-based diffusion · 1.5physical constraints · 1.5markov random field · 1.5geometric constraints · 1.5diffusion · 1.5riemannian diffusion · 1.3representation learning · 1.3threading · 0.2conserved region extraction · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-objective antibody design with constrained preference optimizationabstractAntibody design is crucial for developing therapies against diseases such as cancer and viral infections. Recent deep generative models have significantly advanced computational antibody design, particularly in enhancing binding affinity to target antigens. However, beyond binding affinity, antibodies should exhibit other favorable biophysical properties such as non-antigen binding specificity and low self-association, which are important for antibody developability and clinical safety. To address this challenge, we propose AbNovo, a framework that leverages constrained preference optimization for multi-objective antibody design. First, we pre-train an antigen-conditioned generative model for antibody structure and sequence co-design. Then, we fine-tune the model using binding affinity as a reward while enforcing explicit constraints on other biophysical properties. Specifically, we model the physical binding energy with continuous rewards rather than pairwise preferences and explore a primal-and-dual approach for constrained optimization. Additionally, we incorporate a structure-aware protein language model to mitigate the issue of limited training data. Evaluated on independent test sets, AbNovo outperforms existing methods in metrics of binding affinity such as Rosetta binding energy and evolutionary plausibility, as well as in metrics for other biophysical properties like stability and specificity. Milong Ren, ZaiKai He, Haicang Zhang |
ICLR | 3 |
| 2025 | A survey on deep learning-based algorithms for the traveling salesman problemabstractAbstract This paper presents an overview of deep learning (DL)-based algorithms designed for solving the traveling salesman problem (TSP), categorizing them into four categories: end-to-end construction algorithms, end-to-end improvement algorithms, direct hybrid algorithms, and large language model (LLM)-based hybrid algorithms. We introduce the principles and methodologies of these algorithms, outlining their strengths and limitations through experimental comparisons. End-to-end construction algorithms employ neural networks to generate solutions from scratch, demonstrating rapid solving speed but often yielding subpar solutions. Conversely, end-to-end improvement algorithms iteratively refine initial solutions, achieving higher-quality outcomes but necessitating longer computation times. Direct hybrid algorithms directly integrate deep learning with heuristic algorithms, showcasing robust solving performance and generalization capability. LLM-based hybrid algorithms leverage LLMs to autonomously generate and refine heuristics, showing promising performance despite being in early developmental stages. In the future, further integration of deep learning techniques, particularly LLMs, with heuristic algorithms and advancements in interpretability and generalization will be pivotal trends in TSP algorithm design. These endeavors aim to tackle larger and more complex real-world instances while enhancing algorithm reliability and practicality. This paper offers insights into the evolving landscape of DL-based TSP solving algorithms and provides a perspective for future research directions. Jingyan Sui, Shizhe Ding, Xulin Huang, Boyang Xia, Zhenxin Ding, Liming Xu, Haicang Zhang, Chungong Yu, Dongbo Bu |
Frontiers Comput. Sci. | 9 |
| 2024 | CarbonNovo: Joint Design of Protein Structure and Sequence Using a Unified Energy-based ModelabstractDe novo protein design aims to create novel protein structures and sequences unseen in nature. Recent structure-oriented design methods typically employ a two-stage strategy, where structure design and sequence design modules are trained separately, and the backbone structures and sequences are generated sequentially in inference. While diffusion-based generative models like RFdiffusion show great promise in structure design, they face inherent limitations within the two-stage framework. First, the sequence design module risks overfitting, as the accuracy of the generated structures may not align with that of the crystal structures used for training. Second, the sequence design module lacks interaction with the structure design module to further optimize the generated structures. To address these challenges, we propose CarbonNovo, a unified energy-based model for jointly generating protein structure and sequence. Specifically, we leverage a score-based generative model and Markov Random Fields for describing the energy landscape of protein structure and sequence. In CarbonNovo, the structure and sequence design module communicates at each diffusion step, encouraging the generation of more coherent structure-sequence pairs. Moreover, the unified framework allows for incorporating the protein language models as evolutionary constraints for generated proteins. The rigorous evaluation demonstrates that CarbonNovo outperforms two-stage methods across various metrics, including designability, novelty, sequence plausibility, and Rosetta Energy. Milong Ren, Haicang Zhang |
ICML | 3 |
| 2024 | Antibody Design Using a Score-based Diffusion Model Guided by Evolutionary, Physical and Geometric ConstraintsabstractAntibodies are central proteins in adaptive immune responses, responsible for protecting against viruses and other pathogens. Rational antibody design has proven effective in the diagnosis and treatment of various diseases like cancers and virus infections. While recent diffusion-based generative models show promise in designing antigen-specific antibodies, the primary challenge lies in the scarcity of labeled antibody-antigen complex data and binding affinity data. We present AbX, a new score-based diffusion generative model guided by evolutionary, physical, and geometric constraints for antibody design. These constraints serve to narrow the search space and provide priors for plausible antibody sequences and structures. Specifically, we leverage a pre-trained protein language model as priors for evolutionary plausible antibodies and introduce additional training objectives for geometric and physical constraints like van der Waals forces. Furthermore, as far as we know, AbX is the first score-based diffusion model with continuous timesteps for antibody design, jointly modeling the discrete sequence space and the $\mathrm{SE}(3)$ structure space. Evaluated on two independent testing sets, we show that AbX outperforms other published methods, achieving higher accuracy in sequence and structure generation and enhanced antibody-antigen binding affinity. Ablation studies highlight the clear contributions of the introduced constraints to antibody design. Milong Ren, Haicang Zhang |
ICML | 3 |
| 2024 | GPCR-BSD: a database of binding sites of human G-protein coupled receptors under diverse statesabstractG-protein coupled receptors (GPCRs), the largest family of membrane proteins in human body, involve a great variety of biological processes and thus have become highly valuable drug targets. By binding with ligands (e.g., drugs), GPCRs switch between active and inactive conformational states, thereby performing functions such as signal transmission. The changes in binding pockets under different states are important for a better understanding of drug-target interactions. Therefore it is critical, as well as a practical need, to obtain binding sites in human GPCR structures. We report a database (called GPCR-BSD) that collects 127,990 predicted binding sites of 803 GPCRs under active and inactive states (thus 1,606 structures in total). The binding sites were identified from the predicted GPCR structures by executing three geometric-based pocket prediction methods, fpocket, CavityPlus and GHECOM. The server provides query, visualization, and comparison of the predicted binding sites for both GPCR predicted and experimentally determined structures recorded in PDB. We evaluated the identified pockets of 132 experimentally determined human GPCR structures in terms of pocket residue coverage, pocket center distance and redocking accuracy. The evaluation showed that fpocket and CavityPlus methods performed better and successfully predicted orthosteric binding sites in over 60% of the 132 experimentally determined structures. The GPCR Binding Site database is freely accessible at https://gpcrbs.bigdata.jcmsc.cn . This study not only provides a systematic evaluation of the commonly-used fpocket and CavityPlus methods for the first time but also meets the need for binding site information in GPCR studies. Xiaonong Li, Liangliang Zhou, Chungong Yu, Haicang Zhang, Dongbo Bu, Xinmiao Liang |
BMC Bioinform. | 6 |
| 2023 | Predicting mutational effects on protein-protein binding via a side-chain diffusion probabilistic modelabstractMany crucial biological processes rely on networks of protein-protein interactions. Predicting the effect of amino acid mutations on protein-protein binding is important in protein engineering, including therapeutic discovery. However, the scarcity of annotated experimental data on binding energy poses a significant challenge for developing computational approaches, particularly deep learning-based methods. In this work, we propose SidechainDiff, a novel representation learning-based approach that leverages unlabelled experimental protein structures. SidechainDiff utilizes a Riemannian diffusion model to learn the generative process of side-chain conformations and can also give the structural context representations of mutations on the protein-protein interface. Leveraging the learned representations, we achieve state-of-the-art performance in predicting the mutational effects on protein-protein binding. Furthermore, SidechainDiff is the first diffusion-based generative model for side-chains, distinguishing it from prior efforts that have predominantly focused on the generation of protein backbone structures. Milong Ren, Chungong Yu, Dongbo Bu, Haicang Zhang |
NeurIPS | 6 |
| 2023 | Accurate and efficient protein sequence design through learning concise local environment of residuesabstractMOTIVATION: Computational protein sequence design has been widely applied in rational protein engineering and increasing the design accuracy and efficiency is highly desired. RESULTS: Here, we present ProDESIGN-LE, an accurate and efficient approach to protein sequence design. ProDESIGN-LE adopts a concise but informative representation of the residue's local environment and trains a transformer to learn the correlation between local environment of residues and their amino acid types. For a target backbone structure, ProDESIGN-LE uses the transformer to assign an appropriate residue type for each position based on its local environment within this structure, eventually acquiring a designed sequence with all residues fitting well with their local environments. We applied ProDESIGN-LE to design sequences for 68 naturally occurring and 129 hallucinated proteins within 20 s per protein on average. The designed proteins have their predicted structures perfectly resembling the target structures with a state-of-the-art average TM-score exceeding 0.80. We further experimentally validated ProDESIGN-LE by designing five sequences for an enzyme, chloramphenicol O-acetyltransferase type III (CAT III), and recombinantly expressing the proteins in Escherichia coli. Of these proteins, three exhibited excellent solubility, and one yielded monomeric species with circular dichroism spectra consistent with the natural CAT III protein. AVAILABILITY AND IMPLEMENTATION: The source code of ProDESIGN-LE is available at https://github.com/bigict/ProDESIGN-LE. Bin Huang 0022, Tingwen Fan, Kaiyue Wang, Haicang Zhang, Chungong Yu, Shuyu Nie, Yangshuo Qi, Wei-Mou Zheng, Shiwei Sun, Huaiyi Yang, Dongbo Bu |
Bioinform. | 4 |
| 2021 | FALCON2: a web server for high-quality prediction of protein tertiary structuresabstractBACKGROUND: Accurate prediction of protein tertiary structures is highly desired as the knowledge of protein structures provides invaluable insights into protein functions. We have designed two approaches to protein structure prediction, including a template-based modeling approach (called ProALIGN) and an ab initio prediction approach (called ProFOLD). Briefly speaking, ProALIGN aligns a target protein with templates through exploiting the patterns of context-specific alignment motifs and then builds the final structure with reference to the homologous templates. In contrast, ProFOLD uses an end-to-end neural network to estimate inter-residue distances of target proteins and builds structures that satisfy these distance constraints. These two approaches emphasize different characteristics of target proteins: ProALIGN exploits structure information of homologous templates of target proteins while ProFOLD exploits the co-evolutionary information carried by homologous protein sequences. Recent progress has shown that the combination of template-based modeling and ab initio approaches is promising. RESULTS: In the study, we present FALCON2, a web server that integrates ProALIGN and ProFOLD to provide high-quality protein structure prediction service. For a target protein, FALCON2 executes ProALIGN and ProFOLD simultaneously to predict possible structures and selects the most likely one as the final prediction result. We evaluated FALCON2 on widely-used benchmarks, including 104 CASP13 (the 13th Critical Assessment of protein Structure Prediction) targets and 91 CASP14 targets. In-depth examination suggests that when high-quality templates are available, ProALIGN is superior to ProFOLD and in other cases, ProFOLD shows better performance. By integrating these two approaches with different emphasis, FALCON2 server outperforms the two individual approaches and also achieves state-of-the-art performance compared with existing approaches. CONCLUSIONS: By integrating template-based modeling and ab initio approaches, FALCON2 provides an easy-to-use and high-quality protein structure prediction service for the community and we expect it to enable insights into a deep understanding of protein functions. Lupeng Kong, Fusong Ju, Haicang Zhang, Shiwei Sun, Dongbo Bu |
BMC Bioinform. | 3 |
| 2019 | Constructing effective energy functions for protein structure prediction through broadening attraction-basin and reverse Monte Carlo samplingabstractBACKGROUND: The ab initio approaches to protein structure prediction usually employ the Monte Carlo technique to search the structural conformation that has the lowest energy. However, the widely-used energy functions are usually ineffective for conformation search. How to construct an effective energy function remains a challenging task. RESULTS: Here, we present a framework to construct effective energy functions for protein structure prediction. Unlike existing energy functions only requiring the native structure to be the lowest one, we attempt to maximize the attraction-basin where the native structure lies in the energy landscape. The underlying rationale is that each energy function determines a specific energy landscape together with a native attraction-basin, and the larger the attraction-basin is, the more likely for the Monte Carlo search procedure to find the native structure. Following this rationale, we constructed effective energy functions as follows: i) To explore the native attraction-basin determined by a certain energy function, we performed reverse Monte Carlo sampling starting from the native structure, identifying the structural conformations on the edge of attraction-basin. ii) To broaden the native attraction-basin, we smoothened the edge points of attraction-basin through tuning weights of energy terms, thus acquiring an improved energy function. Our framework alternates the broadening attraction-basin and reverse sampling steps (thus called BARS) until the native attraction-basin is sufficiently large. We present extensive experimental results to show that using the BARS framework, the constructed energy functions could greatly facilitate protein structure prediction in improving the quality of predicted structures and speeding up conformation search. CONCLUSION: Using the BARS framework, we constructed effective energy functions for protein structure prediction, which could improve the quality of predicted structures and speed up conformation search as well. Haicang Zhang, Lupeng Kong, Shiwei Sun, Wei-Mou Zheng, Dongbo Bu |
BMC Bioinform. | 3 |
| 2019 | Predicting protein inter-residue contacts using composite likelihood maximization and deep learningabstractBACKGROUND: Accurate prediction of inter-residue contacts of a protein is important to calculating its tertiary structure. Analysis of co-evolutionary events among residues has been proved effective in inferring inter-residue contacts. The Markov random field (MRF) technique, although being widely used for contact prediction, suffers from the following dilemma: the actual likelihood function of MRF is accurate but time-consuming to calculate; in contrast, approximations to the actual likelihood, say pseudo-likelihood, are efficient to calculate but inaccurate. Thus, how to achieve both accuracy and efficiency simultaneously remains a challenge. RESULTS: In this study, we present such an approach (called clmDCA) for contact prediction. Unlike plmDCA using pseudo-likelihood, i.e., the product of conditional probability of individual residues, our approach uses composite-likelihood, i.e., the product of conditional probability of all residue pairs. Composite likelihood has been theoretically proved as a better approximation to the actual likelihood function than pseudo-likelihood. Meanwhile, composite likelihood is still efficient to maximize, thus ensuring the efficiency of clmDCA. We present comprehensive experiments on popular benchmark datasets, including PSICOV dataset and CASP-11 dataset, to show that: i) clmDCA alone outperforms the existing MRF-based approaches in prediction accuracy. ii) When equipped with deep learning technique for refinement, the prediction accuracy of clmDCA was further significantly improved, suggesting the suitability of clmDCA for subsequent refinement procedure. We further present a successful application of the predicted contacts to accurately build tertiary structures for proteins in the PSICOV dataset. CONCLUSIONS: Composite likelihood maximization algorithm can efficiently estimate the parameters of Markov Random Fields and can improve the prediction accuracy of protein inter-residue contacts. Haicang Zhang, Fusong Ju, Jianwei Zhu, Yujuan Gao, Ziwei Xie, Minghua Deng, Shiwei Sun, Wei-Mou Zheng, Dongbo Bu |
BMC Bioinform. | 1 |
| 2019 | Correction to: Predicting protein inter-residue contacts using composite likelihood maximization and deep learningabstractFollowing publication of the original article [1], the author explained that there are several errors in the original article. Haicang Zhang, Fusong Ju, Jianwei Zhu, Yujuan Gao, Ziwei Xie, Minghua Deng, Shiwei Sun, Wei-Mou Zheng, Dongbo Bu |
BMC Bioinform. | 1 |
| 2017 | Improving protein fold recognition by extracting fold-specific features from predicted residue-residue contactsabstractMOTIVATION: Accurate recognition of protein fold types is a key step for template-based prediction of protein structures. The existing approaches to fold recognition mainly exploit the features derived from alignments of query protein against templates. These approaches have been shown to be successful for fold recognition at family level, but usually failed at superfamily/fold levels. To overcome this limitation, one of the key points is to explore more structurally informative features of proteins. Although residue-residue contacts carry abundant structural information, how to thoroughly exploit these information for fold recognition still remains a challenge. RESULTS: In this study, we present an approach (called DeepFR) to improve fold recognition at superfamily/fold levels. The basic idea of our approach is to extract fold-specific features from predicted residue-residue contacts of proteins using deep convolutional neural network (DCNN) technique. Based on these fold-specific features, we calculated similarity between query protein and templates, and then assigned query protein with fold type of the most similar template. DCNN has showed excellent performance in image feature extraction and image recognition; the rational underlying the application of DCNN for fold recognition is that contact likelihood maps are essentially analogy to images, as they both display compositional hierarchy. Experimental results on the LINDAHL dataset suggest that even using the extracted fold-specific features alone, our approach achieved success rate comparable to the state-of-the-art approaches. When further combining these features with traditional alignment-related features, the success rate of our approach increased to 92.3%, 82.5% and 78.8% at family, superfamily and fold levels, respectively, which is about 18% higher than the state-of-the-art approach at fold level, 6% higher at superfamily level and 1% higher at family level. An independent assessment on SCOP_TEST dataset showed consistent performance improvement, indicating robustness of our approach. Furthermore, bi-clustering results of the extracted features are compatible with fold hierarchy of proteins, implying that these features are fold-specific. Together, these results suggest that the features extracted from predicted contacts are orthogonal to alignment-related features, and the combination of them could greatly facilitate fold recognition at superfamily/fold levels and template-based prediction of protein structures. AVAILABILITY AND IMPLEMENTATION: Source code of DeepFR is freely available through https://github.com/zhujianwei31415/deepfr, and a web server is available through http://protein.ict.ac.cn/deepfr. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jianwei Zhu, Haicang Zhang, Shuaicheng Li 0001, Lupeng Kong, Shiwei Sun, Wei-Mou Zheng, Dongbo Bu |
Bioinform. | 2 |
| 2017 | Improving prediction of burial state of residues by exploiting correlation among residuesabstractBACKGROUND: Residues in a protein might be buried inside or exposed to the solvent surrounding the protein. The buried residues usually form hydrophobic cores to maintain the structural integrity of proteins while the exposed residues are tightly related to protein functions. Thus, the accurate prediction of solvent accessibility of residues will greatly facilitate our understanding of both structure and functionalities of proteins. Most of the state-of-the-art prediction approaches consider the burial state of each residue independently, thus neglecting the correlations among residues. RESULTS: In this study, we present a high-order conditional random field model that considers burial states of all residues in a protein simultaneously. Our approach exploits not only the correlation among adjacent residues but also the correlation among long-range residues. Experimental results showed that by exploiting the correlation among residues, our approach outperformed the state-of-the-art approaches in prediction accuracy. In-depth case studies also showed that by using the high-order statistical model, the errors committed by the bidirectional recurrent neural network and chain conditional random field models were successfully corrected. CONCLUSIONS: Our methods enable the accurate prediction of residue burial states, which should greatly facilitate protein structure prediction and evaluation. Hai'e Gong, Haicang Zhang, Jianwei Zhu, Shiwei Sun, Wei-Mou Zheng, Dongbo Bu |
BMC Bioinform. | 2 |
| 2016 | FALCON@home: a high-throughput protein structure prediction server based on remote homologue recognitionabstractSUMMARY: The protein structure prediction approaches can be categorized into template-based modeling (including homology modeling and threading) and free modeling. However, the existing threading tools perform poorly on remote homologous proteins. Thus, improving fold recognition for remote homologous proteins remains a challenge. Besides, the proteome-wide structure prediction poses another challenge of increasing prediction throughput. In this study, we presented FALCON@home as a protein structure prediction server focusing on remote homologue identification. The design of FALCON@home is based on the observation that a structural template, especially for remote homologous proteins, consists of conserved regions interweaved with highly variable regions. The highly variable regions lead to vague alignments in threading approaches. Thus, FALCON@home first extracts conserved regions from each template and then aligns a query protein with conserved regions only rather than the full-length template directly. This helps avoid the vague alignments rooted in highly variable regions, improving remote homologue identification. We implemented FALCON@home using the Berkeley Open Infrastructure of Network Computing (BOINC) volunteer computing protocol. With computation power donated from over 20,000 volunteer CPUs, FALCON@home shows a throughput as high as processing of over 1000 proteins per day. In the Critical Assessment of protein Structure Prediction (CASP11), the FALCON@home-based prediction was ranked the 12th in the template-based modeling category. As an application, the structures of 880 mouse mitochondria proteins were predicted, which revealed the significant correlation between protein half-lives and protein structural factors. AVAILABILITY AND IMPLEMENTATION: FALCON@home is freely available at http://protein.ict.ac.cn/FALCON/. CONTACT: [email protected], [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Haicang Zhang, Wei-Mou Zheng, Dong Xu 0002, Jianwei Zhu, Kang Ning 0001, Shiwei Sun, Shuaicheng Li 0001, Dongbo Bu |
Bioinform. | 2 |
| 2014 | On the inapproximability of minimizing cascading failures under the deterministic threshold model
Jingjie Liu, Haicang Zhang, Zhiwei Xu 0002 |
Inf. Process. Lett. | 3 |