Hai-Feng Chen

dblp:48/8075 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0002-7496-4182ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021
YearPublicationVenuePosition
2025 ProDualNet: dual-target protein sequence design method based on protein language model and structure model
abstract
Proteins typically interact with multiple partners to regulate biological processes, and peptide drugs targeting multiple receptors have shown strong therapeutic potential, emphasizing the need for multi-target strategies in protein design. However, most current protein sequence design methods focus on interactions with a single receptor, often neglecting the complexity of designing proteins that can bind to two distinct receptors. We introduced Protein Dual-Target Design Network (ProDualNet), a structure-based sequence design method that integrates sequence-structure information from two receptors to design dual-target protein sequences. ProDualNet used a heterogeneous graph network for pretraining and combines noise-augmented single-target data with real dual-target data for fine-tuning. This approach addressed the challenge of limited dual-target protein experimental structures. The efficacy of ProDualNet has been validated across multiple test sets, demonstrating better recovery and success rates compared to other multi-state design methods. In silico evaluation of cases like dual-target allosteric binding and non-overlapping interface binding highlights its potential for designing dual-target binding proteins. Data and code are available at https://github.com/chengliu97/ProDualNet.
Liu Cheng, Xiaochen Cui, Hai-Feng Chen, Zhangsheng Yu
Briefings Bioinform.4
2024 Graphormer supervised de novo protein design method and function validation
abstract
Protein design is central to nearly all protein engineering problems, as it can enable the creation of proteins with new biological functions, such as improving the catalytic efficiency of enzymes. One key facet of protein design, fixed-backbone protein sequence design, seeks to design new sequences that will conform to a prescribed protein backbone structure. Nonetheless, existing sequence design methods present limitations, such as low sequence diversity and shortcomings in experimental validation of the designed functional proteins. These inadequacies obstruct the goal of functional protein design. To improve these limitations, we initially developed the Graphormer-based Protein Design (GPD) model. This model utilizes the Transformer on a graph-based representation of three-dimensional protein structures and incorporates Gaussian noise and a sequence random masks to node features, thereby enhancing sequence recovery and diversity. The performance of the GPD model was significantly better than that of the state-of-the-art ProteinMPNN model on multiple independent tests, especially for sequence diversity. We employed GPD to design CalB hydrolase and generated nine artificially designed CalB proteins. The results show a 1.7-fold increase in catalytic activity compared to that of the wild-type CalB and strong substrate selectivity on p-nitrophenyl acetate with different carbon chain lengths (C2-C16). Thus, the GPD method could be used for the de novo design of industrial enzymes and protein drugs. The code was released at https://github.com/decodermu/GPD.
Junxi Mu, Jamshed Iqbal, Abdul Wadood, Hai-Feng Chen
Briefings Bioinform.9
2024 Phanto-IDP: compact model for precise intrinsically disordered protein backbone generation and enhanced sampling
abstract
The biological function of proteins is determined not only by their static structures but also by the dynamic properties of their conformational ensembles. Numerous high-accuracy static structure prediction tools have been recently developed based on deep learning; however, there remains a lack of efficient and accurate methods for exploring protein dynamic conformations. Traditionally, studies concerning protein dynamics have relied on molecular dynamics (MD) simulations, which incur significant computational costs for all-atom precision and struggle to adequately sample conformational spaces with high energy barriers. To overcome these limitations, various enhanced sampling techniques have been developed to accelerate sampling in MD. Traditional enhanced sampling approaches like replica exchange molecular dynamics (REMD) and frontier expansion sampling (FEXS) often follow the MD simulation approach and still cost a lot of computational resources and time. Variational autoencoders (VAEs), as a classic deep generative model, are not restricted by potential energy landscapes and can explore conformational spaces more efficiently than traditional methods. However, VAEs often face challenges in generating reasonable conformations for complex proteins, especially intrinsically disordered proteins (IDPs), which limits their application as an enhanced sampling method. In this study, we presented a novel deep learning model (named Phanto-IDP) that utilizes a graph-based encoder to extract protein features and a transformer-based decoder combined with variational sampling to generate highly accurate protein backbones. Ten IDPs and four structured proteins were used to evaluate the sampling ability of Phanto-IDP. The results demonstrate that Phanto-IDP has high fidelity and diversity in the generated conformation ensembles, making it a suitable tool for enhancing the efficiency of MD simulation, generating broader protein conformational space and a continuous protein transition path.
Haowei Tong, Zhouyu Lu, Ningjie Zhang, Hai-Feng Chen
Briefings Bioinform.7
2024 Multi-indicator comparative evaluation for deep learning-based protein sequence design methods
abstract
MOTIVATION: Proteins found in nature represent only a fraction of the vast space of possible proteins. Protein design presents an opportunity to explore and expand this protein landscape. Within protein design, protein sequence design plays a crucial role, and numerous successful methods have been developed. Notably, deep learning-based protein sequence design methods have experienced significant advancements in recent years. However, a comprehensive and systematic comparison and evaluation of these methods have been lacking, with indicators provided by different methods often inconsistent or lacking effectiveness. RESULTS: To address this gap, we have designed a diverse set of indicators that cover several important aspects, including sequence recovery, diversity, root-mean-square deviation of protein structure, secondary structure, and the distribution of polar and nonpolar amino acids. In our evaluation, we have employed an improved weighted inferiority-superiority distance method to comprehensively assess the performance of eight widely used deep learning-based protein sequence design methods. Our evaluation not only provides rankings of these methods but also offers optimization suggestions by analyzing the strengths and weaknesses of each method. Furthermore, we have developed a method to select the best temperature parameter and proposed solutions for the common issue of designing sequences with consecutive repetitive amino acids, which is often encountered in protein design methods. These findings can greatly assist users in selecting suitable protein sequence design methods. Overall, our work contributes to the field of protein sequence design by providing a comprehensive evaluation system and optimization suggestions for different methods.
Jinyu Yu, Junxi Mu, Hai-Feng Chen
Bioinform.4
2023 PTMint database of experimentally verified PTM regulation on protein-protein interaction
abstract
MOTIVATION: Post-translational modification (PTM) is an important biochemical process. which includes six most well-studied types: phosphorylation, acetylation, methylation, sumoylation, ubiquitylation and glycosylation. PTM is involved in various cell signaling pathways and biological processes. Abnormal PTM status is closely associated with severe diseases (such as cancer and neurologic diseases) by regulating protein functions, such as protein-protein interactions (PPIs). A set of databases was constructed separately for PTM sites and PPI; however, the resource of regulation for PTM on PPI is still unsolved. RESULTS: Here, we firstly constructed a public accessible database of PTMint (PTMs that are associated with PPIs) (https://ptmint.sjtu.edu.cn/) that contains manually curated complete experimental evidence of the PTM regulation on PPIs in multiple organisms, including Homo sapiens, Arabidopsis thaliana, Caenorhabditis elegans, Drosophila melanogaster, Saccharomyces cerevisiae and Schizosaccharomyces pombe. Currently, the first version of PTMint encompassed 2477 non-redundant PTM sites in 1169 proteins affecting 2371 protein-protein pairs involving 357 diseases. Various annotations were systematically integrated, such as protein sequence, structure properties and protein complex analysis. PTMint database can help to insight into disease mechanism, disease diagnosis and drug discovery associated with PTM and PPI. AVAILABILITY AND IMPLEMENTATION: PTMint is freely available at: https://ptmint.sjtu.edu.cn/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiaokun Hong, Ningshan Li, Jiyang Lv, Yan Zhang 0124, Jing Li 0107, Jian Zhang 0037, Hai-Feng Chen
Bioinform.7
2019 Dynamical important residue network (DIRN): network inference via conformational change
abstract
MOTIVATION: Protein residue interaction network has emerged as a useful strategy to understand the complex relationship between protein structures and functions and how functions are regulated. In a residue interaction network, every residue is used to define a network node, adding noises in network post-analysis and increasing computational burden. In addition, dynamical information is often necessary in deciphering biological functions. RESULTS: We developed a robust and efficient protein residue interaction network method, termed dynamical important residue network, by combining both structural and dynamical information. A major departure from previous approaches is our attempt to identify important residues most important for functional regulation before a network is constructed, leading to a much simpler network with the important residues as its nodes. The important residues are identified by monitoring structural data from ensemble molecular dynamics simulations of proteins in different functional states. Our tests show that the new method performs well with overall higher sensitivity than existing approaches in identifying important residues and interactions in tested proteins, so it can be used in studies of protein functions to provide useful hypotheses in identifying key residues and interactions. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ray Luo 0001, Hai-Feng Chen
Bioinform.3