VLDB 2026 Research / reviewers in the wild / expert
Jun Zhang 0078
dblp:29/4190-78
· DBLP profile ↗
16ranked-venue papers
8as first author
16since 2021 · last 2026
0000-0003-0484-0966ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 7 first-author · 15 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PepCCD: A Contrastive Conditioned Diffusion Framework for Target-Specific Peptide GenerationabstractPeptide-based drug design targeting “undruggable” proteins remains one of the most critical challenges in modern drug discovery. Conventional peptide-discovery pipelines rely on low-throughput experimental screening, which is both time-consuming and prohibitively expensive. Moreover, existing computational approaches for designing peptides against target proteins typically depend on the availability of high-quality structural information. Although recent structure-prediction tools such as AlphaFold3 have achieved breakthroughs in protein modeling, their accuracy for functional interfaces remains limited. The acquisition of high-resolution structures is often expensive, time-intensive, and particularly challenging for targets with dynamic conformations, further restricting the efficient development of peptide therapeutics. Additionally, current sequence-based generative methods follow a paradigm that relies on known templates, which limits the exploration of sequence space and results in generated peptides lacking diversity and novelty. To address these limitations, we propose a contrastive conditioned diffusion framework for target-specific peptide generation, referred to as PepCCD. It employs a contrastive learning strategy between proteins and peptides to extract sequence-based conditioning representations of target proteins, which serve as precise conditions to guide a pre-trained diffusion model to generate peptide sequences with the desired target specificity. Extensive experiments on multiple benchmark target proteins demonstrate that the peptides designed by PepCCD exhibit strong binding affinity and outperform state-of-the-art methods in terms of diversity and generation efficiency. Jun Zhang 0078, Yangyang Zhou, Zexuan Zhu 0001 |
AAAI | 1 |
| 2026 | Molecular-level protein semantic learning via structure-aware coarse-grained language modelingabstractMOTIVATION: Protein language models (PLMs) have emerged as pivotal tools for protein representation, enabling significant advances in structure-function prediction and computational biology. However, current PLMs predominantly rely on fine-grained amino acid sequences as input, treating individual residues as tokens. While this approach facilitates semantic learning at the residue level, it struggles to capture molecular-level semantics, particularly for large proteins, where sequence truncation and inefficient local pattern extraction hinder holistic understanding. The spatial structure of a protein determines its function. Despite the critical role of protein function analysis, coarse-grained protein language frameworks that bridge sequence and structural semantics remain underdeveloped. RESULTS: To fill this gap, we introduce a novel structure-aware coarse-grained protein language that discretizes proteins into local structural patterns derived from their secondary structures. By constructing a vocabulary of these patterns as "words," we represent proteins as compact, structure-aware "sentences" significantly shorter than raw amino acid sequences. We benchmark the proposed coarse-grained language against three state-of-the-art fine-grained protein languages and a classical language modeling method in natural language processing, using two architectures: a lightweight Doc2Vec model and a Transformer-based BERT model, and evaluating performance across diverse downstream tasks, including function prediction, enzyme classification, and interaction identification. The proposed method achieves stable performance across three tasks, especially for long proteins. These results demonstrate that the proposed coarse-grained protein language preserves critical structural and functional semantics and improves molecular-level analysis, offering a promising direction for decoding higher-order biological insights. AVAILABILITY AND IMPLEMENTATION: The data and source code of the proposed method are available at GitHub (https://github.com/bug-0x3f/coarse-grained-protein-language) and Zenodo (DOI: 10.5281/zenodo.17674298). Jun Zhang 0078, Xueer Weng, Zexuan Zhu 0001 |
Bioinform. | 1 |
| 2025 | CFPLM: Improve Protein-RNA Interaction Prediction with a Collaborative Framework Powered by Language ModelsabstractProtein-RNA interactions (PRIs) play pivotal roles in biological processes such as gene regulation, making their prediction essential for therapeutic and mechanistic studies. While traditional wet-lab methods are timeconsuming and challenging, computational approaches offer efficient alternatives. Graph-based methods show promise by capturing both direct interactive domains (protein-RNA interaction) and indirect collaborative domains (functional similarity among proteins/RNAs). However, integrating these domains and learning meaningful node representations remain critical challenges. To address this, we propose CFPLM, a collaborative framework fusing large language models, graph convolutional networks, and cross-attention mechanisms to improve PRI prediction. Experiment results demonstrate that CFPLM achieves robust, state-of-the-art performance across three benchmark datasets. It's anticipated to have applicability to similar other interaction prediction tasks. The data and codes are available at: https://github.com/HuanchaoFeng/CFPLM. Jun Zhang 0078, Huanchao Feng, Hang Wei 0005, Zexuan Zhu 0001 |
BIBM | 1 |
| 2025 | TG-CDDPM: text-guided antimicrobial peptides generation based on conditional denoising diffusion probabilistic modelabstractAntimicrobial peptides (AMPs) have emerged as a promising substitution to antibiotics thanks to their boarder range of activities, less likelihood of drug resistance, and low toxicity. Traditional biochemical methods for AMP discovery are costly and inefficient. Deep generative models, including the long-short term memory model, variational autoencoder model, and generative adversarial model, have been widely introduced to expedite AMP discovery. However, these models tend to suffer from the lack of diversity in generating AMPs. The denoising diffusion probabilistic model serves as a good candidate for solving this issue. We proposed a three-stage Text-Guided Conditional Denoising Diffusion Probabilistic Model (TG-CDDPM) to generate novel and homologous AMPs. In the first two stages, contrastive learning and inferring models are crafted to create better conditions for guiding AMP generation, respectively. In the last stage, a pre-trained conditional denoising diffusion probabilistic model is leveraged to enrich the peptide knowledge and fine-tuned to learn feature representation in downstream. TG-CDDPM was compared to the state-of-the-art generative models for AMP generation, and it demonstrated competitive or better performance with the assistance of text description as supervised information. The membrane penetration capabilities of the identified candidate AMPs by TG-CDDPM were also validated through molecular weight dynamics experiments. Junhang Cao, Jun Zhang 0078, Qiyuan Yu, Junkai Ji, Jianqiang Li 0001, Shan He 0001, Zexuan Zhu 0001 |
Briefings Bioinform. | 2 |
| 2025 | INAB: identify nucleic acid binding domain via cross-modal protein language models and multiscale computationabstractProtein-nucleic acid interactions play a crucial role in biological processes, including gene regulation and editing. Accurately identifying nucleic acid-binding domains in proteins is essential to unravel these interactions, yet traditional experimental methods like X-ray crystallography remain costly and time-intensive. Computational approaches have thus emerged as indispensable tools to complement wet-lab techniques. Here, we introduce a framework for nucleic acid-binding domain prediction by integrating cross-modal protein language models with a multiscale computational architecture. The proposed method leverages a structurally annotated benchmark dataset, which quantifies binding likelihood through hierarchical, proximity-based labels derived from experimental complexes. Evaluations demonstrate that the approach achieves state-of-the-art performance, providing a new insight into the design of multimodal learning systems in protein-nucleic acid interaction analysis and an open resource to accelerate discoveries in functional genomics and drug design. Jun Zhang 0078, Junjie Chen 0004, Zexuan Zhu 0001 |
Briefings Bioinform. | 1 |
| 2025 | HiPHD: Hierarchical Classification for Protein Remote Homology Detection by Incorporating Protein Sequential and Structural InformationabstractProtein remote homology detection is crucial in various biological tasks, such as protein function annotation and structure prediction. Computational methods have been developed to improve efficiency and accuracy of protein homology detection. However, proteins with remote homology usually share similar structures and low sequence identity, resulting in limited performance of sequence/structure alignment-based methods. This study introduces HiPHD, a hierarchical classification framework for protein remote homology detection. HiPHD integrates protein sequential information embedded by protein language models and structural information encoded by graph neural networks, effectively combining spatial and sequential features. Experimental results demonstrate that HiPHD outperforms existing methods in terms of accuracy at all hierarchical levels in both SCOPe and CATH databases. It's anticipated HiPHD will become a valuable tool for protein homology detection and representation learning. Fuchuan Qu, Yijin Zhao, Jun Zhang 0078, Junjie Chen 0004 |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2025 | MDTL-ACP: Anticancer Peptides Prediction Based on Multi-Domain Transfer LearningabstractAnticancer peptides (ACPs) have emerged as one of the most promising therapeutic agents for cancer treatment. They are bioactive peptides featuring broad-spectrum activity and low drug-resistance. The discovery of ACPs via traditional biochemical methods is laborious and costly. Accordingly, various computational methods have been developed to facilitate the discovery of ACPs. However, the data resources and knowledge of ACPs are still very scarce, and only a few of them are clinically verified, which limits the competence of computational methods. To address this issue, in this article, we propose an ACP prediction model based on multi-domain transfer learning, namely MDTL-ACP, to discriminate novel ACPs from plentiful inactive peptides. In particular, we collect abundant antimicrobial peptides (AMPs) from four well-studied peptide domains and extract their inherent features as the input of MDTL-ACP. The features learned from multiple source domains of AMPs are then transferred into the target prediction task of ACPs via artificial neural network-based shared-extractor and task-specific classifiers in MDTL-ACP. The knowledge captured in the transferred features enhances the prediction of ACPs in the target domain. Experimental results demonstrate that MDTL-ACP can outperform the traditional and state-of-the-art ACP prediction methods. Junhang Cao, Wei Zhou 0001, Qiyuan Yu, Junkai Ji, Jun Zhang 0078, Shan He 0001, Zexuan Zhu 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | PAIR: protein-aptamer interaction prediction based on language models and contrastive learning frameworkabstractAptamers are single-stranded DNA or RNA oligonucleotides that selectively bind to specific targets, making them valuable for drug design and diagnostic applications. Identifying the interactions between aptamers and target proteins is crucial for these applications. The systematic evolution of ligands by exponential enrichment process, traditionally used for this purpose, is challenging and time-consuming. The resulting aptamers often suffer from limitations in stability and diversity. Computational approaches have shown promise in aiding the discovery of high-performance aptamers, but existing methods are usually constrained by insufficient training data and limited generalizability. Recently, advancements in pre-training large language models have offered a new avenue to mitigate the dependency on large datasets. In this study, we propose a novel method to predict aptamer-protein interactions using large language models within a contrastive learning framework. Experimental results demonstrate that our method exhibits superior generalization and outperforms existing approaches. This method holds promise as a powerful tool for predicting aptamer-protein interactions. Jun Zhang 0078, Zexuan Zhu 0001 |
BIBM | 1 |
| 2023 | MIX-TPI: a flexible prediction framework for TCR-pMHC interactions based on multimodal representationsabstractMOTIVATION: The interactions between T-cell receptors (TCR) and peptide-major histocompatibility complex (pMHC) are essential for the adaptive immune system. However, identifying these interactions can be challenging due to the limited availability of experimental data, sequence data heterogeneity, and high experimental validation costs. RESULTS: To address this issue, we develop a novel computational framework, named MIX-TPI, to predict TCR-pMHC interactions using amino acid sequences and physicochemical properties. Based on convolutional neural networks, MIX-TPI incorporates sequence-based and physicochemical-based extractors to refine the representations of TCR-pMHC interactions. Each modality is projected into modality-invariant and modality-specific representations to capture the uniformity and diversities between different features. A self-attention fusion layer is then adopted to form the classification module. Experimental results demonstrate the effectiveness of MIX-TPI in comparison with other state-of-the-art methods. MIX-TPI also shows good generalization capability on mutual exclusive evaluation datasets and a paired TCR dataset. AVAILABILITY AND IMPLEMENTATION: The source code of MIX-TPI and the test data are available at: https://github.com/Wolverinerine/MIX-TPI. Zhi-an Huang, Wei Zhou 0001, Junkai Ji, Jun Zhang 0078, Shan He 0001, Zexuan Zhu 0001 |
Bioinform. | 5 |
| 2023 | iDRBP-EL: Identifying DNA- and RNA- Binding Proteins Based on Hierarchical Ensemble LearningabstractIdentification of DNA-binding proteins (DBPs) and RNA-binding proteins (RBPs) from the primary sequences is essential for further exploring protein-nucleic acid interactions. Previous studies have shown that machine-learning-based methods can efficiently identify DBPs or RBPs. However, the information used in these methods is slightly unitary, and most of them only can predict DBPs or RBPs. In this study, we proposed a computational predictor iDRBP-EL to identify DNA- and RNA- binding proteins, and introduced hierarchical ensemble learning to integrate three level information. The method can integrate the information of different features, machine learning algorithms and data into one multi-label model. The ablation experiment showed that the fusion of different information can improve the prediction performance and overcome the cross-prediction problem. Experimental results on the independent datasets showed that iDRBP-EL outperformed all the other competing methods. Moreover, we established a user-friendly webserver iDRBP-EL (http://bliulab.net/iDRBP-EL), which can predict both DBPs and RBPs only based on protein sequences. Ning Wang 0054, Jun Zhang 0078, Bin Liu 0014 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | iDRNA-ITF: identifying DNA- and RNA-binding residues in proteins based on induction and transfer frameworkabstractProtein-DNA and protein-RNA interactions are involved in many biological activities. In the post-genome era, accurate identification of DNA- and RNA-binding residues in protein sequences is of great significance for studying protein functions and promoting new drug design and development. Therefore, some sequence-based computational methods have been proposed for identifying DNA- and RNA-binding residues. However, they failed to fully utilize the functional properties of residues, leading to limited prediction performance. In this paper, a sequence-based method iDRNA-ITF was proposed to incorporate the functional properties in residue representation by using an induction and transfer framework. The properties of nucleic acid-binding residues were induced by the nucleic acid-binding residue feature extraction network, and then transferred into the feature integration modules of the DNA-binding residue prediction network and the RNA-binding residue prediction network for the final prediction. Experimental results on four test sets demonstrate that iDRNA-ITF achieves the state-of-the-art performance, outperforming the other existing sequence-based methods. The webserver of iDRNA-ITF is freely available at http://bliulab.net/iDRNA-ITF. Ning Wang 0054, Ke Yan 0003, Jun Zhang 0078, Bin Liu 0014 |
Briefings Bioinform. | 3 |
| 2022 | PreRBP-TL: prediction of species-specific RNA-binding proteins based on transfer learningabstractMOTIVATION: RNA-binding proteins (RBPs) play crucial roles in post-transcriptional regulation. Accurate identification of RBPs helps to understand gene expression, regulation, etc. In recent years, some computational methods were proposed to identify RBPs. However, these methods fail to accurately identify RBPs from some specific species with limited data, such as bacteria. RESULTS: In this study, we introduce a computational method called PreRBP-TL for identifying species-specific RBPs based on transfer learning. The weights of the prediction model were initialized by pretraining with the large general RBP dataset and then fine-tuned with the small species-specific RPB dataset by using transfer learning. The experimental results show that the PreRBP-TL achieves better performance for identifying the species-specific RBPs from Human, Arabidopsis, Escherichia coli and Salmonella, outperforming eight state-of-the-art computational methods. It is anticipated PreRBP-TL will become a useful method for identifying RBPs. AVAILABILITY AND IMPLEMENTATION: For the convenience of researchers to identify RBPs, the web server of PreRBP-TL was established, freely available at http://bliulab.net/PreRBP-TL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jun Zhang 0078, Ke Yan 0003, Qingcai Chen, Bin Liu 0014 |
Bioinform. | 1 |
| 2022 | IDRBP-PPCT: Identifying Nucleic Acid-Binding Proteins Based on Position-Specific Score Matrix and Position-Specific Frequency Matrix Cross TransformationabstractDNA-binding proteins (DBPs) and RNA-binding proteins (RBPs) are two important nucleic acid-binding proteins (NABPs), which play important roles in biological processes such as replication, translation and transcription of genetic material. Some proteins (DRBPs) bind to both DNA and RNA, also play a key role in gene expression. Identification of DBPs, RBPs and DRBPs is important to study protein-nucleic acid interactions. Computational methods are increasingly being proposed to automatically identify DNA- or RNA-binding proteins based only on protein sequences. One challenge is to design an effective protein representation method to convert protein sequences into fixed-dimension feature vectors. In this study, we proposed a novel protein representation method called Position-Specific Scoring Matrix (PSSM) and Position-Specific Frequency Matrix (PSFM) Cross Transformation (PPCT) to represent protein sequences. This method contains the evolutionary information in PSSM and PSFM, and their correlations. A new computational predictor called IDRBP-PPCT was proposed by combining PPCT and the two-layer framework based on the random forest algorithm to identify DBPs, RBPs and DRBPs. The experimental results on the independent dataset and the tomato genome proved the effectiveness of the proposed method. A user-friendly web-server of IDRBP-PPCT was constructed, which is freely available at http://bliulab.net/IDRBP-PPCT. Ning Wang 0054, Jun Zhang 0078, Bin Liu 0014 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2021 | NCBRPred: predicting nucleic acid binding residues in proteins based on multilabel learningabstractThe interactions between proteins and nucleic acid sequences play many important roles in gene expression and some cellular activities. Accurate prediction of the nucleic acid binding residues in proteins will facilitate the research of the protein functions, gene expression, drug design, etc. In this regard, several computational methods have been proposed to predict the nucleic acid binding residues in proteins. However, these methods cannot satisfactorily measure the global interactions among the residues along protein. Furthermore, these methods are suffering cross-prediction problem, new strategies should be explored to solve this problem. In this study, a new computational method called NCBRPred was proposed to predict the nucleic acid binding residues based on the multilabel sequence labeling model. NCBRPred used the bidirectional Gated Recurrent Units (BiGRUs) to capture the global interactions among the residues, and treats this task as a multilabel learning task. Experimental results on three widely used benchmark datasets and an independent dataset showed that NCBRPred achieved higher predictive results with lower cross-prediction, outperforming 10 existing state-of-the-art predictors. The web-server and a stand-alone package of NCBRPred are freely available at http://bliulab.net/NCBRPred. It is anticipated that NCBRPred will become a very useful tool for identifying nucleic acid binding residues. Jun Zhang 0078, Qingcai Chen, Bin Liu 0014 |
Briefings Bioinform. | 1 |
| 2021 | SMI-BLAST: a novel supervised search framework based on PSI-BLAST for protein remote homology detectionabstractMOTIVATION: As one of the most important and widely used mainstream iterative search tool for protein sequence search, an accurate Position-Specific Scoring Matrix (PSSM) is the key of PSI-BLAST. However, PSSMs containing non-homologous information obviously reduce the performance of PSI-BLAST for protein remote homology. RESULTS: To further study this problem, we summarize three types of Incorrectly Selected Homology (ISH) errors in PSSMs. A new search tool Supervised-Manner-based Iterative BLAST (SMI-BLAST) is proposed based on PSI-BLAST for solving these errors. SMI-BLAST obviously outperforms PSI-BLAST on the Structural Classification of Proteins-extended (SCOPe) dataset. Compared with PSI-BLAST on the ISH error subsets of SCOPe dataset, SMI-BLAST detects 1.6-2.87 folds more remote homologous sequences, and outperforms PSI-BLAST by 35.66% in terms of ROC1 scores. Furthermore, this framework is applied to JackHMMER, DELTA-BLAST and PSI-BLASTexB, and their performance is further improved. AVAILABILITY AND IMPLEMENTATION: User-friendly webservers for SMI-BLAST, JackHMMER, DELTA-BLAST and PSI-BLASTexB are established at http://bliulab.net/SMI-BLAST/, by which the users can easily get the results without the need to go through the mathematical details. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaopeng Jin, Qing Liao 0001, Hang Wei 0005, Jun Zhang 0078, Bin Liu 0014 |
Bioinform. | 4 |
| 2021 | DeepDRBP-2L: A New Genome Annotation Predictor for Identifying DNA-Binding Proteins and RNA-Binding Proteins Using Convolutional Neural Network and Long Short-Term MemoryabstractDNA-binding proteins (DBPs) and RNA-binding proteins (RBPs) are two kinds of crucial proteins, which are associated with various cellule activities and some important diseases. Accurate identification of DBPs and RBPs facilitate both theoretical research and real world application. Existing sequence-based DBP predictors can accurately identify DBPs but incorrectly predict many RBPs as DBPs, and vice versa, resulting in low prediction precision. Moreover, some proteins (DRBPs) interacting with both DNA and RNA play important roles in gene expression and cannot be identified by existing computational methods. In this study, a two-level predictor named DeepDRBP-2L was proposed by combining Convolutional Neural Network (CNN) and the Long Short-Term Memory (LSTM). It is the first computational method that is able to identify DBPs, RBPs and DRBPs. Rigorous cross-validations and independent tests showed that DeepDRBP-2L is able to overcome the shortcoming of the existing methods and can go one further step to identify DRBPs. Application of DeepDRBP-2L to tomato genome further demonstrated its performance. The webserver of DeepDRBP-2L is freely available at http://bliulab.net/DeepDRBP-2L. Jun Zhang 0078, Qingcai Chen, Bin Liu 0014 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |