VLDB 2026 Research / reviewers in the wild / expert
Jian Zhang 0020
dblp:07/314-20
· DBLP profile ↗
16ranked-venue papers
11as first author
7since 2021 · last 2025
0000-0001-7155-7760ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 10 first-author · 4 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Accurate Residue-Level Prediction of Linear Interacting Peptides with Auxiliary Multi-Scale Anchor DetectionabstractIntrinsically disordered regions (IDRs) play essential roles in cellular signaling and regulation, primarily mediating interactions through short peptide motifs. Linear Interacting Peptides (LIPs) are a recently defined class of bindingassociated IDRs that undergo disorder-to-order transitions upon binding. LIPs encompass several well-characterized subclasses, including molecular recognition features (MoRFs) and short linear motifs (SLiMs). However, many current predictors focus exclusively on MoRFs prediction and exhibit limited effectiveness in identifying the broader class of LIPs. To address this limitation, we propose LipPredictor, the first deep learning framework specifically designed for LIP prediction. The model incorporates protein language model embeddings as input to a shared feature extraction module, which is composed of convolutional neural networks and a multi-head attention layer. It further adopts a dual-branch architecture that couples residue-level classification with auxiliary multi-scale anchor detection, enabling accurate identification of LIPs across diverse segment lengths through joint training. These results demonstrate its robustness and effectiveness in predicting LIPs. The source code can be obtained from https://github.com/Chenxi-Xia/LipPredictor. Fuhao Zhang, Chenxi Xia, Mingxin Dong, Min Zeng 0004, Jian Zhang 0020, Min Li 0007 |
BIBM | 5 |
| 2025 | ANGraph: A GNN-Based Performance Prediction Framework for Asynchronous Neuromorphic HardwareabstractDesign space exploration (DSE) through system-level simulation is essential for designing energy-efficient asynchronous neuromorphic hardware, which is increasingly promising in edge AI applications. However, there are significant mismatches between system-level predictions and gate-level simulations, resulting in low precision when predicting performance during the DSE process for asynchronous neuromorphic hardware. To address this issue, we put forward ANGraph, a graph neural network (GNN)-based performance prediction framework for asynchronous neuromorphic hardware. In the ANGraph framework, we transform the intermediate representation of systemlevel simulations into graphs, collect gate-level circuit simulation results to build benchmarks with over one million samples, and train a GNN model to predict hardware latency for asynchronous neuromorphic hardware. Additionally, we use a residual network (ResNet)-based method to predict the power consumption of asynchronous neuromorphic hardware. We evaluate these two models on additional datasets without extra training across different scales, process nodes, and traffic patterns of input data. Compared to the latency predictions from the state-of-the-art simulator, we improve the R -square score by 0.69 and reduce root mean square error (RMSE) by 76% on average across all datasets. We also achieve an R-square score of 0.98 and a mean absolute percentage error (MAPE) of 0.88% for the power consumption prediction task. The benchmarks and models are available at https://github.com/HuaGuaiGuai/ANGraph. Yuan Hua, Jian Zhang 0020, Hong Chen 0002 |
DAC | 2 |
| 2025 | EnrichRBP: an automated and interpretable computational platform for predicting and analysing RNA-binding protein eventsabstractMOTIVATION: Predicting RNA-binding proteins (RBPs) is central to understanding post-transcriptional regulatory mechanisms. Here, we introduce EnrichRBP, an automated and interpretable computational platform specifically designed for the comprehensive analysis of RBP interactions with RNA. RESULTS: EnrichRBP is a web service that enables researchers to develop original deep learning and machine learning architectures to explore the complex dynamics of RBPs. The platform supports 70 deep learning algorithms, covering feature representation, selection, model training, comparison, optimization, and evaluation, all integrated within an automated pipeline. EnrichRBP is adept at providing comprehensive visualizations, enhancing model interpretability, and facilitating the discovery of functionally significant sequence regions crucial for RBP interactions. In addition, EnrichRBP supports base-level functional annotation tasks, offering explanations and graphical visualizations that confirm the reliability of the predicted RNA-binding sites. Leveraging high-performance computing, EnrichRBP provides ultra-fast predictions ranging from seconds to hours, applicable to both pre-trained and custom model scenarios, thus proving its utility in real-world applications. Case studies highlight that EnrichRBP provides robust and interpretable predictions, demonstrating the power of deep learning in the functional analysis of RBP interactions. Finally, EnrichRBP aims to enhance the reproducibility of computational method analyses for RBP sequences, as well as reduce the programming and hardware requirements for biologists, thereby offering meaningful functional insights. AVAILABILITY AND IMPLEMENTATION: EnrichRBP is available at https://airbp.aibio-lab.com/. The source code is available at https://github.com/wangyb97/EnrichRBP, and detailed online documentation can be found at https://enrichrbp.readthedocs.io/en/latest/. Yujian Huang, Jian Zhang 0020, Ka-Chun Wong, Xiangtao Li |
Bioinform. | 6 |
| 2023 | ANAS: Asynchronous Neuromorphic Hardware Architecture Search Based on a System-Level SimulatorabstractEvent-driven asynchronous neuromorphic hardware is emerging for edge computing with high energy efficiency. In order to obtain the architecture with the best hardware performance, we need to search both the numerical and non-numerical design space of asynchronous neuromorphic hardware. However, it is challenging to find an optimal hardware architecture from the non-numerical design space. To address this problem, we propose an asynchronous neuromorphic hardware architecture search (ANAS) method, which uses an evolutionary algorithm to optimize both the numerical and non-numerical design space. Besides, we introduce a configurable asynchronous neuromorphic hardware simulator (CanMore) to offer system-level modeling and performance estimation. Experimental results show that ANAS rivals the best human-designed architecture by 7 × EDP reduction, and offers 2.3 × EDP reduction than the methods that only optimize numerical design space. Jian Zhang 0020, Dexuan Huo, Hong Chen 0002 |
DAC | 1 |
| 2023 | Temporal-Coded Spiking Neural Networks with Dynamic Firing Threshold: Learning with Event-Driven BackpropagationabstractSpiking Neural Networks (SNNs) offer a highly promising computing paradigm due to their biological plausibility, exceptional spatiotemporal information processing capability and low power consumption. As a temporal encoding scheme for SNNs, Time-To-First-Spike (TTFS) encodes information using the timing of a single spike, which allows spiking neurons to transmit information through sparse spike trains and results in lower power consumption and higher computational efficiency compared to traditional rate-based encoding counterparts. However, despite the advantages of the TTFS encoding scheme, the effective and efficient training of TTFS-based deep SNNs remains a significant and open research problem. In this work, we first examine the factors underlying the limitations of applying existing TTFS-based learning algorithms to deep SNNs. Specifically, we investigate issues related to over-sparsity of spikes and the complexity of finding the ‘causal set'. We then propose a simple yet efficient dynamic firing threshold (DFT) mechanism for spiking neurons to address these issues. Building upon the proposed DFT mechanism, we further introduce a novel direct training algorithm for TTFS-based deep SNNs, called DTA-TTFS. This method utilizes event-driven processing and spike timing to enable efficient learning of deep SNNs. The proposed training method was validated on the image classification task and experimental results clearly demonstrate that our proposed method achieves state-of-the-art accuracy in comparison to existing TTFS-based learning algorithms, while maintaining high levels of sparsity and energy efficiency on neuromorphic inference accelerator. Wenjie Wei, Malu Zhang, Hong Qu 0002, Ammar Belatreche, Jian Zhang 0020, Hong Chen 0002 |
ICCV | 5 |
| 2023 | SCAMPER: Accurate Type-Specific Prediction of Calcium-Binding Residues Using Sequence-Derived FeaturesabstractUnderstanding molecular mechanisms involved in calcium-protein interactions and modeling corresponding docking rely on the accurate identification of calcium-binding residues (CaBRs). The defects of experimentally annotating protein functions enhances the development of computational approaches that correctly identify calcium-binding interactions. Studies have reported that current methods severely cross-predict residues that interact with other types of molecules (e.g., nucleic acids, proteins, and small ligands) as CaBRs. In this study, a novel predictor named SCAMPER (Selective CAlciuM-binding PrEdictoR) is proposed for the accurate and specific prediction of CaBRs. SCAMPER is designed using newly compiled dataset with complete UniProt sequences and annotations, which include calcium-binding, nucleic acid-binding, protein-binding, and small ligand-binding residues. We use a novel designed two-layer scheme to perform predictions as well as penalize cross-predictions. Empirical tests on an independent test dataset reveals that the proposed method significantly outperforms state-of-the-art predictors. SCAMPER is proved to be capable of distinguishing CaBRs from different types of metal-ion binding residues. We further perform CaBRs predictions on the whole human proteome, and use the results to hypothesize calcium-binding proteins (CaBPs). The latest experimental verified CaBPs and GO analysis prove the accuracy of our predictions. We implement the proposed method and share the data at http://www.inforstation.com/webservers/SCAMPER/. Jian Zhang 0020, Xingchen Liang, Guifu Yang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2021 | DNAgenie: accurate prediction of DNA-type-specific binding residues in protein sequencesabstractEfforts to elucidate protein-DNA interactions at the molecular level rely in part on accurate predictions of DNA-binding residues in protein sequences. While there are over a dozen computational predictors of the DNA-binding residues, they are DNA-type agnostic and significantly cross-predict residues that interact with other ligands as DNA binding. We leverage a custom-designed machine learning architecture to introduce DNAgenie, first-of-its-kind predictor of residues that interact with A-DNA, B-DNA and single-stranded DNA. DNAgenie uses a comprehensive physiochemical profile extracted from an input protein sequence and implements a two-step refinement process to provide accurate predictions and to minimize the cross-predictions. Comparative tests on an independent test dataset demonstrate that DNAgenie outperforms the current methods that we adapt to predict residue-level interactions with the three DNA types. Further analysis finds that the use of the second (refinement) step leads to a substantial reduction in the cross predictions. Empirical tests show that DNAgenie's outputs that are converted to coarse-grained protein-level predictions compare favorably against recent tools that predict which DNA-binding proteins interact with double-stranded versus single-stranded DNAs. Moreover, predictions from the sequences of the whole human proteome reveal that the results produced by DNAgenie substantially overlap with the known DNA-binding proteins while also including promising leads for several hundred previously unknown putative DNA binders. These results suggest that DNAgenie is a valuable tool for the sequence-based characterization of protein functions. The DNAgenie's webserver is available at http://biomine.cs.vcu.edu/servers/DNAgenie/. Jian Zhang 0020, Sina Ghadermarzi, Akila Katuwawala, Lukasz A. Kurgan |
Briefings Bioinform. | 1 |
| 2020 | Corrigendum to: Comprehensive review and empirical analysis of hallmarks of DNA-, RNA- and protein-binding residues in protein chainsabstractBriefings in Bioinformatics, 2017. https://doi.org/10.1093/bib/bbx168 In the original version of this article, reference 60 incorrectly listed Ben-Tal N as the third author of the paper ‘Local geometry and evolutionary conservation of protein surfaces reveal the multiple recognition patches in protein-protein interactions’. This reference has now been corrected to the below: Laine E, Carbone A. Local geometry and evolutionary conservation of protein surfaces reveal the multiple recognition patches in protein-protein interactions. PLoS Comput Biol 2015;11(12):e1004580. The authors would like to apologize for this error. Jian Zhang 0020, Zhiqiang Ma 0003, Lukasz A. Kurgan |
Briefings Bioinform. | 1 |
| 2020 | Prediction of protein-binding residues: dichotomy of sequence-based methods developed using structured complexes versus disordered proteinsabstractMOTIVATION: There are over 30 sequence-based predictors of the protein-binding residues (PBRs). They use either structure-annotated or disorder-annotated training datasets, potentially creating a dichotomy where the structure-/disorder-specific models may not be able to cross-over to accurately predict the other type. Moreover, the structure-trained predictors were shown to substantially cross-predict PBRs among residues that interact with non-protein partners (nucleic acids and small ligands). We address these issues by performing first-of-its-kind comparative study of a representative collection of disorder- and structure-trained predictors using a comprehensive benchmark set with the structure- and disorder-derived annotations of PBRs (to analyze the cross-over) and the protein-, nucleic acid- and small ligand-binding proteins (to study the cross-predictions). RESULTS: Three predictors provide accurate results: SCRIBER, ANCHOR and disoRDPbind. Some of the structure-trained methods make accurate predictions on the structure-annotated proteins. Similarly, the disorder-trained predictors predict well on the disorder-annotated proteins. However, the considered predictors generally fail to cross-over, with the exception of SCRIBER. Our study also reveals that virtually all methods substantially cross-predict PBRs, except for SCRIBER for the structure-annotated proteins and disoRDPbind for the disorder-annotated proteins. We formulate a novel hybrid predictor, hybridPBRpred, that combines results produced by disoRDPbind and SCRIBER to accurately predict disorder- and structure-annotated PBRs. HybridPBRpred generates accurate results that cross-over structure- and disorder-annotated proteins and produces relatively low amount of cross-predictions, offering an accurate alternative to predict PBRs. AVAILABILITY AND IMPLEMENTATION: HybridPBRpred webserver, benchmark dataset and supplementary information are available at http://biomine.cs.vcu.edu/servers/hybridPBRpred/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jian Zhang 0020, Sina Ghadermarzi, Lukasz A. Kurgan |
Bioinform. | 1 |
| 2020 | PROBselect: accurate prediction of protein-binding residues from proteins sequences via dynamic predictor selectionabstractMOTIVATION: Knowledge of protein-binding residues (PBRs) improves our understanding of protein-protein interactions, contributes to the prediction of protein functions and facilitates protein-protein docking calculations. While many sequence-based predictors of PBRs were published, they offer modest levels of predictive performance and most of them cross-predict residues that interact with other partners. One unexplored option to improve the predictive quality is to design consensus predictors that combine results produced by multiple methods. RESULTS: We empirically investigate predictive performance of a representative set of nine predictors of PBRs. We report substantial differences in predictive quality when these methods are used to predict individual proteins, which contrast with the dataset-level benchmarks that are currently used to assess and compare these methods. Our analysis provides new insights for the cross-prediction concern, dissects complementarity between predictors and demonstrates that predictive performance of the top methods depends on unique characteristics of the input protein sequence. Using these insights, we developed PROBselect, first-of-its-kind consensus predictor of PBRs. Our design is based on the dynamic predictor selection at the protein level, where the selection relies on regression-based models that accurately estimate predictive performance of selected predictors directly from the sequence. Empirical assessment using a low-similarity test dataset shows that PROBselect provides significantly improved predictive quality when compared with the current predictors and conventional consensuses that combine residue-level predictions. Moreover, PROBselect informs the users about the expected predictive quality for the prediction generated from a given input protein. AVAILABILITY AND IMPLEMENTATION: PROBselect is available at http://bioinformatics.csu.edu.cn/PROBselect/home/index. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Fuhao Zhang, Jian Zhang 0020, Min Zeng 0004, Min Li 0007, Lukasz A. Kurgan |
Bioinform. | 3 |
| 2019 | Comprehensive review and empirical analysis of hallmarks of DNA-, RNA- and protein-binding residues in protein chainsabstractProteins interact with a variety of molecules including proteins and nucleic acids. We review a comprehensive collection of over 50 studies that analyze and/or predict these interactions. While majority of these studies address either solely protein-DNA or protein-RNA binding, only a few have a wider scope that covers both protein-protein and protein-nucleic acid binding. Our analysis reveals that binding residues are typically characterized with three hallmarks: relative solvent accessibility (RSA), evolutionary conservation and propensity of amino acids (AAs) for binding. Motivated by drawbacks of the prior studies, we perform a large-scale analysis to quantify and contrast the three hallmarks for residues that bind DNA-, RNA-, protein- and (for the first time) multi-ligand-binding residues that interact with DNA and proteins, and with RNA and proteins. Results generated on a well-annotated data set of over 23 000 proteins show that conservation of binding residues is higher for nucleic acid- than protein-binding residues. Multi-ligand-binding residues are more conserved and have higher RSA than single-ligand-binding residues. We empirically show that each hallmark discriminates between binding and nonbinding residues, even predicted RSA, and that combining them improves discriminatory power for each of the five types of interactions. Linear scoring functions that combine these hallmarks offer good predictive performance of residue-level propensity for binding and provide intuitive interpretation of predictions. Better understanding of these residue-level interactions will facilitate development of methods that accurately predict binding in the exponentially growing databases of protein sequences. Jian Zhang 0020, Zhiqiang Ma 0003, Lukasz A. Kurgan |
Briefings Bioinform. | 1 |
| 2019 | SCRIBER: accurate and partner type-specific prediction of protein-binding residues from proteins sequencesabstractMOTIVATION: Accurate predictions of protein-binding residues (PBRs) enhances understanding of molecular-level rules governing protein-protein interactions, helps protein-protein docking and facilitates annotation of protein functions. Recent studies show that current sequence-based predictors of PBRs severely cross-predict residues that interact with other types of protein partners (e.g. RNA and DNA) as PBRs. Moreover, these methods are relatively slow, prohibiting genome-scale use. RESULTS: We propose a novel, accurate and fast sequence-based predictor of PBRs that minimizes the cross-predictions. Our SCRIBER (SeleCtive pRoteIn-Binding rEsidue pRedictor) method takes advantage of three innovations: comprehensive dataset that covers multiple types of binding residues, novel types of inputs that are relevant to the prediction of PBRs, and an architecture that is tailored to reduce the cross-predictions. The dataset includes complete protein chains and offers improved coverage of binding annotations that are transferred from multiple protein-protein complexes. We utilize innovative two-layer architecture where the first layer generates a prediction of protein-binding, RNA-binding, DNA-binding and small ligand-binding residues. The second layer re-predicts PBRs by reducing overlap between PBRs and the other types of binding residues produced in the first layer. Empirical tests on an independent test dataset reveal that SCRIBER significantly outperforms current predictors and that all three innovations contribute to its high predictive performance. SCRIBER reduces cross-predictions by between 41% and 69% and our conservative estimates show that it is at least 3 times faster. We provide putative PBRs produced by SCRIBER for the entire human proteome and use these results to hypothesize that about 14% of currently known human protein domains bind proteins. AVAILABILITY AND IMPLEMENTATION: SCRIBER webserver is available at http://biomine.cs.vcu.edu/servers/SCRIBER/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jian Zhang 0020, Lukasz A. Kurgan |
Bioinform. | 1 |
| 2018 | Review and comparative assessment of sequence-based predictors of protein-binding residuesabstractUnderstanding of molecular mechanisms that govern protein-protein interactions and accurate modeling of protein-protein docking rely on accurate identification and prediction of protein-binding partners and protein-binding residues. We review over 40 methods that predict protein-protein interactions from protein sequences including methods that predict interacting protein pairs, protein-binding residues for a pair of interacting sequences and protein-binding residues in a single protein chain. We focus on the latter methods that provide residue-level annotations and that can be broadly applied to all protein sequences. We compare their architectures, inputs and outputs, and we discuss aspects related to their assessment and availability. We also perform first-of-its-kind comprehensive empirical comparison of representative predictors of protein-binding residues using a novel and high-quality benchmark data set. We show that the selected predictors accurately discriminate protein-binding and non-binding residues and that newer methods outperform older designs. However, these methods are unable to accurately separate residues that bind other molecules, such as DNA, RNA and small ligands, from the protein-binding residues. This cross-prediction, defined as the incorrect prediction of nucleic-acid- and small-ligand-binding residues as protein binding, is substantial for all evaluated methods and is not driven by the proximity to the native protein-binding residues. We discuss reasons for this drawback and we offer several recommendations. In particular, we postulate the need for a new generation of more accurate predictors and data sets, inclusion of a comprehensive assessment of the cross-predictions in future studies and higher standards of availability of the published methods. Jian Zhang 0020, Lukasz A. Kurgan |
Briefings Bioinform. | 1 |
| 2018 | HEMEsPred: Structure-Based Ligand-Specific Heme Binding Residues Prediction by Using Fast-Adaptive Ensemble Learning SchemeabstractHeme is an essential biomolecule that widely exists in numerous extant organisms. Accurately identifying heme binding residues (HEMEs) is of great importance in disease progression and drug development. In this study, a novel predictor named HEMEsPred was proposed for predicting HEMEs. First, several sequence- and structure-based features, including amino acid composition, motifs, surface preferences, and secondary structure, were collected to construct feature matrices. Second, a novel fast-adaptive ensemble learning scheme was designed to overcome the serious class-imbalance problem as well as to enhance the prediction performance. Third, we further developed ligand-specific models considering that different heme ligands varied significantly in their roles, sizes, and distributions. Statistical test proved the effectiveness of ligand-specific models. Experimental results on benchmark datasets demonstrated good robustness of our proposed method. Furthermore, our method also showed good generalization capability and outperformed many state-of-art predictors on two independent testing datasets. HEMEsPred web server was available at http://www.inforstation.com/HEMEsPred/ for free academic use. Jian Zhang 0020, Haiting Chai, Guifu Yang, Zhiqiang Ma 0003 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2017 | Prediction of bioluminescent proteins by using sequence-derived features and lineage-specific schemeabstractBACKGROUND: Bioluminescent proteins (BLPs) widely exist in many living organisms. As BLPs are featured by the capability of emitting lights, they can be served as biomarkers and easily detected in biomedical research, such as gene expression analysis and signal transduction pathways. Therefore, accurate identification of BLPs is important for disease diagnosis and biomedical engineering. In this paper, we propose a novel accurate sequence-based method named PredBLP (Prediction of BioLuminescent Proteins) to predict BLPs. RESULTS: We collect a series of sequence-derived features, which have been proved to be involved in the structure and function of BLPs. These features include amino acid composition, dipeptide composition, sequence motifs and physicochemical properties. We further prove that the combination of four types of features outperforms any other combinations or individual features. To remove potential irrelevant or redundant features, we also introduce Fisher Markov Selector together with Sequential Backward Selection strategy to select the optimal feature subsets. Additionally, we design a lineage-specific scheme, which is proved to be more effective than traditional universal approaches. CONCLUSION: Experiment on benchmark datasets proves the robustness of PredBLP. We demonstrate that lineage-specific models significantly outperform universal ones. We also test the generalization capability of PredBLP based on independent testing datasets as well as newly deposited BLPs in UniProt. PredBLP is proved to be able to exceed many state-of-art methods. A web server named PredBLP, which implements the proposed method, is free available for academic use. Jian Zhang 0020, Haiting Chai, Guifu Yang, Zhiqiang Ma 0003 |
BMC Bioinform. | 1 |
| 2016 | Identification of DNA-binding proteins using multi-features fusion and binary firefly optimization algorithmabstractBACKGROUND: DNA-binding proteins (DBPs) play fundamental roles in many biological processes. Therefore, the developing of effective computational tools for identifying DBPs is becoming highly desirable. RESULTS: In this study, we proposed an accurate method for the prediction of DBPs. Firstly, we focused on the challenge of improving DBP prediction accuracy with information solely from the sequence. Secondly, we used multiple informative features to encode the protein. These features included evolutionary conservation profile, secondary structure motifs, and physicochemical properties. Thirdly, we introduced a novel improved Binary Firefly Algorithm (BFA) to remove redundant or noisy features as well as select optimal parameters for the classifier. The experimental results of our predictor on two benchmark datasets outperformed many state-of-the-art predictors, which revealed the effectiveness of our method. The promising prediction performance on a new-compiled independent testing dataset from PDB and a large-scale dataset from UniProt proved the good generalization ability of our method. In addition, the BFA forged in this research would be of great potential in practical applications in optimization fields, especially in feature selection problems. CONCLUSIONS: A highly accurate method was proposed for the identification of DBPs. A user-friendly web-server named iDbP (identification of DNA-binding Proteins) was constructed and provided for academic use. Jian Zhang 0020, Haiting Chai, Zhiqiang Ma 0003, Guifu Yang |
BMC Bioinform. | 1 |