EDBT 2026 Demo / reviewers in the wild / expert
Jun Hu 0010
dblp:28/441-10
· DBLP profile ↗
11ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0001-9202-7515ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RLEAAI: improving antibody-antigen interaction prediction using protein language model and sequence order informationabstractAntibody-antigen interactions (AAIs) are a pervasive phenomenon in the natural and are instrumental in the design of antibody-based drugs. Despite the emergence of various deep learning-based methods aimed at enhancing the accuracy of AAIs predictions, most of these approaches overlook the significance of sequence order information. In this study, we propose a new deep learning-based method RLEAAI, to improve the prediction performance of AAIs. In RLEAAI, a sequence order extraction strategy, called Composition of K-Spaced Amino Acid Pairs, is employed to generate feature representation from the feature embedding outputted by a pre-trained protein language model. In order to fully dig out the discrimination information from features, three neural network modules, i.e. convolutional neural network, bidirectional long short-term memory network and recurrent criss-cross attention mechanism, are integrated. Benchmarked results on two independent test sets demonstrate that RLEAAI is capable of achieving an average accuracy of 0.7787 and an average Matthews's correlation coefficient (MCC) value of 0.5552, representing a 5.2% and 15.8% improvement over the start-of-the-art method DeepAAI. Furthermore, the complementary determining regions-sensitivity value calculated on MCC of RLEAAI is 216.4% higher than that of the state-of-the-art method DeepAAI. The standalone package of RLEAAI is freely available at https://github.com/zhouyu9931/RLEAAI.git. Jun Hu 0010, Wen-Yi Zhang |
Briefings Bioinform. | 1 |
| 2024 | Ense-i6mA: Identification of DNA N6-Methyladenine Sites Using XGB-RFE Feature Selection and Ensemble Machine Learningabstract-methyladenine (6mA) is an important epigenetic modification that plays a vital role in various cellular processes. Accurate identification of the 6mA sites is fundamental to elucidate the biological functions and mechanisms of modification. However, experimental methods for detecting 6mA sites are high-priced and time-consuming. In this study, we propose a novel computational method, called Ense-i6mA, to predict 6mA sites. Firstly, five encoding schemes, i.e., one-hot encoding, gcContent, Z-Curve, K-mer nucleotide frequency, and K-mer nucleotide frequency with gap, are employed to extract DNA sequence features. Secondly, eXtreme gradient boosting coupled with recursive feature elimination is applied to remove noisy features for avoiding over-fitting, reducing computing time and complexity. Then, the best subset of features is fed into base-classifiers composed of Extra Trees, eXtreme Gradient Boosting, Light Gradient Boosting Machine, and Support Vector Machine. Finally, to minimize generalization errors, the prediction probabilities of the base-classifiers are aggregated by averaging for inferring the final 6mA sites results. We conduct experiments on two species, i.e., Arabidopsis thaliana and Drosophila melanogaster, to compare the performance of Ense-i6mA against the recent 6mA sites prediction methods. The experimental results demonstrate that the proposed Ense-i6mA achieves area under the receiver operating characteristic curve values of 0.967 and 0.968, accuracies of 91.4% and 92.0%, and Mathew's correlation coefficient values of 0.829 and 0.842 on two benchmark datasets, respectively, and outperforms several existing state-of-the-art methods. Xueqiang Fan, Jun Hu 0010, Zhongyi Guo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | ATPdock: a template-based method for ATP-specific protein-ligand dockingabstractMOTIVATION: Accurately identifying protein-ATP binding poses is significantly valuable for both basic structure biology and drug discovery. Although many docking methods have been designed, most of them require a user-defined binding site and are difficult to achieve a high-quality protein-ATP docking result. It is critical to develop a protein-ATP-specific blind docking method without user-defined binding sites. RESULTS: Here, we present ATPdock, a template-based method for docking ATP into protein. For each query protein, if no pocket site is given, ATPdock first identifies its most potential pocket using ATPbind, an ATP-binding site predictor; then, the template pocket, which is most similar to the given or identified pocket, is searched from the database of pocket-ligand structures using APoc, a pocket structural alignment tool; thirdly, the rough docking pose of ATP (rdATP) is generated using LS-align, a ligand structural alignment tool, to align the initial ATP pose to the template ligand corresponding to template pocket; finally, the Metropolis Monte Carlo simulation is used to fine-tune the rdATP under the guidance of AutoDock Vina energy function. Benchmark tests show that ATPdock significantly outperforms other state-of-the-art methods in docking accuracy. AVAILABILITY AND IMPLEMENTATION: https://jun-csbio.github.io/atpdock/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Liang Rao, Ning-Xin Jia, Jun Hu 0010, Dongjun Yu, Guijun Zhang |
Bioinform. | 3 |
| 2022 | Robust distance metric optimization driven GEPSVM classifier for pattern classification
Liyong Fu, Tian'an Zhang, Jun Hu 0010, Qiaolin Ye, Yong Qi 0002, Dongjun Yu |
Pattern Recognit. | 4 |
| 2022 | Protein-DNA Binding Residue Prediction via Bagging Strategy and Sequence-Based Cube-Format FeatureabstractProtein-DNA interactions play an important role in diverse biological processes. Accurately identifying protein-DNA binding residues is a critical but challenging task for protein function annotations and drug design. Although wet-lab experimental methods are the most accurate way to identify protein-DNA binding residues, they are time consuming and labor intensive. There is an urgent need to develop computational methods to rapidly and accurately predict protein-DNA binding residues. In this study, we propose a novel sequence-based method, named PredDBR, for predicting DNA-binding residues. In PredDBR, for each query protein, its position-specific frequency matrix (PSFM), predicted secondary structure (PSS), and predicted probabilities of ligand-binding residues (PPLBR) are first generated as three feature sources. Secondly, for each feature source, the sliding window technique is employed to extract the matrix-format feature of each residue. Then, we design two strategies, i.e., square root (SR) and average (AVE), to separately transform PSFM-based and two predicted feature source-based, i.e., PSS-based and PPLBR-based, matrix-format features of each residue into three corresponding cube-format features. Finally, after serially combining the three cube-format features, the ensemble classifier is generated via applying bagging strategy to multiple base classifiers built by the framework of 2D convolutional neural network. The computational experimental results demonstrate that the proposed PredDBR achieves an average overall accuracy of 93.7% and a Mathew's correlation coefficient of 0.405 on two independent validation datasets and outperforms several state-of-the-art sequenced-based protein-DNA binding residue predictors. The PredDBR web-server is available at https://jun-csbio.github.io/PredDBR/. Jun Hu 0010, Yan-Song Bai, Lin-Lin Zheng, Ning-Xin Jia, Dongjun Yu, Guijun Zhang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2021 | Accurate multistage prediction of protein crystallization propensity using deep-cascade forest with sequence-based featuresabstractX-ray crystallography is the major approach for determining atomic-level protein structures. Because not all proteins can be easily crystallized, accurate prediction of protein crystallization propensity provides critical help in guiding experimental design and improving the success rate of X-ray crystallography experiments. This study has developed a new machine-learning-based pipeline that uses a newly developed deep-cascade forest (DCF) model with multiple types of sequence-based features to predict protein crystallization propensity. Based on the developed pipeline, two new protein crystallization propensity predictors, denoted as DCFCrystal and MDCFCrystal, have been implemented. DCFCrystal is a multistage predictor that can estimate the success propensities of the three individual steps (production of protein material, purification and production of crystals) in the protein crystallization process. MDCFCrystal is a single-stage predictor that aims to estimate the probability that a protein will pass through the entire crystallization process. Moreover, DCFCrystal is designed for general proteins, whereas MDCFCrystal is specially designed for membrane proteins, which are notoriously difficult to crystalize. DCFCrystal and MDCFCrystal were separately tested on two benchmark datasets consisting of 12 289 and 950 proteins, respectively, with known crystallization results from various experimental records. The experimental results demonstrated that DCFCrystal and MDCFCrystal increased the value of Matthew's correlation coefficient by 199.7% and 77.8%, respectively, compared to the best of other state-of-the-art protein crystallization propensity predictors. Detailed analyses show that the major advantages of DCFCrystal and MDCFCrystal lie in the efficiency of the DCF model and the sensitivity of the sequence-based features used, especially the newly designed pseudo-predicted hybrid solvent accessibility (PsePHSA) feature, which improves crystallization recognition by incorporating sequence-order information with solvent accessibility of residues. Meanwhile, the new crystal-dataset constructions help to train the models with more comprehensive crystallization knowledge. Yiheng Zhu 0001, Jun Hu 0010, Fang Ge, Fuyi Li, Jiangning Song, Yang Zhang 0040, Dongjun Yu |
Briefings Bioinform. | 2 |
| 2021 | Learning to capture dependencies between global features of different convolution layers
Zhangwei Li, Anshun Hu, Jun Hu 0010, Guijun Zhang |
J. Vis. Commun. Image Represent. | 4 |
| 2021 | Protein Structure Prediction Using Population-Based Algorithm Guided by Information EntropyabstractAb initio protein structure prediction is one of the most challenging problems in computational biology. Multistage algorithms are widely used in ab initio protein structure prediction. The different computational costs of a multistage algorithm for different proteins are important to be considered. In this study, a population-based algorithm guided by information entropy (PAIE), which includes exploration and exploitation stages, is proposed for protein structure prediction. In PAIE, an entropy-based stage switch strategy is designed to switch from the exploration stage to the exploitation stage. Torsion angle statistical information is also deduced from the first stage and employed to enhance the exploitation in the second stage. Results indicate that an improvement in the performance of protein structure prediction in a benchmark of 30 proteins and 17 other free modeling targets in CASP. Guijun Zhang, Tengyu Xie, Liu-Jing Wang, Jun Hu 0010 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2020 | scTPA: a web tool for single-cell transcriptome analysis of pathway activation signaturesabstractMOTIVATION: At present, a fundamental challenge in single-cell RNA-sequencing data analysis is functional interpretation and annotation of cell clusters. Biological pathways in distinct cell types have different activation patterns, which facilitates the understanding of cell functions using single-cell transcriptomics. However, no effective web tool has been implemented for single-cell transcriptome data analysis based on prior biological pathway knowledge. RESULTS: Here, we present scTPA, a web-based platform for pathway-based analysis of single-cell RNA-seq data in human and mouse. scTPA incorporates four widely-used gene set enrichment methods to estimate the pathway activation scores of single cells based on a collection of available biological pathways with different functional and taxonomic classifications. The clustering analysis and cell-type-specific activation pathway identification were provided for the functional interpretation of cell types from a pathway-oriented perspective. An intuitive interface allows users to conveniently visualize and download single-cell pathway signatures. Overall, scTPA is a comprehensive tool for the identification of pathway activation signatures for the analysis of single cell heterogeneity. AVAILABILITY AND IMPLEMENTATION: http://sctpa.bio-data.cn/sctpa. CONTACT: [email protected] or [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yaru Zhang, Jun Hu 0010, Fangjie Guo, Meng Zhou 0003, Guijun Zhang, Fulong Yu, Jianzhong Su |
Bioinform. | 3 |
| 2020 | TargetDBP: Accurate DNA-Binding Protein Prediction Via Sequence-Based Multi-View Feature LearningabstractAccurately identifying DNA-binding proteins (DBPs) from protein sequence information is an important but challenging task for protein function annotations. In this paper, we establish a novel computational method, named TargetDBP, for accurately targeting DBPs from primary sequences. In TargetDBP, four single-view features, i.e., AAC (Amino Acid Composition), PsePSSM (Pseudo Position-Specific Scoring Matrix), PsePRSA (Pseudo Predicted Relative Solvent Accessibility), and PsePPDBS (Pseudo Predicted Probabilities of DNA-Binding Sites), are first extracted to represent different base features, respectively. Second, differential evolution algorithm is employed to learn the weights of four base features. Using the learned weights, we weightedly combine these base features to form the original super feature. An excellent subset of the super feature is then selected by using a suitable feature selection algorithm SVM-REF+CBR (Support Vector Machine Recursive Feature Elimination with Correlation Bias Reduction). Finally, the prediction model is learned via using support vector machine on the selected feature subset. We also construct a new gold-standard and non-redundant benchmark dataset from PDB database to evaluate and compare the proposed TargetDBP with other existing predictors. On this new dataset, TargetDBP can achieve higher performance than other state-of-the-art predictors. The TargetDBP web server and datasets are freely available at http://csbio.njust.edu.cn/bioinf/targetdbp/ for academic use. Jun Hu 0010, Yiheng Zhu 0001, Dongjun Yu, Guijun Zhang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2020 | Two-Stage Distance Feature-based Optimization Algorithm for De novo Protein Structure PredictionabstractDe novo protein structure prediction can be treated as a conformational space optimization problem under the guidance of an energy function. However, it is a challenge of how to design an accurate energy function which ensures low-energy conformations close to native structures. Fortunately, recent studies have shown that the accuracy of de novo protein structure prediction can be significantly improved by integrating the residue-residue distance information. In this paper, a two-stage distance feature-based optimization algorithm (TDFO) for de novo protein structure prediction is proposed within the framework of evolutionary algorithm. In TDFO, a similarity model is first designed by using feature information which is extracted from distance profiles by bisecting K-means algorithm. The similarity model-based selection strategy is then developed to guide conformation search, and thus improve the quality of the predicted models. Moreover, global and local mutation strategies are designed, and a state estimation strategy is also proposed to strike a trade-off between the exploration and exploitation of the search space. Experimental results of 35 benchmark proteins show that the proposed TDFO can improve prediction accuracy for a large portion of test proteins. Guijun Zhang, Xiao-Qi Wang, Lai-Fa Ma, Liu-Jing Wang, Jun Hu 0010 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |