VLDB 2026 Research / reviewers in the wild / expert
Jing Hu 0003
dblp:95/6046-3
· DBLP profile ↗
27ranked-venue papers
9as first author
18since 2021 · last 2026
0000-0003-1348-8773ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 26 · 9 first-author · 17 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DBGT-PLA: Dual-Branch Graph-Transformer Fusion for Interpretable Protein- Ligand Affinity PredictionabstractProtein-ligand binding affinity prediction is critical for drug discovery, yet existing methods struggle to jointly model local atomic interactions and global contextual dependencies. To address this, we propose the Interpretable Dual-Branch Graph-Transformer framework for Protein-Ligand Affinity prediction (DBGT-PLA), a novel dual-branch architecture that integrates graph neural network (GNN) with a stability-enhanced Transformer equipped with learnable positional embeddings and a NaN-filtering mechanism that handles potential Not-a-Number (NaN) values arising from numerical instability or data preprocessing. We design a Gated Residual Learning (GRL) Fusion module that performs dimension-wise adaptive integration between local graph topology and global Transformer context. This mechanism enables multi-level feature coordination through a residual path, achieving biophysically consistent alignment between atomic-level interactions and global conformational dependencies. Furthermore, we introduce an edge-level Shapley attribution framework tailored to protein-ligand interaction graphs, quantifying contributions of chemical bonds (e.g., hydrophobic contacts) and non-covalent interactions. Experiments show DBGT-PLA reduces RMSE by 18.3% (from 1.522 to 1.244 on the Holdout Set 2019), outperforming state-of-the-art models. Crucially, our explainability module reveals that the ligand edges dominate affinity predictions, accounting for nearly 70%. This work not only advances predictive accuracy but also offers unprecedented, quantitative insights into interaction determinants, which can guide rational drug optimization. Jing Hu 0003, Junlin Xu, Bo Li 0002 |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | Dynamic Co-Evolution Mechanism and Multi-View Collaborative Distillation Optimization for Semi-Supervised Medical Image SegmentationabstractDespite the ability of semi-supervised medical image segmentation to achieve superior results with limited labeled data and extensive unlabeled data, prevailing methods encounter a significant challenge: the detrimental cycle of pseudo-label noise resulting in model bias and error accumulation, which ultimately degrades performance. In this paper, we propose a differential teacher dynamic co-evolution mechanism: initially, an uncertainty perception-driven confidence fusion module (UPC) is developed, converting the uncertainty region into learning opportunities to maintain essential anatomical structures. The Dynamic Update Mechanism (DUM) is concurrently proposed to adaptively select and enhance the teacher models utilizing the pixel-level confidence graph to mitigate noise propagation. Furthermore, we propose multi-view collaborative distillation optimization (MCD), which facilitates bidirectional knowledge transfer between teachers and students via multi-level knowledge distillation, and addresses mistake accumulation to transcend the noise-induced local decision-making boundary. Experiments on the LA and ACDC datasets demonstrate that our method markedly outperforms current semi-supervised segmentation methods. Haoyu Yin, Jing Hu 0003 |
BIBM | 2 |
| 2025 | Low Rank Representation Based on Pseudo-label Learning
Wei-Jia Liu, Jing Hu 0003, Bo Li 0002 |
ICIC (9) | 2 |
| 2025 | scSAMAC: saliency-adjusted masking induced attention contrastive learning for single-cell clusteringabstractSingle-cell sequencing technology has enabled researchers to study cellular heterogeneity at the cell level. To facilitate the downstream analysis, clustering single-cell data into subgroups is essential. However, the high dimensionality, sparsity, and dropout events of the data make the clustering challenging. Currently, many deep learning methods have been proposed. Nevertheless, they either fail to fully utilize pairwise distances information between similar cells, or do not adequately capture their feature correlations. They cannot also effectively handle high-dimensional sparse data. Therefore, they are not suitable for high-fidelity clustering, leading to difficulties in analyzing the clear cell types required for downstream analysis. The proposed scSAMAC method integrates contrastive learning and negative binomial losses into a variational autoencoder, extracting features via contrastive unit similarity while preserving the intrinsic characteristics. This enhances the robustness and generalization during the clustering. In the contrastive learning, it constructs a mask module by adopting a negative sample generation method with gene feature saliency adjustment, which selects features more influential in the clustering phase and simulates data missing events. Additionally, it develops a novel loss, which consists of a soft k-means loss, a Wasserstein distance, and a contrastive loss. This fully utilizes data information and improves clustering performance. Furthermore, a multi-head attention mechanism module is applied to the latent variables at each layer of autoencoder to enhance feature correlation, integration, and information repair. Experimental results demonstrate that scSAMAC outperforms several state-of-the-art clustering methods. Bo Li 0002, Yongkang Zhao, Jing Hu 0003, Xiaolong Zhang 0002 |
Briefings Bioinform. | 3 |
| 2025 | An Efficient Targeted Drug Design Method Using Framework of Multiscale Encoder-DecoderabstractThis paper describes a targeted drug design method based on a framework of multiscale encoder-decoder. Encoders are used to encode target gene and protein features. A decoder is used to design drugs based on target features. This method fuses target gene and protein information for targeted drug design, and invokes effective feature extraction strategies. A multilevel gene feature extraction (MGFE) is proposed to extract multilevel target features by extracting base and codon features in gene expression. The process of extracting features from nucleotide sequences by MGFE based gene encoder simulates the process of gene transcription and translation. Meanwhile, a multi-embedding protein feature extraction (MPFE) is proposed to extract target protein features from amino acid sequences. The MPFE based protein encoder includes three embedding layers which provides a unique linear layer for each amino acid. According to structural characteristics of proteins, amino acids with different positions but the same type are embedded into the same embedding vector without location encoding. Finally, a gated recurrent unit based drug decoder is used to decode gene and protein features, and creates new targeted drugs. The experiments adminstrate that the proposed method outperforms the previous ones in terms of validity, novelty and binding affinity. Xiaoli Lin, Jing Hu 0003, Jun Pang 0002, Xiaolong Zhang 0002 |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2024 | A Multi Target Drug Design Method Based on Target Protein Sequence and Feature SimilarityabstractThis paper describes a multi target drug design method based on features of the target proteins. The multi target drug which inhibit multiple proteins have prospective applications, but design is difficult. In this study, the target protein sequences are embedded to obtain target features. The features of each target are independently encoded as target latent vectors, and the features of multiple targets are jointly encoded as target similarity latent vectors. Based on the features of targets and the similarity features among the targets, multi target drugs are efficiently designed. In the experiments, the designed multi target drugs can be docked with target proteins with a proprietary molecular structure according to the requirements of different target pocket structures. The excellent fitness among the molecular structure of multi target drugs and the protein structure of multiple targets confirms the performance of the proposed method. In the molecular docking part of the experi¬ments, the binding affinity of designed multi target drugs is far better than that in the previous studies, which administrates the better performance of this work. Xiaoli Lin, Jing Hu 0003, Xiaolong Zhang 0002 |
BIBM | 3 |
| 2024 | KGRLFF: Detecting Drug-Drug Interactions Based on Knowledge Graph Representation Learning and Feature FusionabstractAccurate prediction of drug-drug interactions (DDIs) plays an important role in improving the efficiency of drug development and ensuring the safety of combination therapy. Most existing models rely on a single source of information to predict DDIs, and few models can perform tasks on biomedical knowledge graphs. This paper proposes a new hybrid method, namely Knowledge Graph Representation Learning and Feature Fusion (KGRLFF), to fully exploit the information from the biomedical knowledge graph and molecular structure of drugs to better predict DDIs. KGRLFF first uses a Bidirectional Random Walk sampling method based on the PageRank algorithm (BRWP) to obtain higher-order neighborhood information of drugs in the knowledge graph, including neighboring nodes, semantic relations, and higher-order information associated with triple facts. Then, an embedded representation learning model named Knowledge Graph-based Cyclic Recursive Aggregation (KGCRA) is used to learn the embedded representations of drugs by recursively propagating and aggregating messages with drugs as both the source and destination. In addition, the model learns the molecular structures of the drugs to obtain the structured features. Finally, a Feature Representation Fusion Strategy (FRFS) was developed to integrate embedded representations and structured feature representations. Experimental results showed that KGRLFF is feasible for predicting potential DDIs. Xiaoli Lin, Zhuang Yin, Xiaolong Zhang 0002, Jing Hu 0003 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2023 | Drug-Target Interaction Prediction Based on Drug Subgraph Fingerprint Extraction Strategy and Subgraph Attention Mechanism
Xiaolong Zhang 0002, Xiaoli Lin, Jing Hu 0003 |
ADMA (3) | 4 |
| 2023 | DU-DANet: Efficient 3D Automatic Brain Tumor Segmentation Based on Dual Attention
Zhenhua Cai, Xiaoli Lin, Xiaolong Zhang 0002, Jing Hu 0003 |
ICIC (3) | 4 |
| 2023 | De Novo Drug Design Using Unified Multilayer Simple Recurrent Unit Model
Zonghao Li, Jing Hu 0003, Xiaolong Zhang 0002 |
ICIC (3) | 2 |
| 2023 | An Efficient Drug Design Method Based on Drug-Target Affinity
Xiaolong Zhang 0002, Xiaoli Lin, Jing Hu 0003 |
ICIC (3) | 4 |
| 2023 | Identifying Drug-Target Interactions Through a Combined Graph Attention Mechanism and Self-attention Sequence Embedding Model
Jing Hu 0003, Xiaolong Zhang 0002 |
ICIC (3) | 2 |
| 2023 | Drug-Target Affinity Prediction Based on Self-attention Graph Pooling and Mutual Interaction Neural Network
Jing Hu 0003, Xiaolong Zhang 0002 |
ICIC (3) | 2 |
| 2022 | Drug-Target Binding Affinity Prediction Based on Graph Neural Networks and Word2vec
Minghao Xia, Jing Hu 0003, Xiaolong Zhang 0002, Xiaoli Lin |
ICIC (2) | 2 |
| 2022 | Drug-Target Affinity Prediction Based on Multi-channel Graph Convolution
Jing Hu 0003, Xiaolong Zhang 0002 |
ICIC (2) | 2 |
| 2022 | Unsupervised Prediction Method for Drug-Target Interactions Based on Structural Similarity
Xiaoli Lin, Jing Hu 0003, Wenquan Ding |
ICIC (2) | 3 |
| 2021 | Prediction of hot spots in protein-protein interaction by Nine-Pipeline & Ensemble Learning strategyabstractThis paper proposes a NPEL (Nine-Pipeline & Ensemble Learning) strategy based on machine Learning algorithm to predict protein-protein interaction hotspots by training amino acid composition, surface area, amino acid chains and other complex/interface-related structural information. We applied Random Forest, Linear Svm, KNN, Gaussian Naive Bayes, Multi-layer Perceptron Neural Network, Adaboost, XGBoost etc. nine machine learning algorithms combination into an independent pipeline to predict protein hot spots, and the final results are optimized through voting and stacking scheme. In the stacking result of XGBoost and Logistic Regression, the highest accuracy is 0.8462 and improve the indicators of the pipeline results greatly. Jing Hu 0003, Zonghao Li, Xiaolong Zhang 0002, Nansheng Chen |
BIBM | 1 |
| 2021 | Improve hot region prediction by analyzing different machine learning algorithmsabstractBACKGROUND: In the process of designing drugs and proteins, it is crucial to recognize hot regions in protein-protein interactions. Each hot region of protein-protein interaction is composed of at least three hot spots, which play an important role in binding. However, it takes time and labor force to identify hot spots through biological experiments. If predictive models based on machine learning methods can be trained, the drug design process can be effectively accelerated. RESULTS: The results show that different machine learning algorithms perform similarly, as evaluating using the F-measure. The main differences between these methods are recall and precision. Since the key attribute of hot regions is that they are packed tightly, we used the cluster algorithm to predict hot regions. By combining Gaussian Naïve Bayes and DBSCAN, the F-measure of hot region prediction can reach 0.809. CONCLUSIONS: In this paper, different machine learning models such as Gaussian Naïve Bayes, SVM, Xgboost, Random Forest, and Artificial Neural Network are used to predict hot spots. The experiment results show that the combination of hot spot classification algorithm with higher recall rate and clustering algorithm with higher precision can effectively improve the accuracy of hot region prediction. Jing Hu 0003, Longwei Zhou, Bo Li 0002, Xiaolong Zhang 0002, Nansheng Chen |
BMC Bioinform. | 1 |
| 2019 | Identification of protein hot regions by integrated machine learning algorithmabstractDiscovering hot regions in protein-protein interaction is important for understanding the interactions between proteins, while because of the complexity and time-consuming of experimental methods, the computational prediction method can be very helpful to improve the efficiency to predict hot regions. In previous researches, some models are based on a single aspect, such as structure, energy, and sequence, each aspect has its advantage and limitations. In this paper, a new method that combing structure-based classification, energy-based clustering and sequence-based conservation in evolution is proposed. This method makes full use of three aspects of protein information and compensates for the limitations of using one single aspect information. Experimental results show that the proposed method significantly improves the prediction performance of hot regions. Jing Hu 0003, Haomin Gan, Xiaolong Zhang 0002, Nansheng Chen |
BIBM | 1 |
| 2019 | Improving Hot Region Prediction by Combining Gaussian Naive Bayes and DBSCAN
Jing Hu 0003, Longwei Zhou, Xiaolong Zhang 0002, Nansheng Chen |
ICIC (2) | 1 |
| 2018 | Analysis of hot regions prediction in PPI with different amino acid mutation using machine learning algorithm
Jing Hu 0003, Haomin Gan, Xiaolong Zhang 0002, Nansheng Chen |
BIBM | 1 |
| 2018 | Accurate Prediction of Hot Spots with Greedy Gradient Boosting Decision Tree
Haomin Gan, Jing Hu 0003, Xiaolong Zhang 0002, Jiafu Zhao |
ICIC (2) | 2 |
| 2017 | Protein Hot Regions Feature Research Based on Evolutionary Conservation
Jing Hu 0003, Xiaoli Lin, Xiaolong Zhang 0002 |
ICIC (2) | 1 |
| 2017 | Classification of Hub Protein and Analysis of Hot Regions in Protein-Protein Interactions
Xiaoli Lin, Xiaolong Zhang 0002, Jing Hu 0003 |
ICIC (2) | 3 |
| 2015 | Prediction of hot regions in protein-protein interaction by density-based incremental clustering with parameter selectionabstractThis paper studies how to select input parameters in density clustering of hot region prediction. There are two parameters radius and density in density-based incremental clustering. We firstly fix density and enumerate radius to find a pair of parameters which leads to maximum number of clusters, and then we fix radius and enumerate density to find another pair of parameters which leads to maximum number of clusters. Experiment results show that the proposed method using both two pairs of parameters provides better prediction performance than the other method, and compare these two predictive results, the result by fixing radius and enumerating density have slightly higher prediction accuracy than that by fixing density and enumerating radius. Jing Hu 0003, Xiaolong Zhang 0002 |
BIBM | 1 |
| 2015 | Testing whether hot regions in protein-protein interactions are conserved in different speciesabstractWe examined the conservation of hot regions in protein-protein interaction in evolution for the first time. We studied sequence conservation of hot regions in different species annotated in the database ASEdb, applying the BLOSUM matrix to construct conservation scoring function of hot regions. The experimental results showed that there is high correlation of conservation between hot regions in different species. We further tested the conservation of hot regions in different species annotated in the latest SKEMPI database using the same method. Our experimental results suggest that there is obvious conservation in hot regions. Jing Hu 0003, Xiaolong Zhang 0002 |
BIBM | 1 |
| 2015 | Identification of Hot Regions in Protein Interfaces: Combining Density Clustering and Neighbor Residues Improves the Accuracy
Jing Hu 0003, Xiaolong Zhang 0002 |
ICIC (2) | 1 |