EDBT 2026 Demo / reviewers in the wild / expert
Han Wang 0028
dblp:67/1771-28
· DBLP profile ↗
17ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0002-4302-1886ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 6 first-author · 15 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SSGraphDTI: A Drug-Target Interaction Prediction Method Integrated Structural and Dynamic Systemic Biology AttributesabstractDrug-Target Interaction (DTI) is a crucial aspect of pharmaceutical development. However, biochemical experiments are prohibitively expensive to identify these interactions on a large scale, while the computational approach is still on the way to making a highly reliable prediction. For the purpose of promoting prediction accuracy, drug-related molecular networks are gradually introduced to this task to furnish valuable information. We hypothesized that integrating structural and systemic biological attributes could effectively enhance the performance of DTI prediction and proposed a novel DTI prediction model, SSGraphDTI, which integrated two aforementioned attributes. Specifically, the structural attributes of drugs and targets are extracted using independent convolutional neural network based models from the Simplified Molecular Input Line Entry System of drugs and the amino acid sequences of targets, respectively. Meanwhile, the systemic biological attributes of drug-target pairs are obtained through graph representation learning on the dynamically constructed heterogeneous drug-target interaction network. SSGraphDTI was meticulously trained and rigorously tested on the benchmark Dataset_DrugBank, achieving an improvement of approximately 1.0% across five metrics compared to recent comparable methods. These results underscore the potential of combining both structural and systemic information for accurate DTI prediction. Benefiting from the fact that the input consists solely of structural data without requiring interaction information, the model effectively addresses the "cold-start problem" in drug discovery. Furthermore, by extracting systemic attributes directly from the dynamically constructed DTI networks, the model maintains strong predictive performance even when data is limited. Haotian Guan, Tian Bai 0002, Jingtong Zhao, Han Wang 0028 |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | HyperDTI-Lite: Hyperbolic Geometry for Enhanced Drug-Target Interaction Prediction via Heterogeneous Feature FusionabstractAccurate prediction of drug-target interactions (DTIs) is crucial for accelerating drug discovery and repurposing efforts. However, existing computational models face significant challenges in effectively integrating heterogeneous features and capturing the complex hierarchical relationships inherent in biological systems. To address these limitations, we propose HyperDTI-Lite, a novel and efficient DTI prediction framework that integrates homologous heterogeneous features, including structural features and physicochemical properties, with semantic embeddings obtained from pretrained language models. By projecting these diverse features into a unified hyperbolic embedding space through a hyperbolic multilayer perceptron, HyperDTILite effectively captures complex dependencies between drug and target features. Extensive experiments on the benchmark DrugBank dataset demonstrate that HyperDTI-Lite consistently outperforms state-of-the-art methods across multiple evaluation metrics, achieving superior predictive accuracy. Comprehensive ablation studies validate the effectiveness of each key component in our framework. This work highlights the advantages of hyperbolic geometry for modeling complex biological data and provides a robust foundation for future drug discovery research. Haotian Guan, Tian Bai 0002, Chuande Yang, Jingtong Zhao, Han Wang 0028 |
BIBM | 6 |
| 2025 | A Pan-Tissue Epigenetic Aging Clock Across the Human Lifespan: Robust Age Prediction from DNA Methylation ProfilesabstractDNA methylation serves as a pivotal epigenetic mechanism that dynamically regulates gene expression, thereby influencing development, aging, and the onset of various diseases. In recent years, epigenetic clocks based on DNA methylation patterns have emerged as powerful tools for estimating biological age. However, many existing models still face limitations in predictive accuracy and offer limited insight into the molecular underpinnings of age-related disorders. To address these challenges, we developed a high-precision methylation-based aging clock using genome-wide DNA methylation data from the Illumina 450K array, encompassing a wide age range and multiple tissue types. Leveraging a deep learning framework based on the Mamba architecture, we constructed an age prediction model that achieves state-of-the-art performance in biological age estimation. In addition to enhanced predictive accuracy, the model offers a valuable framework for evaluating the efficacy of anti-aging interventions and exploring the epigenetic mechanisms that underlie aging and disease progression. Zhejun Kuang, Kangwei Geng, Han Wang 0028 |
BIBM | 5 |
| 2025 | CoLA-DTA: Cross-Modal Latent Alignment for Generalization-Enhanced Drug-Target Affinity PredictionabstractDrug discovery, a cornerstone of modern biomedical research, is critical for developing targeted therapies against complex diseases. However, many existing computational approaches for DTA prediction suffer from limited accuracy and poor generalization across diverse target families and chemical spaces. To address this, we propose CoLA-DTA, a cross-modal alignment framework designed for generalizationenhanced drug-target affinity prediction. CoLA-DTA integrates multiple modalities of protein and drug representations, including sequence embeddings, 3D structural graphs, and molecular graphs. Specifically, pretrained language models and graph neural networks extract semantic, spatial, and topological features from protein sequences, drug SMILES, and their respective graph representations. A Gated Fusion module effectively combines multi-source features, followed by a bidirectional cross-attention mechanism that captures finegrained intermolecular interactions. Experimental results on benchmark datasets demonstrate that CoLA-DTA outperforms state-of-the-art models in both accuracy and generalization, particularly under cold-start settings. A case study on protooncogene tyrosine-protein kinase ROS(PTKR) further demonstrated CoLA-DTA's potential for drug discovery. CoLADTA's source code, pre-trained models, and data preprocessing code are available at https://github.com/NENUBioCompute/CoLA-DTA.git. Xike Ouyang, Han Wang 0028 |
BIBM | 5 |
| 2025 | Breast Cancer Diagnosis via Small Raman Spectroscopy Samples Using a Transformer-Based ModelabstractBreast cancer remains a major global health challenge, while more efficient and low-cost diagnosis technologies are expected to explore, while Raman spectroscopy has been employed for breast cancer detection, current analytical methods relying on characteristic peak intensity or position extraction fail to capture comprehensive patterns in the original spectral data. In this field, deep learning approaches have been explored, with one-dimensional convolutional neural networks (1D-CNNs) achieving the highest accuracy. However, existing deep learning methods face challenges in effectively identifying the global characteristics of Raman spectra. Therefore, this study proposes a lightweight Transformer architecture designed to effectively capture global spectral features while preserving local characteristics, thereby enhancing the accuracy of breast cancer diagnosis. Using a four-head attention mechanism for spectral analysis, the model employs Raman spectroscopy data as input to develop a binary classification model to distinguish normal from cancerous cytopathological states. This study provides a promising technical framework for the integration of Raman spectroscopy with deep learning, with significant potential for clinical application. Mingyue Ma, Yue Du, Han Wang 0028 |
BIBM | 7 |
| 2025 | Improving generalizability of drug-target binding prediction by pre-trained multi-view molecular representationsabstractMOTIVATION: Most drugs start on their journey inside the body by binding the right target proteins. This is the reason that numerous efforts have been devoted to predicting the drug-target binding during drug development. However, the inherent diversity among molecular properties, coupled with limited training data availability, poses challenges to the accuracy and generalizability of these methods beyond their training domain. RESULTS: In this work, we proposed a neural networks construction for high accurate and generalizable drug-target binding prediction, named Pre-trained Multi-view Molecular Representations (PMMR). The method uses pre-trained models to transfer representations of target proteins and drugs to the domain of drug-target binding prediction, mitigating the issue of poor generalizability stemming from limited data. Then, two typical representations of drug molecules, Graphs and SMILES strings, are learned respectively by a Graph Neural Network and a Transformer to achieve complementarity between local and global features. PMMR was evaluated on drug-target affinity and interaction benchmark datasets, and it derived preponderant performance contrast to peer methods, especially generalizability in cold-start scenarios. Furthermore, our state-of-the-art method was indicated to have the potential for drug discovery by a case study of cyclin-dependent kinase 2. AVAILABILITY AND IMPLEMENTATION: https://github.com/NENUBioCompute/PMMR. Xike Ouyang, Yannuo Feng, Han Wang 0028 |
Bioinform. | 6 |
| 2024 | Multimodal Medical Image Feature Representation and Fusion for AD Early DiagnosisabstractAlzheimer's disease (AD) is progressive and gets worse with time. Early detection is an effective treatment for the disease, which can timely implement effective interventions. The multi-modal medical image contains more comprehensive and abundant information about the disease, which can reflect the different aspects of AD and reduce the potential bias in single-modal diagnosis. However, an effective and reliable solution is still lacking, which can integrate multimodal neuroimaging effectively for the precision diagnosis and treatment of AD. In this work, we proposed a feature representation and fusion model of MRI and PET scans based on a multiscale convolution network and cross-attention for AD early diagnosis, named camAD. Compared with the existing models, camAD performs better with fewer pre-processing steps. Using the Alzheimer’s disease neuroimaging initiative (ADNI) datasets, we demonstrated that integrating multi-modality data outperforms single-modality models in accuracy, specificity, sensitivity, AUC, and F1 scores. Our models exhibited significantly good accuracy and generality in four classification tasks, which provided a promising way to understand the underlying mechanisms of the disorder's progression. Xueliang Bai, Lanxin Xu, Han Wang 0028, Guixia Liu |
BIBM | 5 |
| 2024 | DeepCRC: Precision Colorectal Cancer Diagnosis Using Deep Learning and Metagenomic BiomarkersabstractColorectal cancer (CRC) is one of the leading causes of cancer-related deaths globally, underscoring the critical need for early detection. Microbial biomarker detection offers significant advantages in revealing the physiological and pathological mechanisms of CRC, as well as in early screening. It provides not only deep insights into disease progression but also enables the possibility of non-invasive early diagnosis. The utilization of metagenomic shotgun sequencing (MGS) technology and the bioBakery workflow to analyze gut microbiome samples from six independent cohorts (n = 1291), including 645 CRC patients and 646 healthy controls (HC). Metagenomic analysis revealed a significant reduction in microbial diversity and identified 117 microbial species with altered abundance in CRC patients (p < 0.05). The DeepCRC model, incorporating key biomarkers and clinical metadata, achieved 97.70% accuracy and an AUC of 0.982 on an independent test set. SHAP analysis indicated that microbial species such as Ruminococcaceae unclassified SGB15260 and Clostridium leptum impact CRC significantly. These bacteria are associated with CRC through mechanisms involving inflammation and gut barrier dysfunction. The superior performance of the DeepCRC diagnostic model is expected to advance the development of non-invasive CRC screening methods. Source code is available at https://github.com/NENUBioCompute/DeepCRC. Han Wang 0028, Yannuo Feng, Kexin Shen, Chang Lu 0015, Zhongshi Xie |
BIBM | 1 |
| 2024 | PocketLG: Protein Binding Pocket Identification Using Protein Language Model and Graph TransformerabstractProtein pocket is a special region on the surface of a protein that commonly interacts with other molecules, especially small molecule compounds. Accurately identifying and understanding the structure character of protein pockets, that can accelerate the development of new drugs, as well improve existing drugs to reveal disease mechanisms. In this study, we propose a deep learning method named PocketLG, to identify the potential binding pockets integrating a Protein Language model and Graph Transformer. The method constructs the pocket identification task as a regression problem, which can more accurately identify the real pocket boundaries. In addition, We propose a data sampling strategy and construct a pipeline to support the data sampling effort. This strategy has the benefit of augmenting the data, allowing the model to better learn the intrinsic characteristics of the pocket structure, even for pockets with unknown boundaries This method can sample the unknown pocket by dividing different sample sizes and then determine the boundary of the unknown pocket by the regression value, it can be a direct help for downstream tasks. Experimental results show that our deep learning model can identify binding pockets on proteins with 91 percent success rate. This work provides a new technical route for protein binding pocket prediction research, which can greatly contribute to the development of the pharmaceutical industry. PocketLG source code, pre-trained models and data preprocessing code are available at https://github.com/NENUBioCompute/PocketLG. Heng Chang, Zehua Sun, Han Wang 0028 |
BIBM | 6 |
| 2023 | Probing Transmembrane Proteins Binding Domain via Multi-level Molecule LearningabstractThe study of transmembrane proteins (TMPs) and their binding activities holds significant importance in the pharmaceutical industry. Due to their physicochemical properties, known binding information regarding TMPs remains comparatively scarce and prevented researchers to dig information from known samples. However, research into general binding structure basis can circumvent this barrier and provide more mechanism insights, which we previously demonstrated its existence and named it as TMPs binding Domain. In this study, we try to discover the TMPs binding domain more precisely. Through atomic-level heterogenous graph convolutions, we significantly improved the classification performance of binding domains. This lays the algorithmic groundwork for utilizing binding domains in the study of TMPs binding activities and further boost the drug target research or new drug development. Yihang Bao, Yuanzhao Guo, Guan Ning Lin, Zehua Sun, Han Wang 0028 |
BIBM | 6 |
| 2023 | MSCAP: DNA Methylation Age Predictor based on Multiscale Convolutional Neural NetworkabstractDNA methylation can reflect age-related issues in individuals, and age prediction tools developed based on DNA methylation are called Epigenetic clocks. Epigenetic clocks also serve as potent tools for enhancing researchers’ understanding of the aging process and aim to contribute to the advancement of aging research. The current prevailing approach involves creating clocks using Elastic Net regression modeling, wherein high-age-correlation CpG sites were selected as features in advance. However, pre-selecting CpG sites in advance may inadvertently overlook interactions between distant CpG sites. Deep learning has significant advantages in dealing with vast feature sets due to the tremendous opportunities that the development of deep learning presents for research. Deep learning architectures have demonstrated greater efficiency in acquiring and representing large numbers of intricate features compared to traditional machine learning models. In this study, we proposed a deep learning model based on a multiscale convolutional neural network (MSCNN) for pan-tissue age prediction, called MSCAP. First, for the model to efficiently predict datasets from different sequencing platforms, we chose a common intersection of sites from the Illumina 27K and Illumina 450K platforms (25,789 CpG sites), then trained and tested the model on 93 datasets. Different tissues of the human body are used in the datasets for training and testing. And independently tested on 7 independent datasets against an Epigenetic clock developed by Horvath in 2013 with high prediction accuracy. The results showed that MSCAP could effectively extract key features from a vast number of potential features without advanced feature selection. In addition, MSCAP outperformed the Horvath clock on all seven independent test datasets. This research helped boost the development of epigenetic clock methods and reinforced the value of deep learning in computational biology. Han Wang 0028, Ruirui Cai, Xizeng Zong, Zhiquan He |
BIBM | 1 |
| 2023 | Predicting Protein-Ligand Binding Affinity with Multi-Scale Structural FeaturesabstractPredicting protein-ligand binding affinity is important in areas such as drug discovery, gene regulation and signal transduction. The DTA(Drug-Target Affinity) method based on protein structure can not only effectively compensates for the lack of binding information, but also more in line with real biological processes. Although the structure-based DTA methods have achieved good performance, the existing methods still have the problem of only considering single-scale structural features and ignoring multi-scale structural features. In order to solve this problem, we propose the MSSDTA (Multi-Scale Structural Representation Drug-Target Affinity Prediction), which extracts multi-scale protein features by integrating the surface node features and structural node features of proteins. At the same time, the drug representation network is used to fuse the 2D molecular structure characteristics and chemical characteristics of the drug to effectively distinguish the drug molecules with similar planar structures. Finally, the affinity prediction network is used to generate protein-ligand binding affinity scores. We verify the performance of this model on the PDBbind v.2019 dataset. The experimental results show that the proposed method achieves excellent performance. Han Wang 0028, Jingtong Zhao, Shengkun Wang, Zhiquan He, Xike Ouyang |
BIBM | 1 |
| 2022 | Predicting Compound-Protein Interaction by Deepening the Systemic Background via Molecular Network Feature EmbeddingabstractIdentifying compound-protein interactions (CPI) is crucial for drug screening, drug repurposing, and combination therapy studies. The performance of CPI prediction depends heavily on the features extracted from compounds and target proteins. The existing prediction methods use different feature combinations, but both molecular-based and network-based models have the problem of incomplete feature representations. Therefore, completely integrating the relevant features of CPI would be an effective way to solve the existing problem. This study proposed a novel model named MCPI, which integrated the PPI (protein-protein interaction) network, CCI (compound-compound interaction) network, and structure features of CPI to improve prediction performance. We compared our model with other existing methods for predicting CPI on public datasets. The experimental results showed that MCPI outperformed the peer methods. In addition, in response to the SARS-CoV-2 pandemic, we applied the model to search for potential inhibitors among FDA-approved drugs and validated the prediction results through the literature. This work may also provide potential guidance for drug development. Han Wang 0028, Hangxu Zhu, Ming Liu 0024, Dong Xu 0002 |
BIBM | 1 |
| 2022 | An Improved Topology Prediction of Alpha-Helical Transmembrane Protein Based on Deep Multi-Scale Convolutional Neural NetworkabstractAlpha-helical proteins ( αTMPs) are essential in various biological processes. Despite their tertiary structures are crucial for revealing complex functions, experimental structure determination remains challenging and costly. In the past decades, various sequence-based topology prediction methods have been developed to bridge the gap between the sequences and structures by characterizing the structural features, but significant improvements are still required. Deep learning brings a great opportunity for its powerful representation learning capability from limited original data. In this work, we improved our αTMP topology prediction method DMCTOP using deep learning, which composed of two deep convolutional blocks to simultaneously extract local and global contextual features. Consequently, the inputs were simplified to reflect the original features of the sequence, including a protein sequence feature and an evolutionary conservation feature. DMCTOP can efficiently and accurately identify all topological types and the N-terminal orientation for an αTMP sequence. To validate the effectiveness of our method, we benchmarked DMCTOP against 13 peer methods according to the whole sequence, the transmembrane segment and the traditional criterion in testing experiments. All the results reveal that our method achieved the highest prediction accuracy and outperformed all the previous methods. The method is available at https://icdtools.nenu.edu.cn/dmctop. Jiawen Yu, Zhe Liu 0030, Han Wang 0028, Zhiqiang Ma 0003, Dong Xu 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2021 | Discover the Binding Domain of Transmembrane Proteins Based on Structural UniversalityabstractTransmembrane proteins (TMPs) serve as drug targets for more than half of the drugs currently available in the market. However, it had not been clearly explained how they realize their drug effects through multiple complex molecules bindings actions. Research into TMPs bindings and corresponding structural basis will provide key information for drug research and new drug development. In this study, we defined the binding domain of TMPs according to the binding region investigation of multiple conjugate types. A 3D deep learning model was architected to discover the structural universality inside those domains. The experimental results proved such binding domains existing on the surface of TMPs, and they are structural specific distinguishing to the surface regions without any binding activities. This work provides a new theoretical basis for TMPs binding research and can greatly boost the development of the drug industry. Yihang Bao, Fei He 0003, Weixi Wang, Han Wang 0028, Minglong Dong |
BIBM | 4 |
| 2020 | SeqTMPPI: Sequence-Based Transmembrane Protein Interaction PredictionabstractTransmembrane proteins (TMPs) play important roles in diverse cellular processes, they are the most common drug targets. The interaction between TMPs and nontransmembrane proteins (nonTMPs) is an important part to form the pathway crossing biomembranes, which is directly related to TMPs associated signal transduction, substance transport, and drug metabolism. However, it is difficult for biologists to identify interactions between TMPs and nonTMPs by wet-lab experiments since TMPs are embedded in the phospholipid bilayer while most nonTMPs are soluble proteins. Predicting protein-protein interactions (PPIs) has always been a hot topic in bioinformatics. However, TMP related PPIs occupy a small proportion of those researches, and those methods designed for soluble protein PPI may have limitations to predict the interaction of TMPs and non-TMPs due to the difference between the aqueous and lipid senvironments. There still lack of deep-learning-based predictors using primary sequence information on TMP-nonTMP interactions. In this work, we constructed a benchmark dataset for TMP-nonTMP interactions, adopted the one-hot vector to encode protein sequence pairs and built a Convolutional Neural Network (CNN) based model, SeqTMPPI, predicting the TMP-nonTMP interactions. The experimental results indicated that our method achieved a good performance on an independent testing where the Matthews Correlation Coefficient (MCC) achieved at 0.7066. The predicted interactions were further analyzed in the scope of distribution of the protein family and species. Materials related are available in https://github.com/JulseJiang/SeqTMPPI. Han Wang 0028, Jiuhong Jiang, Qiufen Chen, Chang Lu 0015, Zhiqiang Ma 0003 |
BIBM | 1 |
| 2019 | DMCTOP: Topology Prediction of Alpha-Helical Transmembrane Protein Based on Deep Multi-Scale Convolutional Neural NetworkabstractAlpha-helical transmembrane proteins ($\alpha \text{TMPs}$) belong to an important category of integral membranes. Their structures are highly valuable in relevant research, but costly to solve experimentally. Sequence-based topology prediction provides a practical computational approach to characterize the structure features. Although much progress had been made in the past decade, there is significant room for improvement in predicting the topology structure. Deep learning brings a great opportunity for its capability of mining new features from data. In this work, we propose a novel$\alpha \text{TMP}$topology prediction method DMCTOP using a Deep Multi-Scale Convolutional Neural Network (DMCNN), which composes of two deep convolutional blocks to extract local and global contextual features. Consequently, the inputs of DMCTOP is simplified to a protein sequence feature and an evolutionary conservation feature. DMCTOP can efficiently and accurately identify all topological types and the N-terminal orientation for an$\alpha \text{TMP}$sequence. In the testing experiments, the prediction accuracy was calculated according to the whole sequence, the transmembrane segments and the traditional criterion. Our state-of-the-art method achieved the highest prediction accuracy compared to all the previous methods. The standalone tool is available at https://github.com/NENUBioCompute/DMCTOP. Han Wang 0028, Jiawen Yu, Dong Xu 0002 |
BIBM | 1 |