Han Zhang 0017

dblp:26/4189-17 · DBLP profile ↗
← Back
36ranked-venue papers
3as first author
26since 2021 · last 2026
0000-0001-8498-3451ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 26 · 2 first-author · 16 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MSRA-Diff: A diffusion model with multi-scale region-aware alignment for medical image fusion
Yu Cheng 0027, Xiongwen Quan, Mingjing Han, Han Zhang 0017
Pattern Recognit.4
2026 HopGAT: A multi-hop graph attention network with heterophily and degree awareness
Han Zhang 0017, Mingjing Han
Pattern Recognit.1
2026 SMFusion: Semantic-Preserving Fusion of Multimodal Medical Images for Enhanced Clinical Diagnosis
abstract
Multimodal medical image fusion plays a crucial role in medical diagnosis by integrating complementary information from different modalities to enhance image readability and clinical applicability. However, existing methods mainly follow computer vision standards for feature extraction and fusion strategy formulation, overlooking the rich semantic information inherent in medical images. To address this limitation, we propose a novel semantic-guided medical image fusion approach that, for the first time, incorporates medical prior knowledge into the fusion process. Specifically, we construct a publicly available multimodal medical image-text dataset, upon which text descriptions generated by BiomedGPT are encoded and semantically aligned with image features in a high-dimensional space via a semantic interaction alignment module. During this process, a cross attention based linear transformation automatically maps the relationship between textual and visual features to facilitate comprehensive learning. The aligned features are then embedded into a text-injection module for further feature-level fusion. Unlike traditional methods, we further generate diagnostic reports from the fused images to assess the preservation of medical information. Additionally, we design a medical semantic loss function to enhance the retention of textual cues from the source images. Experimental results on test datasets demonstrate that the proposed method achieves superior performance in both qualitative and quantitative evaluations while preserving more critical medical information.
Haozhe Xiang, Han Zhang 0017, Yu Cheng 0027, Xiongwen Quan, Wanwan Huang
IEEE J. Biomed. Health Informatics2
2025 scStarCorrect: A StarGAN-Based Adversarial Model for Reference-Guided Batch Correction in Single-Cell RNA-Seq Data
abstract
Batch effects remain a major challenge in single-cell RNA-seq data, which often mask true biological variation across datasets. Most existing correction methods rely on pairwise mutual alignment or statistical integration, which can be inappropriate when batch sizes are imbalanced or when a consistent reference domain is desired. To address this problem, we propose scStarCorrect, a StarGAN-based adversarial model for reference-guided batch correction in single-cell data. The model includes a shared variational autoencoder conditioned on batch labels, a domain-aware discriminator that enforces alignment to a fixed reference batch. Unlike traditional mutual alignment methods that suffer from imbalance, scStarCorrect performs one-way batch correction by projecting all batches into the more accurate reference domain, enabling consistent and interpretable integration. We jointly optimize reconstruction loss, adversarial loss, domain classification loss, and variational regularization to learn batch-invariant embeddings. Experimental results on real single-cell datasets demonstrate that scStarCorrect effectively removes batch-specific variation while preserving cell type structure, offering a robust solution for multi-batch alignment and downstream analyses.
Weizhen Gu, Han Zhang 0017, Mingjing Han
BIBM2
2025 A comprehensive graph neural network method for predicting triplet motifs in disease-drug-gene interactions
abstract
MOTIVATION: The drug-disease, gene-disease, and drug-gene relationships, as high-frequency edge types, describe complex biological processes within the biomedical knowledge graph. The structural patterns formed by these three edges are the graph motifs of (disease, drug, gene) triplets. Among them, the triangle is a steady and important motif structure in the network, and other various motifs different from the triangle also indicate rich semantic relationships. However, existing methods only focus on the triangle representation learning for classification, and fail to further discriminate various motifs of triplets. A comprehensive method is needed to predict the various motifs within triplets, which will uncover new pharmacological mechanisms and improve our understanding of disease-gene-drug interactions. Identifying complex motif structures within triplets can also help us to study the structural properties of triangles. RESULTS: We consider the seven typical motifs within the triplets and propose a novel graph contrastive learning-based method for triplet motif prediction (TriMoGCL). TriMoGCL utilizes a graph convolutional encoder to extract node features from the global network topology. Next, node pooling and edge pooling extract context information as the triplet features from global and local views. To avoid the redundant context information and motif imbalance problem caused by dense edges, we use node and class-prototype contrastive learning to denoise triplet features and enhance discrimination between motifs. The experiments on two different-scale knowledge graphs demonstrate the effectiveness and reliability of TriMoGCL in identifying various motif types. In addition, our model reveals new pharmacological mechanisms, providing a comprehensive analysis of triplet motifs. AVAILABILITY AND IMPLEMENTATION: Codes and datasets are available at https://github.com/zhanglabNKU/TriMoGCL and https://doi.org/10.5281/zenodo.14633572.
Chuanze Kang, Zonghuan Liu, Han Zhang 0017
Bioinform.3
2024 TMGE: a multi-view graph embedding model for prediction task of genes associated with neurodegenerative diseases
abstract
Predicting genes associated with diseases is of great significance for the early diagnosis and efficient treatment of neurodegenerative diseases. However, with the increasing complexity of omics data, it is difficult to integrate many aspects of relationships and effectively mine the genes associated with neurodegenerative diseases. To address this problem, we propose a novel multi-view graph embedding model, named TMGE, which can effectively integrate multiple networks to learn gene representations for prediction task of genes associated with neurodegenerative diseases. First, we design view-specific and view-shared encoders to extract view-specific content and capture consistent information across views; this enables the separation and preservation of diverse information within multi-view data for more accurate representations. Secondly, we introduce cross-linked encoders to enhance the integration of multi-view information, enabling the imputation of missing data across views by capturing cross-view dependencies. Furthermore, we employ adversarial learning to reduce the distribution discrepancy between diverse views at the global level, and then optimize the mean square error of local pairs with the help of the cross-linked encoder structure to achieve local alignment. We conduct experiments to verify the effectiveness of TMGE for predicting genes associated with neurodegenerative diseases using two disease datasets of Alzheimer and Huntington.
Yuanxiang Jiang, Mingjing Han, Yanbin Yin, Han Zhang 0017
BIBM4
2024 BERMAD: batch effect removal for single-cell RNA-seq data using a multi-layer adaptation autoencoder with dual-channel framework
abstract
MOTIVATION: Removal of batch effect between multiple datasets from different experimental platforms has become an urgent problem, since single-cell RNA sequencing (scRNA-seq) techniques developed rapidly. Although there have been some methods for this problem, most of them still face the challenge of under-correction or over-correction. Specifically, handling batch effect in highly nonlinear scRNA-seq data requires a more powerful model to address under-correction. In the meantime, some previous methods focus too much on removing difference between batches, which may disturb the biological signal heterogeneity of datasets generated from different experiments, thereby leading to over-correction. RESULTS: In this article, we propose a novel multi-layer adaptation autoencoder with dual-channel framework to address the under-correction and over-correction problems in batch effect removal, which is called BERMAD and can achieve better results of scRNA-seq data integration and joint analysis. First, we design a multi-layer adaptation architecture to model distribution difference between batches from different feature granularities. The distribution matching on various layers of autoencoder with different feature dimensions can result in more accurate batch correction outcome. Second, we propose a dual-channel framework, where the deep autoencoder processing each single dataset is independently trained. Hence, the heterogeneous information that is not shared between different batches can be retained more completely, which can alleviate over-correction. Comprehensive experiments on multiple scRNA-seq datasets demonstrate the effectiveness and superiority of our method over the state-of-the-art methods. AVAILABILITY AND IMPLEMENTATION: The code implemented in Python and the data used for experiments have been released on GitHub (https://github.com/zhanglabNKU/BERMAD) and Zenodo (https://zenodo.org/records/10695073) with detailed instructions.
Xiangxin Zhan, Yanbin Yin, Han Zhang 0017
Bioinform.3
2024 KAGNN: Graph neural network with kernel alignment for heterogeneous graph learning
Mingjing Han, Han Zhang 0017
Knowl. Based Syst.2
2024 A Dual-Modality Complex-Valued Fusion Method for Predicting Side Effects of Drug-Drug Interactions Based on Graph Neural Network
abstract
Predicting potential side effects of drug-drug interactions (DDIs), which is a major concern in clinical treatment, can increase therapeutic efficacy. In recent studies, how to use the multi-modal drug features is important for DDI prediction. Thus, it remains a challenge to explore an efficient computational method to achieve the feature fusion cross- and intra-modality. In this paper, we propose a dual-modality complex-valued fusion method (DMCF-DDI) for predicting the side effects of DDIs, using the form and properties of complex-vector to enhance the representations of DDIs. Firstly, DMCF-DDI applies two Graph Convolutional Network (GCN) encoders to learn molecular structure and topological features from fingerprint and knowledge graphs, respectively. Secondly, an asymmetric skip connection (ASC) uses distinct semantic-level features to construct the complex-valued drug pair representations (DPRs). Then, the complex-vector multiplication is used as a fusion operator to obtain the fine-grained DPRs. Finally, we calculate the prediction probability of DDIs by Hermitian inner product in the complex space. Compared with other methods, DMCF-DDI achieves superior performance in all situations using a fusion operator with the lowest parameter numbers. For the case study, we select six diseases and common side effects in clinical treatment to verify identification ability of our model. We also prove the advantage of ASC and complex-valued fusion can achieve to align the cross-modal fused positive DPRs through a comprehensive analysis on the phase-modulus distribution histogram of DPRs. In the end, we explain the reason for alignment based on the similarity of features and node neighbors.
Chuanze Kang, Han Zhang 0017, Yanbin Yin
IEEE J. Biomed. Health Informatics2
2023 Hierarchical Semantic Augmentation Graph Neural Network for Drug-Disease Association Prediction
abstract
As an essential step in drug intervention discovery, predicting the drug-disease associations (DDAs) explores the potential therapeutic associations in given dugs and diseases. Since the various links in drugs and diseases contain high-order relations and complex therapeutic semantics, Graph Neural Networks (GNNs) have been introduced to DDA predictions and achieved great success. However, most previous approaches require the nodes of given drugs and diseases to have smooth attributes, which is difficult to meet in practical applications. Besides, GNN-based models suffer from the problem of semantic confusion for DDA prediction in heterogeneous graph. These challenges limit the model validity to discover therapeutic semantics in drug-disease networks. To address these challenges in DDA, we propose a novel graph neural network model called HSAGNN to augment node semantics hierarchically with three key steps by applying semantic-guided idea of SGNN method, including topological embedding learning, attribute completion, and semantic-guided aggregation. HSAGNN first learns the topological embedding and adopts the learned topological relationships to complete missing attributes with attention mechanism, which allows the node to contain richer information for neighbor aggregation. Then, the model aggregates the neighbor information with semantic-guided aggregation in both node and semantic levels. Here, HSAGNN injects the learned common knowledge as jumping knowledge to alleviate the semantic confusion. We evaluate the model in DDA tasks with various baselines and explore the model validity with extensive studies. The experimental results show that HSAGNN can discover the potential therapeutic associations by augmented semantics.
Mingjing Han, Yanbin Yin, Han Zhang 0017
BIBM3
2023 Unsupervised Feature Selection by Fusing Spectral Clustering and Locality Preserving Projection
abstract
Due to feature redundancy in high dimensional data, the unsupervised feature selection methods for dimension reduction have attracted considerable attention. The current feature selection frameworks consider the global information of data, but ignore mutual screening of global and local information, and there have been no important breakthroughs on this approach research recently. We propose a novel feature selection method based on iterative optimization between the pseudo label matrix from spectral clustering and the local projection information (SNUFS), and then prove the convergence of the method. The pseudo label matrix and the local projection are designed in objective function for mutual screening and guiding the regularization feature selection by iterative approach. Our method selects features most relevant to the pseudo label and preserves the local structure of original data from feature selection matrix, where alternate iteration of different optimization items including pseudo label matrix achieve mutual screening. For this method, we give the objective function, iterative optimization fusion approach and convergence analysis in detail. Furthermore, we use K-Nearest Neighbor (KNN) and K-means to implement locality preserving projection and obtain two specific algorithms. Experiments on four real-world datasets in different fields demonstrate that our algorithms can effectively improve the accuracy of feature selection. In particular, our algorithm of KNN implementation is more effective and outperforms other major algorithms.
Xiongwen Quan, Mingjing Han, Xia Guo, Han Zhang 0017, Yanbin Yin
BIBM5
2023 scMHNN: a novel hypergraph neural network for integrative analysis of single-cell epigenomic, transcriptomic and proteomic data
abstract
Technological advances have now made it possible to simultaneously profile the changes of epigenomic, transcriptomic and proteomic at the single cell level, allowing a more unified view of cellular phenotypes and heterogeneities. However, current computational tools for single-cell multi-omics data integration are mainly tailored for bi-modality data, so new tools are urgently needed to integrate tri-modality data with complex associations. To this end, we develop scMHNN to integrate single-cell multi-omics data based on hypergraph neural network. After modeling the complex data associations among various modalities, scMHNN performs message passing process on the multi-omics hypergraph, which can capture the high-order data relationships and integrate the multiple heterogeneous features. Followingly, scMHNN learns discriminative cell representation via a dual-contrastive loss in self-supervised manner. Based on the pretrained hypergraph encoder, we further introduce the pre-training and fine-tuning paradigm, which allows more accurate cell-type annotation with only a small number of labeled cells as reference. Benchmarking results on real and simulated single-cell tri-modality datasets indicate that scMHNN outperforms other competing methods on both cell clustering and cell-type annotation tasks. In addition, we also demonstrate scMHNN facilitates various downstream tasks, such as cell marker detection and enrichment analysis.
Wei Li 0184, Bin Xiang, Fan Yang 0081, Yu Rong 0001, Yanbin Yin, Jianhua Yao 0001, Han Zhang 0017
Briefings Bioinform.7
2023 Semantic-guided graph neural network for heterogeneous graph embedding
Mingjing Han, Han Zhang 0017, Wei Li 0184, Yanbin Yin
Expert Syst. Appl.2
2023 Graph representation learning via redundancy reduction
Mengyao He, Han Zhang 0017, Chuanze Kang, Wei Li 0184, Mingjing Han
Neurocomputing3
2023 Graph pooling via Dual-view Multi-level Infomax
Han Zhang 0017, Mengyao He, Wei Li 0184, Chuanze Kang, Mingjing Han
Knowl. Based Syst.2
2022 Mutual Information Estimation-Based Disentangled Representation Network for Medical Image Fusion
abstract
Deep learning-based method for medical image fusion has become a hot topic in recent years. However, they ignore the expression of the most important features in image fusion and only extract the general features for medical image fusion, which will restrict the expression of unique information on the fusion image. To address this restriction, we propose a novel disentangled representation network for medical image fusion with mutual information estimation, which extract the disentangled features of medical image fusion, i.e., the shared and exclusive features between multi-model medical images. In our method, we use the cross mutual information method to obtain the shared features of each modality pair, which enforce the fusion network to achieve the maximum of mutual information estimation for multi-modal medical images. The exclusive features are extracted by the adversarial objective method and it constrains the fusion network with the optimization to the minimum of mutual information estimation between shared and exclusive features. These disentangled features effectively take the interpretative advantages and make the fusion image retaining more details from source images as well as improving the visual quality of fusion image. Our method has achieved better results than several state-of-the-art methods. Both qualitative and quantitative experiments have proved the superiority of our method.
Wanwan Huang, Han Zhang 0017, Yanbin Yin
BIBM2
2022 LR-GNN: a graph neural network based on link representation for predicting molecular associations
abstract
In biomedical networks, molecular associations are important to understand biological processes and functions. Many computational methods, such as link prediction methods based on graph neural networks (GNNs), have been successfully applied in discovering molecular relationships with biological significance. However, it remains a challenge to explore a method that relies on representation learning of links for accurately predicting molecular associations. In this paper, we present a novel GNN based on link representation (LR-GNN) to identify potential molecular associations. LR-GNN applies a graph convolutional network (GCN)-encoder to obtain node embedding. To represent associations between molecules, we design a propagation rule that captures the node embedding of each GCN-encoder layer to construct the LR. Furthermore, the LRs of all layers are fused in output by a designed layer-wise fusing rule, which enables LR-GNN to output more accurate results. Experiments on four biomedical network data, including lncRNA-disease association, miRNA-disease association, protein-protein interaction and drug-drug interaction, show that LR-GNN outperforms state-of-the-art methods and achieves robust performance. Case studies are also presented on two datasets to verify the ability to predict unknown associations. Finally, we validate the effectiveness of the LR by visualization.
Chuanze Kang, Han Zhang 0017, Shenwei Huang, Yanbin Yin
Briefings Bioinform.2
2022 MGEGFP: a multi-view graph embedding method for gene function prediction based on adaptive estimation with GCN
abstract
In recent years, a number of computational approaches have been proposed to effectively integrate multiple heterogeneous biological networks, and have shown impressive performance for inferring gene function. However, the previous methods do not fully represent the critical neighborhood relationship between genes during the feature learning process. Furthermore, it is difficult to accurately estimate the contributions of different views for multi-view integration. In this paper, we propose MGEGFP, a multi-view graph embedding method based on adaptive estimation with Graph Convolutional Network (GCN), to learn high-quality gene representations among multiple interaction networks for function prediction. First, we design a dual-channel GCN encoder to disentangle the view-specific information and the consensus pattern across diverse networks. By the aid of disentangled representations, we develop a multi-gate module to adaptively estimate the contributions of different views during each reconstruction process and make full use of the multiplexity advantages, where a diversity preservation constraint is designed to prevent the over-fitting problem. To validate the effectiveness of our model, we conduct experiments on networks from the STRING database for both yeast and human datasets, and compare the performance with seven state-of-the-art methods in five evaluation metrics. Moreover, the ablation study manifests the important contribution of the designed dual-channel encoder, multi-gate module and the diversity preservation constraint in MGEGFP. The experimental results confirm the superiority of our proposed method and suggest that MGEGFP can be a useful tool for gene function prediction.
Wei Li 0184, Han Zhang 0017, Minghe Li, Mingjing Han, Yanbin Yin
Briefings Bioinform.2
2022 HDMC: a novel deep learning-based framework for removing batch effects in single-cell RNA-seq data
abstract
MOTIVATION: With the development of single-cell RNA sequencing (scRNA-seq) techniques, increasingly more large-scale gene expression datasets become available. However, to analyze datasets produced by different experiments, batch effects among different datasets must be considered. Although several methods have been recently published to remove batch effects in scRNA-seq data, two problems remain to be challenging and not completely solved: (i) how to reduce the distribution differences of different batches more accurately; and (ii) how to align samples from different batches to recover the cell type clusters. RESULTS: We proposed a novel deep-learning approach, which is a hierarchical distribution-matching framework assisted with contrastive learning to address these two problems. Firstly, we design a hierarchical framework for distribution matching based on a deep autoencoder. This framework employs an adversarial training strategy to match the global distribution of different batches. This provides an improved foundation to further match the local distributions with a maximum mean discrepancy-based loss. For local matching, we divide cells in each batch into clusters and develop a contrastive learning mechanism to simultaneously align similar cluster pairs and keep noisy pairs apart from each other. This allows to obtain clusters with all cells of the same type (true positives), and avoid clusters with cells of different type (false positives). We demonstrate the effectiveness of our method on both simulated and real datasets. Results show that our new method significantly outperforms the state-of-the-art methods and has the ability to prevent overcorrection. AVAILABILITY AND IMPLEMENTATION: The python code to generate results and figures in this article is available at https://github.com/zhanglabNKU/HDMC, the data underlying this article is also available at this github repository. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiao Wang 0099, Jia Wang 0051, Han Zhang 0017, Shenwei Huang, Yanbin Yin
Bioinform.3
2022 Domain-aware multi-modality fusion network for generalized zero-shot learning
Jia Wang 0051, Xiao Wang 0099, Han Zhang 0017
Neurocomputing3
2022 Multiple kernel learning for label relation and class imbalance in multi-label learning
Mingjing Han, Han Zhang 0017
Inf. Sci.2
2021 Predicting lncRNA-protein interactions based on graph autoencoders and collaborative training
abstract
Long non-coding RNAs(lncRNAs) play an important role in various biological processes. lncRNAs usually perform their molecular functions by interacting with proteins. Therefore, it is essential to predict potential lncRNA-protein associations for disease prevention and disease treatment. Label-propagation-based methods are widely used for predicting associations among biological entities. However, in these approaches, similarity computation and label propagation are separate procedures, which decrease the effectiveness of label propagation. Moreover, the prediction accuracy of existing models is also not ideal. In this paper, we proposed an end-to-end deep learning lncRNA-protein Interaction predictor through Graph Autoencoders and Collaborative training (LPIGAC). Different from previous studies, our model implemented two graph autoencoders on lncRNA graph and protein graph respectively, and trained these two graph autoencoders collaboratively. Graph autoencoders on lncRNA graph and protein graph are competent to reconstruct score matrix through initial association matrix, which is equivalent to propagate labels on graphs. This end-to-end framework can strengthen the robustness and precision of the label propagation procedure. Cross validations indicate that LPIGAC outperforms current lncRNA-protein associations prediction methods. Case studies demonstrate that LPIGAC is competent to detect potential lncRNA-protein associations. Source code of our paper is available at https://github.com/zhanglabNKULPIGAC.
Zhuangwei Shi, Han Zhang 0017, Yanbin Yin
BIBM3
2021 Adversarial Dual-Channel Variational Graph Autoencoder for Synthetic Lethality Prediction in Human Cancers
abstract
Synthetic Lethality (SL) is a type of vital gene interaction that can lead to various human diseases including cancers. Therefore, SL gene pair prediction can aid in the prevention and treatment of cancer. A number of computational approaches, especially Graph Neural Network (GNN) based methods, have been proposed for this link prediction problem on the graph. However, these GNN-based methods only consider embedding as deterministic vectors and do not take data distribution into account. Here we propose an Adversarial Dual-Channel Variational Graph Autoencoder based on semi-implicit variational inference for SL prediction in human cancers. We consider node embedding as a random variable that has an explicit Gaussian distribution. Then we design a dual-channel GCN encoder to inject stochasticity into the distribution parameters and allow latent embedding to exceed the Gaussian distribution. This hierarchical scheme leads to a more flexible posterior of latent embedding and enhances the model representation capacity. To further obtain a robust and stable representation, an adversarial module is devised for variance regularization. Experimental results compared with other state-of-the-art methods confirm the effectiveness of our proposed method. Moreover, we conduct a case study to demonstrate that our model can be very useful to predict novel SL pairs.
Wei Li 0184, Han Zhang 0017, Jian Liu 0040, Yanbin Yin
BIBM2
2021 Novel GAN Inversion Model with Latent Space Constraints for Face Reconstruction
Jinglong Yang, Xiongwen Quan, Han Zhang 0017
ICONIP (3)3
2021 A representation learning model based on variational inference and graph autoencoder for predicting lncRNA-disease associations
abstract
BACKGROUND: Numerous studies have demonstrated that long non-coding RNAs are related to plenty of human diseases. Therefore, it is crucial to predict potential lncRNA-disease associations for disease prognosis, diagnosis and therapy. Dozens of machine learning and deep learning algorithms have been adopted to this problem, yet it is still challenging to learn efficient low-dimensional representations from high-dimensional features of lncRNAs and diseases to predict unknown lncRNA-disease associations accurately. RESULTS: We proposed an end-to-end model, VGAELDA, which integrates variational inference and graph autoencoders for lncRNA-disease associations prediction. VGAELDA contains two kinds of graph autoencoders. Variational graph autoencoders (VGAE) infer representations from features of lncRNAs and diseases respectively, while graph autoencoders propagate labels via known lncRNA-disease associations. These two kinds of autoencoders are trained alternately by adopting variational expectation maximization algorithm. The integration of both the VGAE for graph representation learning, and the alternate training via variational inference, strengthens the capability of VGAELDA to capture efficient low-dimensional representations from high-dimensional features, and hence promotes the robustness and preciseness for predicting unknown lncRNA-disease associations. Further analysis illuminates that the designed co-training framework of lncRNA and disease for VGAELDA solves a geometric matrix completion problem for capturing efficient low-dimensional representations via a deep learning approach. CONCLUSION: Cross validations and numerical experiments illustrate that VGAELDA outperforms the current state-of-the-art methods in lncRNA-disease association prediction. Case studies indicate that VGAELDA is capable of detecting potential lncRNA-disease associations. The source code and data are available at https://github.com/zhanglabNKU/VGAELDA .
Zhuangwei Shi, Han Zhang 0017, Xiongwen Quan, Yanbin Yin
BMC Bioinform.2
2021 ATTCry: Attention-based neural network model for protein crystallization prediction
Jianzhao Gao, Zhuangwei Shi, Han Zhang 0017
Neurocomputing4
2020 Bayesian Multi-scale Convolutional Neural Network for Motif Occupancy Identification
abstract
Convolutional neural network (CNN) has been successfully used for the identification of motif occupancy. However, the CNN architecture requires varying length instead of fixed-length filters due to different motif lengths. Moreover, plain neural networks with single point estimation for weights suffer from over-fitting, which is more likely to occur as increasing parameters for multi-scale modeling.Hence, we have designed a Bayesian Multi-scale CNN. The model employs convolutional filters of different scales to extract latent features of DNA sequence, and incorporates Bayesian architecture which regards multi-scale weights as random variables. We further stack two sequential convolutional operations for mean and variance respectively, and apply Bayes by Back prop for posterior estimation of weights. Results have shown that our method not only improved the prediction performance for motif occupancy identification, but also prevented over-fitting due to the capability of Bayesian neural network. The model has also developed a measure of uncertainty estimation for model assessment.
Wei Li 0184, Han Zhang 0017, Xiongwen Quan, Jing Xu 0008, Yanbin Yin
BIBM3
2020 GDASC: a GPU parallel-based web server for detecting hidden batch factors
abstract
SUMMARY: We developed GDASC, a web version of our former DASC algorithm implemented with GPU. It provides a user-friendly web interface for detecting batch factors. Based on the good performance of DASC algorithm, it is able to give the most accurate results. For two steps of DASC, data-adaptive shrinkage and semi-non-negative matrix factorization, we designed parallelization strategies facing convex clustering solution and decomposition process. It runs more than 50 times faster than the original version on the representative RNA sequencing quality control dataset. With its accuracy and high speed, this server will be a useful tool for batch effects analysis. AVAILABILITY AND IMPLEMENTATION: http://bioinfo.nankai.edu.cn/gdasc.php. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiao Wang 0099, Haidong Yi, Jia Wang 0051, Zhandong Liu, Yanbin Yin, Han Zhang 0017
Bioinform.6
2020 eCAMI: simultaneous classification and motif identification for enzyme annotation
abstract
MOTIVATION: Carbohydrate-active enzymes (CAZymes) are extremely important to bioenergy, human gut microbiome, and plant pathogen researches and industries. Here we developed a new amino acid k-mer-based CAZyme classification, motif identification and genome annotation tool using a bipartite network algorithm. Using this tool, we classified 390 CAZyme families into thousands of subfamilies each with distinguishing k-mer peptides. These k-mers represented the characteristic motifs (in the form of a collection of conserved short peptides) of each subfamily, and thus were further used to annotate new genomes for CAZymes. This idea was also generalized to extract characteristic k-mer peptides for all the Swiss-Prot enzymes classified by the EC (enzyme commission) numbers and applied to enzyme EC prediction. RESULTS: This new tool was implemented as a Python package named eCAMI. Benchmark analysis of eCAMI against the state-of-the-art tools on CAZyme and enzyme EC datasets found that: (i) eCAMI has the best performance in terms of accuracy and memory use for CAZyme and enzyme EC classification and annotation; (ii) the k-mer-based tools (including PPR-Hotpep, CUPP and eCAMI) perform better than homology-based tools and deep-learning tools in enzyme EC prediction. Lastly, we confirmed that the k-mer-based tools have the unique ability to identify the characteristic k-mer peptides in the predicted enzymes. AVAILABILITY AND IMPLEMENTATION: https://github.com/yinlabniu/eCAMI and https://github.com/zhanglabNKU/eCAMI. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jing Xu 0008, Han Zhang 0017, Jinfang Zheng, Philippe Dovoedo, Yanbin Yin
Bioinform.2
2019 Antimicrobial peptide identification using multi-scale convolutional network
abstract
BACKGROUND: Antibiotic resistance has become an increasingly serious problem in the past decades. As an alternative choice, antimicrobial peptides (AMPs) have attracted lots of attention. To identify new AMPs, machine learning methods have been commonly used. More recently, some deep learning methods have also been applied to this problem. RESULTS: In this paper, we designed a deep learning model to identify AMP sequences. We employed the embedding layer and the multi-scale convolutional network in our model. The multi-scale convolutional network, which contains multiple convolutional layers of varying filter lengths, could utilize all latent features captured by the multiple convolutional layers. To further improve the performance, we also incorporated additional information into the designed model and proposed a fusion model. Results showed that our model outperforms the state-of-the-art models on two AMP datasets and the Antimicrobial Peptide Database (APD)3 benchmark dataset. The fusion model also outperforms the state-of-the-art model on an anti-inflammatory peptides (AIPs) dataset at the accuracy. CONCLUSIONS: Multi-scale convolutional network is a novel addition to existing deep neural network (DNN) models. The proposed DNN model and the modified fusion model outperform the state-of-the-art models for new AMP discovery. The source code and data are available at https://github.com/zhanglabNKU/APIN.
Jing Xu 0008, Yanbin Yin, Xiongwen Quan, Han Zhang 0017
BMC Bioinform.5
2019 A disease-related gene mining method based on weakly supervised learning model
abstract
BACKGROUND: Predicting disease-related genes is helpful for understanding the disease pathology and the molecular mechanisms during the disease progression. However, traditional methods are not suitable for screening genes related to the disease development, because there are some samples with weak label information in the disease dataset and a small number of genes are known disease-related genes. RESULTS: We designed a disease-related gene mining method based on the weakly supervised learning model in this paper. The method is separated into two steps. Firstly, the differentially expressed genes are screened based on the weakly supervised learning model. In the model, the strong and weak label information at different stages of the disease progression is fully utilized. The obtained differentially expressed gene set is stable and complete after the algorithm converges. Then, we screen disease-related genes in the obtained differentially expressed gene set using transductive support vector machine based on the difference kernel function. The difference kernel function can map the input space of the original Huntington's disease gene expression dataset to the difference space. The relation between the two genes can be evaluated more accurately in the difference space and the known disease-related gene information can be used effectively. CONCLUSIONS: The experimental results show that the disease-related gene mining method based on the weakly supervised learning model can effectively improve the precision of the disease-related gene prediction compared with other excellent methods.
Han Zhang 0017, Xueting Huo, Xia Guo, Xiongwen Quan
BMC Bioinform.1
2019 Flexible Non-Negative Matrix Factorization to Unravel Disease-Related Genes
abstract
Recently, non-negative matrix factorization (NMF) has been shown to perform well in the analysis of omics data. NMF assumes that the expression level of one gene is a linear additive composition of metagenes. The elements in metagene matrix represent the regulation effects and are restricted to non-negativity. However, according to the real biological meaning, there are two kinds of regulation effects, i.e., up-regulation and down-regulation. Few methods based on NMF have considered this biological meaning. Therefore, we designed a flexible non-negative matrix factorization (FNMF) algorithm by further considering the biological meaning of gene expression data. It allows negative numbers in the metagene matrix, and negative numbers represent down-regulation effects. We separated gene expression data into disease-driven gene expression and background gene expression. Subsequently, we computed disease-driven gene relative expression, and a ranked list of genes was obtained. The top ranked genes are considered to be involved in some disease-related biological processes. Experimental results on two real-world gene expression data demonstrate the feasibility and effectiveness of FNMF. Compared with conventional disease-related gene identification algorithms, FNMF has superior performance in analyzing gene expression data of diseases with complex pathology.
Han Zhang 0017, Xiongwen Quan
IEEE ACM Trans. Comput. Biol. Bioinform.2
2018 A Disease-related Gene Mining Method Based On Weakly Supervised Learning Model
Han Zhang 0017, Xueting Huo, Xia Guo, Xiongwen Quan
BIBM1
2018 Detecting hidden batch factors through data-adaptive adjustment for biological effects
abstract
Motivation: Batch effects are one of the major source of technical variations that affect the measurements in high-throughput studies such as RNA sequencing. It has been well established that batch effects can be caused by different experimental platforms, laboratory conditions, different sources of samples and personnel differences. These differences can confound the outcomes of interest and lead to spurious results. A critical input for batch correction algorithms is the knowledge of batch factors, which in many cases are unknown or inaccurate. Hence, the primary motivation of our paper is to detect hidden batch factors that can be used in standard techniques to accurately capture the relationship between gene expression and other modeled variables of interest. Results: We introduce a new algorithm based on data-adaptive shrinkage and semi-Non-negative Matrix Factorization for the detection of unknown batch effects. We test our algorithm on three different datasets: (i) Sequencing Quality Control, (ii) Topotecan RNA-Seq and (iii) Single-cell RNA sequencing (scRNA-Seq) on Glioblastoma Multiforme. We have demonstrated a superior performance in identifying hidden batch effects as compared to existing algorithms for batch detection in all three datasets. In the Topotecan study, we were able to identify a new batch factor that has been missed by the original study, leading to under-representation of differentially expressed genes. For scRNA-Seq, we demonstrated the power of our method in detecting subtle batch effects. Availability and implementation: DASC R package is available via Bioconductor or at https://github.com/zhanglabNKU/DASC. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Haidong Yi, Ayush T. Raman, Han Zhang 0017, Genevera I. Allen, Zhandong Liu
Bioinform.3
2017 pHMM-tree: phylogeny of profile hidden Markov models
abstract
Protein families are often represented by profile hidden Markov models (pHMMs). Homology between two distant protein families can be determined by comparing the pHMMs. Here we explored the idea of building a phylogeny of protein families using the distance matrix of their pHMMs. We developed a new software and web server (pHMM-tree) to allow four major types of inputs: (i) multiple pHMM files, (ii) multiple aligned protein sequence files, (iii) mixture of pHMM and aligned sequence files and (iv) unaligned protein sequences in a single file. The output will be a pHMM phylogeny of different protein families delineating their relationships. We have applied pHMM-tree to build phylogenies for CAZyme (carbohydrate active enzyme) classes and Pfam clans, which attested its usefulness in the phylogenetic representation of the evolutionary relationship among distant protein families. Availability and Implementation: This software is implemented in C/C ++ and is available at http://cys.bios.niu.edu/pHMM-Tree/source/. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Luyang Huo, Han Zhang 0017, Xueting Huo, Yasong Yang, Xueqiong Li
Bioinform.2
2017 Identify Huntington's disease associated genes based on restricted Boltzmann machine with RNA-seq data
abstract
BACKGROUND: Predicting disease-associated genes is helpful for understanding the molecular mechanisms during the disease progression. Since the pathological mechanisms of neurodegenerative diseases are very complex, traditional statistic-based methods are not suitable for identifying key genes related to the disease development. Recent studies have shown that the computational models with deep structure can learn automatically the features of biological data, which is useful for exploring the characteristics of gene expression during the disease progression. RESULTS: In this paper, we propose a deep learning approach based on the restricted Boltzmann machine to analyze the RNA-seq data of Huntington's disease, namely stacked restricted Boltzmann machine (SRBM). According to the SRBM, we also design a novel framework to screen the key genes during the Huntington's disease development. In this work, we assume that the effects of regulatory factors can be captured by the hierarchical structure and narrow hidden layers of the SRBM. First, we select disease-associated factors with different time period datasets according to the differentially activated neurons in hidden layers. Then, we select disease-associated genes according to the changes of the gene energy in SRBM at different time periods. CONCLUSIONS: The experimental results demonstrate that SRBM can detect the important information for differential analysis of time series gene expression datasets. The identification accuracy of the disease-associated genes is improved to some extent using the novel framework. Moreover, the prediction precision of disease-associated genes for top ranking genes using SRBM is effectively improved compared with that of the state of the art methods.
Han Zhang 0017, Feng Duan 0006, Xiongwen Quan
BMC Bioinform.2