VLDB 2026 Research / reviewers in the wild / expert
Quan Zou 0001
dblp:01/755-1
· DBLP profile ↗
276ranked-venue papers
12as first author
205since 2021 · last 2026
0000-0001-6406-1142ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 230 · 8 first-author · 180 since 2021Artificial intelligence and machine learning · 30 · 4 first-author · 19 since 2021Databases, data management, data science and information retrieval · 7 · 4 since 2021Software engineering, systems software and programming languages · 3Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLM-Guided Label Propagation with Hypergraph for Drug Repurposing
Yuqing Qian, Yijie Ding, Quan Zou 0001 |
ICIC (30) | 3 |
| 2026 | TAXICF: Efficient Index Construction and Accurate Metagenomic Classification with Interleaved Cuckoo Filters
Qinzhong Tian, Pinglu Zhang, Quan Zou 0001 |
ICIC (30) | 3 |
| 2026 | High-Quality Large-Scale Viral Genome Multiple Sequence Alignment with Updated HAlign4
Pinglu Zhang, Qinzhong Tian, Yixiao Zhai, Quan Zou 0001 |
ICIC (30) | 5 |
| 2026 | Quantum computing applications in drug discoveryabstractIn early drug discovery, virtual screening based on deep learning, virtual screening based on molecular docking, and molecular dynamics are three widely used computational strategies, but they always face a trade-off between throughput, search stability, and physical fidelity. This article discusses how quantum computing can be integrated into these processes under the constraints of Noisy Intermediate-Scale Quantum (NISQ). At present, the most realistic role of quantum computing is not the complete replacement of classical processes, but modular coprocessing for selected decision-sensitive subroutines. In the screening of deep learning, quantum modules are mainly inserted into selected components of the model. In predictive models, they are used to enhance representation learning or feature extraction. In generative models, they serve as priors or generators. In docking screening, quantum integration is suitable for specific substeps such as site recognition, pose search, and flexible docking. In molecular dynamics, representative examples include ground state ab initio molecular dynamics, annealer-based trajectory propagation, and excited state molecular dynamics, while most large-scale sampling is still done by classical methods. The actual problem in these scenarios is not whether the quantum module can be inserted, but whether it can provide repeatable gains related to decision-making under the constraints of actual running time and resources. Therefore, we emphasize strong classical baselines, reliable ranking and calibration, transparent resource reporting, and evaluation at downstream decision points as key criteria for assessing progress in the near term. Leyi Wei, Henry H. Y. Tong, Quan Zou 0001 |
Briefings Bioinform. | 4 |
| 2026 | MuFGPS: enhancing liquid-liquid phase separation protein prediction through multi-level features and ensemble learningabstractLiquid-liquid phase separation (LLPS) is a key mechanism driving the assembly of membrane-less organelles and is increasingly recognized for its involvement in essential cellular functions and various diseases. However, existing computational approaches largely rely on sequence-level descriptors and often fail to explicitly incorporate structural topology information, limiting their ability to capture the complex determinants of LLPS behavior. Accurate identification of LLPS-capable proteins remains challenging due to their sequence diversity and complex structural determinants. Here, we present MuFGPS (Multi-level Feature Graph-based Predictor for Phase-Separating proteins), a predictive framework integrating sequence-derived physicochemical features, Define Secondary Structure of Proteins-annotated secondary structures, and graph-based structural embeddings from AlphaFold residue contact maps via a multi-head Graph Attention Network. Class imbalance is addressed using Synthetic Minority Oversampling Technique (SMOTE), and classification is performed through a stacking ensemble of Random Forest, XGBoost, and LightGBM. Benchmarks against six representative methods demonstrate that MuFGPS achieves superior performance across all metrics, with notable gains in F1-score and matthews correlation coefficient (MCC). Ablation analyses confirm the synergistic contributions of structural features and ensemble learning to accuracy and robustness. MuFGPS offers a scalable and high-accuracy framework for proteome-wide LLPS protein prediction. Lei Xian, Quan Zou 0001, Ren Qi, Mengting Niu, Yansu Wang |
Briefings Bioinform. | 2 |
| 2026 | A deep adversarial network model for multi-task analysis of single-cell omics dataabstractSingle-cell multi-omics data reveal complex cellular states and deepen our understanding of tissue cell phenotypes and functions. However, data analysis remains challenging due to the discrete nature and high noise level of the data, as well as the lack of modality. Here, we propose scMultiNet, a multi-task deep adversarial neural network that can integrate different tasks to analyze single-cell multi-modal data. In particular, we achieve joint training of multi-modal integration and cross-modal prediction tasks by introducing a cross-modal bi-prediction module and a multi-head self-attention module. Data denoising is further enhanced by integrating an indicator matrix that constrains and precisely reconstructs the original expression values. Extensive simulations and real data experiments demonstrate that scMultiNet outperforms existing state-of-the-art methods in dimensionality reduction, visualization, clustering, batch elimination, data denoising, multi-modal integration, single-cell cross-modality translation, and in revealing cell type-specific biological insights. In addition, we demonstrate that scMultiNet can effectively transfer the complex relationships between modalities from one batch to another. In summary, scMultiNet stands as a comprehensive end-to-end framework, ideally suited for analyzing single-cell multi-omics data. Junlin Xu, Yajie Meng, Shuting Jin, Changcheng Lu, Feifei Cui, Xiangzheng Fu, Quan Zou 0001, Xiangxiang Zeng |
Briefings Bioinform. | 9 |
| 2026 | DynaTCR: dynamic hard-negative ensemble graph learning improves TCR-epitope binding predictionabstractMOTIVATION: T-cell receptors (TCRs) recognize antigenic peptides presented by major histocompatibility complex (MHC) molecules and are central to adaptive immunity. Computational prediction of TCR-epitope binding (TEB) can accelerate immunotherapy development, yet remains hampered by limited labeled data, false-negative noise in unobserved pairs, and over-smoothing in graph-based models. RESULTS: We present DynaTCR, a dynamic graph ensemble learning framework for TEB prediction. DynaTCR encodes TCR and epitope sequences with protein language model embeddings and organizes them into a bipartite interaction graph. A graph regularization-variance-preserving aggregation (GR-VPA) encoder stabilizes message propagation and alleviates over-smoothing, while a global attention layer captures long-range dependencies. Multiple base learners are trained with iteratively updated hard-negative samples to reduce false-negative predictions. Under the StrictTCR evaluation protocol on four public datasets, DynaTCR achieves AUC improvements of 4.0-8.2 percentage points over the strongest existing method and up to 15.8 percentage points in AUPR. On the most stringently curated dataset, DynaTCR attains an AUC of 95.1%. Furthermore, on an independent structure-derived test set, DynaTCR achieves the highest AUC (72.6%) among all compared methods, demonstrating its robustness and effectiveness for TEB prediction and candidate prioritization. AVAILABILITY: Source code and data can be downloaded from: https://github.com/2014402680/TEB/. Xiangzheng Fu, Xinyu Zhang 0012, Linlin Zhuo, Dong-Sheng Cao 0001, Quan Zou 0001 |
Bioinform. | 6 |
| 2026 | HMA-GCA: hybrid manifold augmentation and gated cross-attention for circRNA-miRNA interaction predictionabstractMOTIVATION: Circular RNAs (circRNAs) interact with microRNAs (miRNAs) to regulate gene expression and influence disease progression. However, traditional models tend to overlook the significant contributions of certain features when dealing with diverse sequence information, resulting in the inability to capture some deep topological structures and thus leaving room for improvement in prediction performance. RESULTS: We propose HMA-GCA, a novel framework that integrates hybrid manifold augmentation and gated cross-attention for CMI prediction. The model first constructs multi-scale descriptors by combining sequence-derived features (K-mer, CTD, Doc2Vec) and topological features (Role2Vec, node degree, neighborhood proximity). It then applies PCA for global linear projection and UMAP for local nonlinear manifold learning, enhancing feature representations while preserving intrinsic data geometry. A channel-wise gated cross-attention mechanism dynamically controls the injection of miRNA information into circRNA representations. Extensive experiments on three benchmark datasets show that HMA-GCA consistently outperforms state-of-the-art methods across multiple metrics. To ensure interpretability, we conducted SHAP analysis to quantify the contribution of each feature type, revealing that sequence-derived features and topological similarities are the most influential. Ablation studies confirm the necessity of each module, while case studies demonstrate that top-ranked predictions are supported by literature evidence. Overall, HMA-GCA not only achieves state-of-the-art predictive performance but also provides interpretable insights into the molecular features. AVAILABILITY AND IMPLEMENTATION: The source code and data are freely available at https://github.com/Lixunwind/Prediction-circ-mi-by-Gate.git. The implementation is based on Python and the required dependencies are listed in the repository. Yunzhou Hu, Yansu Wang, Yifeng Bai, Lei Xu 0047, Quan Zou 0001, Chunyu Wang 0002, Mengting Niu |
Bioinform. | 5 |
| 2026 | EAGP: an efficient generative augmentation framework for phage protein classification under severe class imbalanceabstractMOTIVATION: The accurate classification of phage proteins is critical for advancing bacteriophage research. Despite the proliferation of machine learning approaches in this domain, the persistent issue of data imbalance continues to hinder performance, particularly for rare protein sequences. Previous attempts to address this by re-weighting minority classes have faced limitations due to insufficient feature extraction capabilities. RESULTS: In this paper, we introduce EAGP, a novel approach that integrates a generative model-functionally equivalent to a WGAN yet tailored for one-dimensional data-with the Evolutionary Scale Modeling (ESM) protein large language model for robust feature extraction. EAGP exhibits exceptional performance in binary classification and protein function annotation tasks. Crucially, our method not only improves overall classification efficacy but also significantly alleviates the performance degradation typically observed in minority classes. AVAILABILITY AND IMPLEMENTATION: The data and code underlying this article are available in GitHub at https://github.com/Innerly/EAGP and have been archived on Zenodo at https://doi.org/10.5281/zenodo.19928069. Jiaru Li, Yansu Wang, Quan Zou 0001, Hongling Zhu |
Bioinform. | 4 |
| 2026 | AutoGERN: single-cell RNA-seq gene regulatory network inference via explicit link modeling and adaptive architecturesabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) enables transcriptome-wide profiling at single-cell resolution, revealing heterogeneous regulatory programs and making gene regulatory network (GRN) inference both central and challenging. Recent graph neural network (GNN)-based approaches for GRN inference typically model edges only implicitly (for example, via concatenated node embeddings), which limits their ability to capture complex regulatory dependencies. In addition, distributional shifts across scRNA-seq datasets make a single fixed GNN architecture poorly suited for broad generalization. RESULTS: We present AutoGERN, a GNN framework tailored for GRN inference from scRNA-seq data. AutoGERN explicitly models regulatory information in the message-passing space and learns expressive link (edge) embeddings, which a lightweight multilayer perceptron uses to score gene-gene regulatory associations. To enhance flexibility and representational power, AutoGERN employs dual message-passing spaces (within-layer and cross-layer) and integrates a robust AutoGNN-based architecture search to adapt the network design to differing dataset distributions. Extensive experiments on multiple real scRNA-seq datasets demonstrate that AutoGERN consistently achieves superior performance and robustness compared with state-of-the-art baselines. AVAILABILITY AND IMPLEMENTATION: The code and data of AutoGERN are available on GitHub at https://github.com/JChander/AutoGERN and on Zenodo at https://doi.org/10.5281/zenodo.18659807. Jiacheng Wang 0009, Yaojia Chen, Quan Zou 0001, Ximei Luo |
Bioinform. | 3 |
| 2026 | Predicting enhancer-promoter interactions using a stacking-based ensemble strategyabstractMOTIVATION: Enhancer-promoter interactions (EPIs) are essential for gene regulation and disease progression. Recent studies have shown that distal enhancers can regulate target genes through interactions with nearby promoters, providing important insights into transcriptional regulation mechanisms. Although high-throughput experimental techniques have enabled large-scale identification of EPIs, these methods are often costly and time-consuming. In addition, existing computational approaches still face challenges in effectively integrating heterogeneous feature representations from different cell lines. RESULTS: We propose a stacked ensemble framework for EPI prediction that integrates feature representations from diverse cell line datasets using multiple machine learning algorithms. The extracted complementary patterns are further combined by an XGBoost classifier to improve robustness against overfitting. Experiments on six independent datasets show that the proposed method achieves superior accuracy and generalization compared with existing EPI prediction models, with an average AUROC of 0.909 while maintaining computational efficiency. AVAILABILITY: The source code and its archived release are available at GitHub and Zenodo. The Zenodo archive provides a versioned snapshot of the repository: https://zenodo.org/records/19952998. Zhichao Xiao, Haibo Ji, Quan Zou 0001, Yijie Ding, Liang Yu 0002 |
Bioinform. | 3 |
| 2026 | Contrastive learning in both structure and function spaces improve drug-target interaction predictionabstractBACKGROUND : Identifying drug–target interactions (DTIs) is essential in drug discovery and repositioning. Recently, deep learning has become the mainstream methodology for DTI prediction. However, the scarcity of three-dimensional structural data has forced almost all methods to predict drug-target interactions with low-dimensional data, thereby constraining their overall performance. METHODS : In tackling this challenge, we introduce a novel approach, CLSF-DTI. CLSF-DTI incorporates high-dimensional structural and functional information into the drug and protein features through contrastive learning during the feature extraction stage. This ensures that the model no longer solely focuses on sequence information, leading to a more precise modeling outcome. RESULTS : Experiments on five benchmark datasets demonstrate that CLSF-DTI achieves the best overall performance among five state-of-the-art baselines. Through ablation studies, we further prove that the contrastive module enhances the predictive performance and generalization ability of CLSF-DTI. Moreover, CLSF-DTI successfully identified some ligands for the protein PKA-Cα in drug screening experiments. CONCLUSIONS : This study proposed a contrastive learning model CLSF-DTI that integrates structural and functional similarity. It outperforms existing methods in drug-target interaction prediction and has stronger generalization ability. However, the handling of unbalanced data and long-distance dependencies still needs to be improved in the future. The data and source code are available at https://github.com/ZhangLab312/CLSF_DTI . Yongqing Zhang 0001, Shuwen Xiong, Zixuan Wang 0025, Quan Zou 0001 |
BMC Bioinform. | 7 |
| 2026 | AdaptIPs: A dual-channel deep learning framework integrating protein language model representations and transfer learning for phosphorylation-site prediction
Aoyun Geng, Yanfei Qu, Junlin Xu, Yajie Meng, Quan Zou 0001, Feifei Cui |
Eng. Appl. Artif. Intell. | 6 |
| 2026 | Explainable multiscale representation learning for anticancer peptide prediction
Yongqing Zhang 0001, Zhigan Zhou, Yugui Xu, Jin Wu 0002, Quan Zou 0001, Lei Xu 0047 |
Eng. Appl. Artif. Intell. | 8 |
| 2026 | CNNCaps-DBP: Leveraging protein language models with attention-augmented convolution for DNA-binding protein prediction
Ziyuan Yan, Aoyun Geng, Yazi Li, Jiajing Wang, Junlin Xu, Yajie Meng, Leyi Wei, Quan Zou 0001, Feifei Cui |
Neural Networks | 8 |
| 2026 | MCFusion-DDI: Multimodal cross-attention fusion of local-global features and latent drug associations for explainable DDI prediction
Yongqing Zhang 0001, Yugui Xu, Zhigan Zhou, Jin Wu 0002, Quan Zou 0001, Lei Xu 0047 |
Neural Networks | 6 |
| 2026 | Enhancing anticancer peptide discovery: A fusion-centric framework with conditional diffusion for prediction and generationabstractAnticancer peptides (ACPs) are short bioactive sequences that selectively target tumor cells with minimal toxicity, positioning them as promising candidates for next-generation cancer therapies. However, existing computational models face limitations in sequence representation and class imbalance. To address these challenges, we propose UACD-ACPs, a unified fusion-driven framework that integrates a diffusion-inspired noise-conditioned classifier for ACP prediction and a diffusion-based peptide generation module with cancer-type-aware organization for targeted downstream screening. The classification module integrates ProtBERT-based semantic embeddings with physicochemical descriptors via the Multiscale Embedding Compression Strategy (MECS) and a diffusion-inspired noise-conditioned encoder, substantially enhancing predictive robustness and accuracy, particularly under challenging imbalanced multi-class settings. In the generative pipeline, we introduce a denoising diffusion-based generative framework augmented by two novel fusion modules: the Bitemporal Fusion Module (BFM) and the Temporal Feature Attention Module (TFAM). These modules perform multi-scale temporal and semantic fusion to promote the generation of structurally coherent and functionally relevant peptide candidates. Experimental results demonstrate that UACD-ACPs outperforms state-of-the-art methods in terms of accuracy, F1-score, and AUC-ROC. The generated peptides exhibit favorable physicochemical properties, diverse secondary structures, and strong structural stability, as validated by molecular dynamics simulations and membrane-binding analyses. Overall, this study highlights the potential of fusion-driven diffusion-based frameworks for alleviating class imbalance and data heterogeneity in anticancer peptide modeling, paving the way for scalable and biologically grounded ACP discovery. Binyu Li, Xin Zhang 0103, Prayag Tiwari, Quan Zou 0001, Yijie Ding, Xiaoyi Guo |
PLoS Comput. Biol. | 5 |
| 2026 | CFGSCDSA: Predicting circRNA-drug sensitivity associations based on collaborative feature learning and graph structure learningabstractMOTIVATION: The expression of circular RNAs (circRNAs) has been shown to be strongly correlated with drug sensitivity in human cells. However, experimental validation using wet-lab techniques is costly and inefficient, leaving a substantial portion of circRNA-drug sensitivity associations undiscovered. Therefore, improving the prediction efficiency of circRNA and sensitivity associations remains critical. METHODS: Here, we describe a method that integrates collaborative feature learning and graph structure learning to predict associations between circRNAs and drug sensitivity (CFGSCDSA). Specifically, collaborative learning integrated heterogeneous features from diverse data sources, thereby addressing the issue of data sparsity. Furthermore, graph structure learning with a confidence-guided pseudo-labeling strategy was employed to mitigate the detrimental effect of excessive negative samples. Results: Experimental evaluation revealed that CFGSCDSA attained superior performance compared to all competing models. Moreover, case studies provided further evidence of its capability to accurately predict both novel associations and new drug-related links. Quan Zou 0001, Chunyu Wang 0002, Mengting Niu |
PLoS Comput. Biol. | 2 |
| 2026 | SAD: Sparse-Aware Diffusion Model for Single-Cell Gene Expression CompletionabstractSingle-cell RNA sequencing (scRNA-seq) is entering an era of foundation models that accept the complete gene atlas as input, yet most current datasets cover only 10-12 k genes and contain numerous technical zeros, severely limiting the generalization of these models in downstream tasks. To address this, we pioneer the gene-completion task for scRNA-seq and present SAD, a diffusion-based framework tailored to extremely sparse data, capable of completing genes and correcting sparsity bias under high missing rates. Unlike imputation or reconstruction methods that rely on the i.i.d. assumption, SAD's completion paradigm can generate gene entries originally absent from the expression profile, be aware of and rectify sparsity-distribution bias, and supply foundation models with consistent, reliable inputs of more than 30 k genes. Extensive benchmarks show that SAD significantly outperforms existing methods across multiple completion metrics, particularly in extreme scenarios with missing rates above 80%. This provides a data foundation for reusing missing scRNA-seq information and for precision-medicine applications. Yixin Xiang, Zixuan Wang 0025, Mengwen Liu, Quan Zou 0001, Yongqing Zhang 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2026 | SMCC: A Novel Clustering Method for Single- and Multi-Omics Data Based on Co-Regularized Network FusionabstractClustering is a common technique for statistical data analysis and is essential for developing precision medicine. Numerous computational methods have been proposed for integrating multi-omics data to identify cancer subtypes. However, most existing clustering models based on network fusion fail to preserve the consistency of the distribution of the data before and after fusion. Motivated by this observation, we would like to measure and minimize the distribution difference between networks, which may not be in the same space, to improve the performance of data fusion. We were therefore motivated to develop a flexible clustering model, based on network fusion, that minimizes the distribution difference between the data before and after fusion by co-regularization; the model can be applied to both single- and multi-omics data. We propose a new network fusion model for single- and multi-omics data clustering for identifying cancer or cell subtypes based on co-regularized network fusion (SMCC). SMCC integrates low-rank subspace representation and entropy to fuse networks. In addition, it measures and minimizes the distribution difference between the similarity networks and the fusion network by co-regularization. The model can both reduce the noise interference in the source data and make the statistical characteristics of the fusion result closer to those of the source data. We evaluated the clustering performance of SMCC across 16 real single- and multi-omics dataset. The experimental results demonstrated that SMCC is superior to 17 state-of-the-art clustering methods. Moreover, it is effective for identifying cancer or cell subtypes, thereby promoting the development of precision medicine. Sha Tian, Yushan Qiu, Quan Zou 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2026 | DeepNhKcr: Explainable Deep Learning Framework for the Prediction of Crotonylation Sites of Non-Histone Lysine in Plants Based on Pre-Trained Protein Language ModelabstractLysine crotonylation (Kcr) is an important protein modification occurring after translation in biology, serving an essential function in a range of biological processes in both plants and animals, including the regulation of gene expression, the maintenance of cellular metabolic balance, and the enhancement of photosynthesis. Exploring the detection of Kcr sites is essential for uncovering their biological functions. Nonetheless, conventional experimental approaches for detection are often time-consuming, expensive, and hindered by various technical constraints, making the precise identification of Kcr sites a significant challenge. This study seeks to develop a computational approach for the rapid and accurate prediction of Kcr sites in plant non-histone proteins. We introduce a novel deep learning framework named DeepNhKcr, which integrates the protein language model (ESM2) with a bidirectional long short-term memory (BiLSTM) network. To address the challenge of data imbalance, the model replaces the conventional cross-entropy loss with the focal loss function. In addition, DeepNhKcr combines advanced deep learning approaches with traditional protein encoding strategies to enable effective feature extraction and integration. This method not only significantly boosts the accuracy of predicting Kcr sites in non-histone proteins of plants. but also provides interpretability, shedding light on the potential links between key sequence characteristics and their biological roles. DeepNhKcr delivers outstanding results, surpassing existing machine learning and deep learning models, and demonstrating excellent performance in both five-fold cross-validation and independent test experiments. Moreover, the model integrates interpretability analysis techniques to investigate the connections between important sequence features and their biological roles. DeepNhKcr acts as a powerful method for detecting Kcr sites in plant non-histone proteins and is anticipated to greatly advance future studies in plant Kcr site prediction. Zhenjie Luo, Aoyun Geng, Junlin Xu, Yajie Meng, Shankai Yan, Leyi Wei, Qingchen Zhang 0001, Quan Zou 0001, Feifei Cui |
IEEE Trans. Comput. Biol. Bioinform. | 10 |
| 2026 | DeepR2OM: Accurate Recognition for RNA 2′-O-Methylation Sites in Human Genome Using Deep Learningabstract2'-O-methylation (2OM) of ribose is a widespread RNA modification that significantly impacts RNA stability, structure, and function. Accurately predicting 2OM sites is crucial for understanding RNA's biological functions and related pathologies. Traditional detection methods pose challenges such as resource intensiveness, potential RNA sample damage, and high costs. However, recent advancements in machine learning, particularly deep learning techniques, offer rapid and cost-effective prediction solutions. In this study, we introduce DeepR2OM, a novel method integrating feature selection and deep learning for 2OM sites prediction. DeepR2OM encodes sequences using eight RNA descriptors, employs feature selection algorithms to reduce dimensions, and then utilizes a deep learning network for training. After evaluating various deep learning architectures, we selected Convolutional Neural Network (CNN), Multi-Head Self-Attention mechanism, and Deep Neural Network (DNN) as our final prediction models. Experimental results demonstrate DeepR2OM's effectiveness, achieving 87.1% accuracy (ACC), 85.5% recall rate (Recall), 87.9% precision (PRE), and a Matthews correlation coefficient (MCC) of 75.7% on an independent test set. This tool serves as a valuable resource for exploring the functional and bioinformatic aspects of 2OM sites. Shun Gao, Ziyuan Yan, Feifei Cui, Leyi Wei, Qingchen Zhang 0001, Quan Zou 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 7 |
| 2026 | Multi-View Clustering With Cauchy-Schwarz Mutual Information MaximinabstractInformation bottleneck (IB) leverages information theory to guide the learning process of deep multi-view clustering (MVC). It optimizes the trade-off between multi-view compression and preservation by minimizing and maximizing mutual information (MI). Although existing deep MVC based on IB has witnessed great achievements, they usually resort to variational inference to estimate the MI lower bound, which typically introduces estimation errors and results in an unstable lower bound of MI. In this study, we propose a novel Cauchy-Schwarz Mutual Information Maximin (CS-MIM) method, which directly estimates MI with closed-form expressions without requiring variational inference, possessing explicit multi-view information modeling capabilities. Specifically, we first present a non-parametric MI estimation method with Cauchy-Schwarz (CS) divergence, which leverages multi-kernel Gram matrices to capture distributional similarities and avoids the approximation errors introduced by the neural estimators of variational inference. Then, based on the new estimation method, a MI maximin mechanism is devised to parameterize the IB principle with analytical gradients, which facilitates effective compression of multi-view data while preserving the relevant features. Finally, we design a cross-view adaptive attention (CAA) mechanism constrained by the MI based on CS divergence, which further captures the complementarity across views under the guidance of the fused multi-view representation. Extensive evaluation results on 12 public available datasets demonstrate that the CS-MIM remarkably outperforms existing SOTA approaches. Zhen Tian 0004, Enlai Ouyang, Quan Zou 0001 |
IEEE Trans. Image Process. | 3 |
| 2026 | BloodPatrol: Revolutionizing Blood Cancer Diagnosis - Advanced Real-Time Detection Leveraging Deep Learning & Cloud TechnologiesabstractCloud computing and Internet of Things (IoT) technologies are gradually becoming the technological changemakers in cancer diagnosis. Blood cancer is an aggressive disease affecting the blood, bone marrow, and lymphatic system, and its early detection is crucial for subsequent treatment. Flow cytometry has been widely studied as a commonly used method for detecting blood cancer. However, the high computation and resource consumption severely limit its practical application, especifically in regions with limited medical and computational resources. In this study, with the help of cloud computing and IoT technologies, we develop a novel blood cancer dynamic monitoring diagnostic model named BloodPatrol based on an intelligent feature weight fusion mechanism. The proposed model is capable of capturing the dual-view importance relationship between cell samples and features, greatly improving prediction accuracy and significantly surpassing previous models. Besides, benefiting from the powerful processing ability of cloud computing, BloodPatrol can run on a distributed network to efficiently process large-scale cell data, which provides immediate and scalable blood cancer diagnostic services. Jinhang Wei, Longyue Wang, Zhecheng Zhou, Linlin Zhuo, Xiangxiang Zeng, Xiangzheng Fu, Quan Zou 0001, Keqin Li 0001, Zhongjun Zhou |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | UPDMIA: Unified Processing of Diverse Medical Imaging Data via a Multi-Channel EfficientNet ArchitectureabstractPulmonary medical image analysis has become increasingly challenged by the overwhelming volume and heterogeneity of CT, MRI, PET, X-ray and histopathology data, leading to time-consuming workflows and error-prone manual interpretation. To address this, we propose UPDMIA, a unified deep-learning pipeline built on a structurally enhanced EfficientNet-B7 backbone, coupled with an intelligent modality-adaptive classification head and a comprehensive multimodal data augmentation suite. All input images are preprocessed to$224 \times 224$resolution, normalized, and subjected to probabilistic augmentation (flips, rotations, cropping, Gaussian blur, color jitter, random erasing). Modality detection then dynamically routes features through lightweight, task-specific classifiers. Experimental evaluation across five modalities demonstrates strong performance: CT (F1 =$96.99 \%)$, MRI$(\mathrm{F} 1=87.40 \%)$, PET$(\mathrm{F} 1=93.35 \%)$, X-ray ($\mathrm{F} 1= 95.49 \%$) and histological slides ($F 1=97.60 \%$), with overall precision and recall above 88 % for all domains. This unified approach matches or exceeds specialized single-modality models while simplifying deployment. Binyun Yang, Quan Zou 0001, Yansu Wang |
BIBM | 2 |
| 2025 | SeqAlignXGBoost: Sequence Alignment and Feature Selection for m1A Modification Site Identification
Yizheng Wang, Yijie Ding, Quan Zou 0001 |
ICIC (28) | 3 |
| 2025 | ReAlign-Star: An Optimized Realignment Method for Multiple Sequence Alignment, Targeting Star Algorithm Tools
Yixiao Zhai, Pinglu Zhang, Yi Liu 0112, Quan Zou 0001 |
ICIC (26) | 4 |
| 2025 | ET-PROTACs: modeling ternary complex interactions using cross-modal learning and ternary attention for accurate PROTAC-induced degradation predictionabstractMOTIVATION: Accurately predicting the degradation capabilities of proteolysis-targeting chimeras (PROTACs) for given target proteins and E3 ligases is important for PROTAC design. The distinctive ternary structure of PROTACs presents a challenge to traditional drug-target interaction prediction methods, necessitating more innovative approaches. While current state-of-the-art (SOTA) methods using graph neural networks (GNNs) can discern the molecular structure of PROTACs and proteins, thus enabling the efficient prediction of PROTACs' degradation capabilities, they rely heavily on limited crystal structure data of the POI-PROTAC-E3 ternary complex. This reliance underutilizes rich PROTAC experimental data and neglects intricate interaction relationships within ternary complexes. RESULTS: In this study, we propose a model based on cross-modal strategy and ternary attention technology, ET-PROTACs, to predict the targeted degradation capabilities of PROTACs. Our model capitalizes on the strengths of cross-modal methods by using equivariant GNN graph neural networks to process the graph structure and spatial coordinates of PROTAC molecules concurrently while utilizing sequence-based methods to learn the protein sequence information. This integration of cross-modal information is cohesively harnessed and channeled into a ternary attention mechanism, specially tailored for the unique structure of PROTACs, enabling the congruent modeling of both PROTAC and protein modalities. Experimental results demonstrate that the ET-PROTACs model outperforms existing SOTA methods. Moreover, visualizing attention scores illuminates crucial residues and atoms pivotal in specific POI-PROTAC-E3 interactions, thus offering invaluable insights and guidance for future pharmaceutical research. AVAILABILITY AND IMPLEMENTATION: The codes of our model are available at https://github.com/GuanyuYue/ET-PROTACs. Guanyu Yue, Li Wang 0145, Quan Zou 0001, Xiangzheng Fu, Dong-Sheng Cao 0001 |
Briefings Bioinform. | 6 |
| 2025 | GSTRPCA: irregular tensor singular value decomposition for single-cell multi-omics data clusteringabstractSingle-cell multi-omics refers to the various types of biological data at the single-cell level. These data have enabled insight and resolution to cellular phenotypes, biological processes, and developmental stages. Current advances hold high potential for breakthroughs by integrating multiple different omics layers. However, singlecell multi-omics data usually have different feature dimensions and direct or indirect relationships. How to keep the data structure of these different data and extract hidden relationships is a major challenge for omics data integration, and effective integration models are urgently needed. In this paper, we propose an irregular tensor decomposition model (GSTRPCA) based on tensor robust principal component analysis (TRPCA). We developed a weighted threshold model for the decomposition of irregular tensor data by combining low-rank and sparsity constraints, which requires that the low-dimensional embeddings of the data remain lowrank and sparse. The major advantage of the GSTRPCA algorithm is its ability to keep the original data structure and explore hidden related features among omics data. For GSTRPCA, we also designed an effective algorithm that theoretically guarantees global convergence for the tensor decomposition. The computational experiments on irregular tensor datasets demonstrate that GSTRPCA significantly outperformed the state-of-the-art methods and hence confirm the superiority of GSTRPCA in clustering single-cell multiomics data. To our knowledge, this is the first tensor decomposition method for irregular tensor data to keep the data structure and hence improve the clustering performance for single-cell multi-omics data. GSTRPCA is a Matlabbased algorithm, and the code is available from https://github.com/GGL-B/GSTRPCA. Lu-Bin Cui, Guiliang Guo, Michael Kwok-Po Ng, Quan Zou 0001, Yushan Qiu |
Briefings Bioinform. | 4 |
| 2025 | An overview of computational methods in single-cell transcriptomic cell type annotationabstractThe rapid accumulation of single-cell RNA sequencing data has provided unprecedented computational resources for cell type annotation, significantly advancing our understanding of cellular heterogeneity. Leveraging gene expression profiles derived from transcriptomic data, researchers can accurately infer cell types, sparking the development of numerous innovative annotation methods. These methods utilize a range of strategies, including marker genes, correlation-based matching, and supervised learning, to classify cell types. In this review, we systematically examine these annotation approaches based on transcriptomics-specific gene expression profiles and provide a comprehensive comparison and categorization of these methods. Furthermore, we focus on the main challenges in the annotation process, especially the long-tail distribution problem arising from data imbalance in rare cell types. We discuss the potential of deep learning techniques to address these issues and enhance model capability in recognizing novel cell types within an open-world framework. Zixuan Wang 0025, Quan Zou 0001, Yongqing Zhang 0001 |
Briefings Bioinform. | 5 |
| 2025 | Assessment and applications of joint profiling of single-cell chromatin accessibility and transcriptomeabstractJoint profiling technologies combining single-cell chromatin accessibility (CA) and transcriptome sequencing enable cellular heterogeneity analysis from both gene and cis-regulatory element perspectives, greatly advancing molecular biology at a cellular resolution. These techniques have been used to construct gene regulatory networks across diverse cell types and biological tissues, contributing significantly to the mapping of cell developmental trajectories. In this review, we summarize existing single-cell joint profiling methods for CA and transcriptomics and systematically evaluate the data quality of each modality using consistent criteria: the median number of genes detected per cell (RNA) and the median number of accessible peaks per cell (ATAC). Furthermore, we examine relevant bioinformatics tools and highlight their applications in various omics research contexts. Finally, we discuss the current limitations of joint profiling technologies, prospects for future improvement, the extensibility of computational tools, and the potential for co-assaying with additional omics data. Jiechen Wang, Quan Zou 0001, Ximei Luo |
Briefings Bioinform. | 4 |
| 2025 | AttenRNA: multi-scale deep attentive model with RNA feature variability analysisabstractAccurate identification of diverse RNA types, including messenger RNAs (mRNAs), long non-coding RNAs (lncRNAs), and circular RNAs (circRNAs), is essential for understanding their roles in gene regulation, disease progression, and epigenetic modification. Existing studies have primarily focused on binary classification tasks, such as distinguishing lncRNAs from mRNAs or identifying specific circRNAs, often overlooking the complex sequence patterns shared across multiple RNA types. To address this limitation, we developed AttenRNA, a multi-class classification model that integrates multi-scale k-mer embeddings and attention mechanisms to simultaneously differentiate between various RNA classes. AttenRNA achieved high weighted F1 scores of 89.8% and 89.6% on the validation and test sets, respectively, demonstrating strong classification performance and robustness. Dimensionality reduction using Uniform Manifold Approximation and Projection further confirmed the model's ability to learn discriminative features among RNA types. Additionally, AttenRNA exhibited strong generalization ability on cross-species data, achieving weighted F1 scores of 83.89% and 83.38% on the mouse RNA validation and test sets, respectively. These results suggest that AttenRNA offers a reliable and scalable solution for systematic RNA function analysis. Quan Zou 0001, Chao Zhan |
Briefings Bioinform. | 2 |
| 2025 | ST-GCP: a graph convolutional network model with contrastive consistency and permutation for spatial transcriptomicsabstractSpatial transcriptomics (STs) technology is a powerful technique that simultaneously preserves gene expression profiles and spatial information, enabling deeper exploration of tissue organization and function. However, many existing computational approaches often rely on labeled ST data and overlook the rich spatial information, resulting in limited representations and suboptimal clustering. In this paper, we propose ST-GCP, a self-supervised graph representation learning framework for ST data, which incorporates a structure-feature perturbation mechanism. First, ST-GCP applies feature-level random permutation of the gene expression matrix and random edge dropout in the spatial neighbor network, creating two complementary augmented graph views of ST data. ST-GCP then employs a two-layer graph convolutional network (GCN) encoder-decoder to extract spatial representations and reconstruct gene expression. Finally, a cosine-similarity-based contrastive objective aligns the view-specific representations, and the overall loss jointly optimizes reconstruction fidelity and contrastive consistency, thereby coupling graph topology with transcriptomic profiles in a shared low-dimensional space. Experimental results on multiple ST datasets demonstrate that ST-GCP can uncover biologically meaningful patterns, such as tumor heterogeneity, brain developmental architecture, and cellular developmental trajectories. Yajie Meng, Xianfang Tang, Feifei Cui, Xiangzheng Fu, Quan Zou 0001, Junlin Xu |
Briefings Bioinform. | 8 |
| 2025 | TargetCLP: clathrin proteins prediction combining transformed and evolutionary scale modeling-based multi-view features via weighted feature integration approachabstractClathrin proteins, key elements of the vesicle coat, play a crucial role in various cellular processes, including neural function, signal transduction, and endocytosis. Disruptions in clathrin protein functions have been associated with a wide range of diseases, such as Alzheimer's, neurodegeneration, viral infection, and cancer. Therefore, correctly identifying clathrin protein functions is critical to unravel the mechanism of these fatal diseases and designing drug targets. This paper presents a novel computational method, named TargetCLP, to precisely identify clathrin proteins. TargetCLP leverages four single-view feature representation methods, including two transformed feature sets (PSSM-CLBP and RECM-CLBP), one qualitative characteristics feature, and one deep-learned-based embedding using ESM. The single-view features are integrated based on their weights using differential evolution, and the BTG feature selection algorithm is utilized to generate a more optimal and reduced subset. The model is trained using various classifiers, among which the proposed SnBiLSTM achieved remarkable performance. Experimental and comparative results on both training and independent datasets show that the proposed TargetCLP offers significant improvements in terms of both prediction accuracy and generalization to unseen data, furthering advancements in the research field. Matee Ullah, Shahid Akbar, Kashif Ahmad Khan, Quan Zou 0001 |
Briefings Bioinform. | 5 |
| 2025 | stHGC: a self-supervised graph representation learning for spatial domain recognition with hybrid graph and spatial regularizationabstractAdvancements in spatial transcriptomics (ST) technology have enabled the analysis of gene expression while preserving cellular spatial information, greatly enhancing our understanding of cellular interactions within tissues. Accurate identification of spatial domains is crucial for comprehending tissue organization. However, the effective integration of spatial location and gene expression still faces significant challenges. To address this challenge, we propose a novel self-supervised graph representation learning framework named stHGC for identifying spatial domains. Firstly, a hybrid neighbor graph is constructed by integrating different similarity metrics to represent spatial proximity and high-dimensional gene expression features. Secondly, a self-supervised graph representation learning framework is introduced to learn the representation of spots in ST data. Within this framework, the graph attention mechanism is utilized to characterize relationships between adjacent spots, and the self-supervised method ensures distinct representations for non-neighboring spots. Lastly, a spatial regularization constraint is employed to enable the model to retain the structural information of spatial neighbors. Experimental results demonstrate that stHGC outperforms state-of-the-art methods in identifying spatial domains across ST datasets with different resolutions. Furthermore, stHGC has been proven to be beneficial for downstream tasks such as denoising and trajectory inference, showcasing its scalability in handling ST data. Runqing Wang, Qiguo Dai, Xiaodong Duan, Quan Zou 0001 |
Briefings Bioinform. | 4 |
| 2025 | BridgeSyn: a bridging fusion framework for drug combination synergy predictionabstractDrug combination is a promising therapeutic strategy for complex diseases. However, only a small fraction of potential drug combinations exhibit true synergistic effects, making the prediction of drug synergy a critical yet challenging task. In this study, we propose BridgeSyn, a novel bridge fusion framework for drug synergy prediction. BridgeSyn leverages the knowledge from pretrained biological language models to enrich both drug compound and cell line representations. We introduce a bridging fusion mechanism that employs a set of shared latent tokens derived from global features, serving as a semantic interface to effectively fuse the representations of drug pairs and cell lines. By combining biological prior knowledge with this fusion strategy, BridgeSyn can capture complex biological interactions and achieve superior prediction results. Extensive experiments on two public datasets demonstrate that BridgeSyn consistently outperforms existing computation methods. Suwan Mao, Quan Zou 0001, Xi Su, Junjie Wang 0005, Ximei Luo |
Briefings Bioinform. | 4 |
| 2025 | FORAlign: accelerating gap-affine DNA pairwise sequence alignment using FOR-blocks based on Four Russians approach with linear space complexityabstractPairwise sequence alignment (PSA) serves as the cornerstone in computational bioinformatics, facilitating multiple sequence alignment and phylogenetic analysis. This paper introduces the FORAlign algorithm, leveraging the Four Russians algorithm with identical upper-bound time and space complexity as the Hirschberg divide-and-conquer PSA algorithm, aimed at accelerating Hirschberg PSA algorithm in parallel. Particularly notable is its capability to achieve up to 16.79 times speedup when aligning sequences with low sequence similarity, compared to the conventional Needleman-Wunsch PSA method using non-heuristic methods. Empirical evaluations underscore FORAlign's superiority over existing wavefront alignment (WFA) series software, especially in scenarios characterized by low sequence similarity during PSA tasks. Our method is capable of directly aligning monkeypox sequences with other sequences using non-heuristic methods. The algorithm was implemented within the FORAlign library, providing functionality for PSA and foundational support for multiple sequence alignment and phylogenetic trees. The FORAlign library is freely available at https://github.com/malabz/FORAlign. Yanming Wei, Tong Zhou 0016, Yixiao Zhai, Liang Yu 0002, Quan Zou 0001 |
Briefings Bioinform. | 5 |
| 2025 | Computational toxicology in drug discovery: applications of artificial intelligence in ADMET and toxicity predictionabstractToxicity risk assessment plays a crucial role in determining the clinical success and market potential of drug candidates. Traditional animal-based testing is costly, time-consuming, and ethically controversial, which has led to the rapid development of computational toxicology. This review surveys over 20 ADMET prediction platforms, categorizing them into rule/statistical-based methods, machine learning (ML) methods, and graph-based methods. We also summarize major toxicological databases into four types: chemical toxicity, environmental toxicology, alternative toxicology, and biological toxin databases, highlighting their roles in model training and validation. Furthermore, we review recent advancements in ML and artificial intelligence (AI) applied to toxicity prediction, covering acute toxicity, organ-specific toxicities, and carcinogenicity. The field is transitioning from single-endpoint predictions to multi-endpoint joint modeling, incorporating multimodal features. We also explore the application of generative modeling techniques and interpretability frameworks to improve the accuracy and credibility of predictions. Additionally, we discuss the use of network toxicology in evaluating the safety of traditional Chinese medicines (TCMs) and the potential of large language models (LLMs) in literature mining, knowledge integration, and molecular toxicity prediction. Finally, we address current challenges, including data quality, model interpretability, and causal inference, and propose future directions such as multi-omics integration, interpretable AI models, and domain-specific LLMs, aiming to provide more efficient and precise technical support for preclinical toxicity assessments in drug development. Jiangyan Zhang, Yuncong Zhang, Junyang Huang, Liping Ren, Chuantao Zhang, Quan Zou 0001, Yang Zhang 0125 |
Briefings Bioinform. | 7 |
| 2025 | Differentiable graph clustering with structural grouping for single-cell RNA-seq dataabstractMOTIVATION: Clustering cells into subpopulations is one of the most crucial tasks in single-cell RNA sequencing (scRNA-seq) data analysis, which provides support for biological research at cellular level. With the development of graph neural networks, deep graph clustering approaches have achieved excellent performance by modeling the topological relationships between cells. However, existing approaches rely on cell node and its neighbors to obtain the cell feature representation, which ignore the graph cluster structure hidden in scRNA-seq data. Besides, how to bridge the heterogeneous gap between cell node feature and its structural information remains a highly challenging problem. RESULTS: Here, we propose a novel differentiable graph clustering with structural grouping (DGCSG) for scRNA-seq data, which incorporates graph cluster information into deep graph clustering model by designing a differentiable clustering mechanism to learn clustering-friendly representation. Firstly, an interactive module is devised to dynamically transfer node representations learned by autoencoder (AE) to graph attention autoencoder (GATE) in layer-by-layer manner. Then, to characterize graph cluster information, a differentiable clustering mechanism is proposed to transform K-way normalized cuts from a discrete optimization problem into differentiable learning objective through spectral relaxation, which jointly optimizes the GATE by allocating more attention scores to nodes in the same graph cluster. Finally, a decoupled self-supervised optimization is proposed, which guides the representation learning of AE and GATE in the interactive module. Extensive evaluations on 14 scRNA-seq benchmarks verify the superiority of DGCSG compared with state-of-the-art baselines. AVAILABILITY AND IMPLEMENTATION: The code associated with this work is available on GitHub (https://github.com/Xiaoqiang-Yan/DGCSG). Shike Du, Quan Zou 0001, Zhen Tian 0004 |
Bioinform. | 3 |
| 2025 | Enhancing and accelerating cell type deconvolution of large-scale spatial transcriptomics slices with dual network modelabstractMOTIVATION: Cell type deconvolution deciphers spatial distribution of mRNA transcripts at single cell level by integrating single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics data to infer mixture of cell types of spots in slices. Current algorithms are criticized for neglecting connection between scRNA-seq and spatial transcriptomics data, as well as time-consuming, hampering their application to large-scale datasets. RESULTS: In this study, we propose a joint learning nonnegative matrix factorization algorithm for fast cell type deconvolution (aka jMF2D), which integrates scRNA-seq and spatial transcriptomics data with network models. To bridge scRNA-seq and spatial transcriptomics data, jMF2D jointly learns cell type similarity network to enhance quality of signatures of cell types, thereby promoting accuracy and efficiency of deconvolution. Experiments demonstrate that jMF2D outperforms state-of-the-art baselines in terms of accuracy by saving about 90% running time on various datasets generated by different platforms. Furthermore, it can also facilitates the identification of spatial domains and bio-marker genes, providing an efficient and effective model for analyzing spatial transcriptomics data. AVAILABILITY AND IMPLEMENTATION: The software is coded using python, and is free available for academic https://github.com/xkmaxidian/jMF2D. Yuhong Zha, Shaoqing Feng, Quan Zou 0001, Xiaoke Ma 0001 |
Bioinform. | 4 |
| 2025 | ReAlign-P: a vertical iterative realignment method for protein multiple sequence alignmentabstractMOTIVATION: Reliable protein multiple sequence alignment (MSA) is essential for downstream biomedical research and directly impacts the accuracy of analytical results. However, protein sequences often exhibit low similarity and complex alignment patterns, and existing general alignment tools frequently fall short in terms of accuracy. Many current realignment methods are outdated, suffering from issues such as code obsolescence and inadequate precision. As a result, there is a pressing need for realignment methods that can better address these challenges. RESULTS: This study introduces ReAlign-P, a realignment tool designed specifically for protein MSA. ReAlign-P first divides the initial alignment into three regions and applies a novel vertical iterative realignment strategy to optimize the more conserved middle region. This method is by default compatible with MUSCLE5 for realignment, leading to a significant improvement in accuracy. We evaluated initial alignments generated using 10 different MSA parameter configurations across four protein benchmark datasets. The results demonstrate that ReAlign-P consistently outperforms or matches the quality of the initial alignments in all cases. In contrast, RASCAL-the only other currently functional protein realignment tool-sometimes even reduces alignment quality. ReAlign-P not only delivers more substantial improvements but also exhibits greater stability, effectively addressing the gap in available protein realignment tools. AVAILABILITY AND IMPLEMENTATION: The source code and test data for ReAlign-P are available on GitHub (https://github.com/malabz/ReAlign-P). Yixiao Zhai, Pinglu Zhang, Quan Zou 0001, Ximei Luo |
Bioinform. | 3 |
| 2025 | Predicting circRNA-disease associations with shared units and multi-channel attention mechanismsabstractMOTIVATION: Circular RNAs (circRNAs) have been identified as key players in the progression of several diseases; however, their roles have not yet been determined because of the high financial burden of biological studies. This highlights the urgent need to develop efficient computational models that can predict circRNA-disease associations, offering an alternative approach to overcome the limitations of expensive experimental studies. Although multi-view learning methods have been widely adopted, most approaches fail to fully exploit the latent information across views, while simultaneously overlooking the fact that different views contribute to varying degrees of significance. RESULTS: This study presents a method that combines multi-view shared units and multichannel attention mechanisms to predict circRNA-disease associations (MSMCDA). MSMCDA first constructs similarity and meta-path networks for circRNAs and diseases by introducing shared units to facilitate interactive learning across distinct network features. Subsequently, multichannel attention mechanisms were used to optimize the weights within similarity networks. Finally, contrastive learning strengthened the similarity features. Experiments on five public datasets demonstrated that MSMCDA significantly outperformed other baseline methods. Additionally, case studies on colorectal cancer, gastric cancer, and nonsmall cell lung cancer confirmed the effectiveness of MSMCDA in uncovering new associations. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at https://github.com/zhangxue2115/MSMCDA.git. Quan Zou 0001, Mengting Niu, Chunyu Wang 0002 |
Bioinform. | 2 |
| 2025 | metaTP: a meta-transcriptome data analysis pipeline with integrated automated workflowsabstractBACKGROUND: The accessibility of sequencing technologies has enabled meta-transcriptomic studies to provide a deeper understanding of microbial ecology at the transcriptional level. Analyzing omics data involves multiple steps that require the use of various bioinformatics tools. With the increasing availability of public microbiome datasets, conducting meta-analyses can reveal new insights into microbiome activity. However, the reproducibility of data is often compromised due to variations in processing methods for sample omics data. Therefore, it is essential to develop efficient analytical workflows that ensure repeatability, reproducibility, and the traceability of results in microbiome research. RESULTS: We developed metaTP, a pipeline that integrates bioinformatics tools for analyzing meta-transcriptomic data comprehensively. The pipeline includes quality control, non-coding RNA removal, transcript expression quantification, differential gene expression analysis, functional annotation, and co-expression network analysis. To quantify mRNA expression, we rely on reference indexes built using protein-coding sequences, which help overcome the limitations of database analysis. Additionally, metaTP provides a function for calculating the topological properties of gene co-expression networks, offering an intuitive explanation for correlated gene sets in high-dimensional datasets. The use of metaTP is anticipated to support researchers in addressing microbiota-related biological inquiries and improving the accessibility and interpretation of microbiota RNA-Seq data. CONCLUSIONS: We have created a conda package to integrate the tools into our pipeline, making it a flexible and versatile tool for handling meta-transcriptomic sequencing data. The metaTP pipeline is freely available at: https://github.com/nanbei45/metaTP . Limuxuan He, Quan Zou 0001, Yansu Wang |
BMC Bioinform. | 2 |
| 2025 | SGPS-IMR: Efficiently inferring microbial resistance using self-supervised graph perturbation strategy
Linlin Zhuo, Zhecheng Zhou, Xiangzheng Fu, Quan Zou 0001 |
Expert Syst. Appl. | 5 |
| 2025 | Computational approaches for circRNA-disease association prediction: a reviewabstractAbstract Circular RNA (circRNA) is a covalently closed RNA molecule formed by back splicing. The role of circRNAs in posttranscriptional gene regulation provides new insights into several types of cancer and neurological diseases. CircRNAs are associated with multiple diseases and are emerging biomarkers in cancer diagnosis and treatment. The associations prediction is one of the current research hotspots in the field of bioinformatics. Although research on circRNAs has made great progress, the traditional biological method of verifying circRNA-disease associations is still a great challenge because it is a difficult task and requires much time. Fortunately, advances in computational methods have made considerable progress in circRNA research. This review comprehensively discussed the functions and databases related to circRNA, and then focused on summarizing the calculation model of related predictions, detailed the mainstream algorithm into 4 categories, and analyzed the advantages and limitations of the 4 categories. This not only helps researchers to have overall understanding of circRNA, but also helps researchers have a detailed understanding of the past algorithms, guide new research directions and research purposes to solve the shortcomings of previous research. Mengting Niu, Yaojia Chen, Chunyu Wang 0002, Quan Zou 0001, Lei Xu 0047 |
Frontiers Comput. Sci. | 4 |
| 2025 | APDCA: An accurate and effective method for predicting associations between RBPs and AS-events during epithelial-mesenchymal transitionabstractMOTIVATION: Epithelial-mesenchymal transition (EMT) plays a key role in cancer metastasis by promoting changes in adhesion and motility. RNA-binding proteins (RBPs) regulate alternative splicing (AS) during EMT, enabling a single gene to produce multiple protein isoforms that affect tumor progression. Disruption of RBP-AS interactions may disrupt the progress of diseases like cancer. Despite the importance of RBP-AS relationships in EMT, few computational methods predict these associations. Existing models struggle in sparse settings with limited known associations. To improve performance, we incorporate both sparsity constraints and heterogeneous biological data to infer RBP-AS associations. RESULT: We propose a new method based on Accelerated Proximal DC Algorithm (APDCA) for predicting RBP-AS associations. In particular, APDCA combines sparse low-rank matrix factorization with a Difference-of-Convex (DC) optimization framework and uses extrapolation to improve convergence. A key feature of APDCA is the use of a sparsity constraint, which filters out noise and highlights key associations. In addition, integrating multiple related data sources with direct or indirect relationships can help in reaching a more comprehensive view of RBPs and AS events and to reduce the impact of false positives associated with individual data sources. we prove that our proposed algorithm is convergent under some conditions and the experimental results have illustrated that APDCA outperforms six baseline methods in both AUC and AUPR. A case study on the RBP QKI shows that the top predictions are verified by the OncoSplicing database. Thus, APDCA provides a fast, interpretable, and scalable tool for discovering post-transcriptional regulatory interactions. Yangsong He, Zheng-Jian Bai, Wai-Ki Ching, Quan Zou 0001, Yushan Qiu |
PLoS Comput. Biol. | 4 |
| 2025 | TXSelect: A multi-task learning model to identify secretory effectorsabstractSecretory effectors from pathogenic microorganisms significantly influence pathogen survival and pathogenicity by manipulating host signalling, immune responses, and metabolic processes. However, because of sequence and structural heterogeneity among bacterial effectors, accurately classifying multiple types simultaneously remains challenging. Therefore, we developed TXSelect, a multi-task learning framework that simultaneously classifies TXSE (types I, II, III, IV and VI secretory effectors) using a shared backbone network with task-specific heads. TXSelect integrates the protein embedding features of evolutionary scale modelling (ESM), particularly the N-terminal mean, with classical descriptors to effectively capture complementary information. These descriptors include distance-based residue (DR) and split amino acid composition general (SC-PseAAC-General). Rigorous evaluation identified ESM N-terminal mean + DR + SC-PseAAC as the optimal feature combination, achieving high accuracy (validation F1 = 0.867, test F1 = 0.8645) and robust generalization. Comprehensive assessments and visualization with Uniform Manifold Approximation and Projection further validated the discriminative capability and interpretability of the model. TXSelect provides an efficient computational tool for accurately classifying bacterial effectors, supporting deeper biological understanding and potential therapeutic development. Quan Zou 0001, Chao Zhan |
PLoS Comput. Biol. | 3 |
| 2025 | Identifying the DNA methylation preference of transcription factors using ProtBERT and SVMabstractTranscription factors (TFs) can affect gene expression by binding to certain specific DNA sequences. This binding process of TFs may be modulated by DNA methylation. A subset of TFs that serve as methylation readers preferentially binds to certain methylated DNA and is defined as TFPM. The identification of TFPMs enhances our understanding of DNA methylation's role in gene regulation. However, their experimental identification is resource-demanding. In this study, we propose a novel two-step computational approach to classify TFs and TFPMs. First, we employed a fine-tuned ProtBERT model to differentiate between the classes of TFs and non-TFs. Second, we combined the Reduced Amino Acid Category (RAAC) with K-mer and SVM to predict the potential of TFs to bind to methylated DNA. Comparative experiments demonstrate that our proposed methods outperform all existing approaches and emphasize the efficiency of our computational framework in classifying TFs and TFPMs. Cross-species validation on an independent mouse dataset further demonstrates the generalizability of our proposed framework In addition, we conducted predictions on all human transcription factors and found that most of the top 20 proteins belong to the Krueppel C2H2-type Zinc-finger family. So far, some studies have demonstrated a partial correlation between this family and DNA methylation and confirmed the preference of some of its members, thereby showing the robustness of our approach. Quan Zou 0001, Antony Stalin, Ximei Luo |
PLoS Comput. Biol. | 2 |
| 2025 | Decoupled contrastive multi-view clustering with adaptive false negative elimination for cancer subtypingabstractCancer's heterogeneity necessitates precise subtype identification for effective diagnosis and treatment, which can be achieved by integrating multi-omics data to reveal distinct molecular characteristics and enable personalized therapies. Recently, significant efforts have been made through contrastive clustering methods to efficiently identify cancer subtypes. However, existing approaches remain limited in effectively capturing inter- and intra-view relationships in multi-omics data. Additionally, most cancer subtyping methods often rely on random sampling to construct negative pairs, which may inadvertently engender false negatives. To overcome these challenges, we propose a novel end-to-end self-supervised learning model named Decoupled Contrastive Multi-view Clustering with adaptive false negative elimination (DCMC). Specifically, DCMC adopts a multi-view clustering architecture that facilitates intra- and inter-view contrastive learning across distinct embedding spaces, allowing view-specific information to be preserved while maintaining cross-view consistency. We further introduce an adaptive false negative elimination framework to progressively screen potential false negatives. Finally, pseudo-label rectification is applied to enhance the quality of the learned representations and further refine the clustering process. DCMC is evaluated on 10 commonly used cancer datasets against 19 state-of-the-art methods, with experimental results validating its superior performance. In the Liver Hepatocellular Carcinoma case study, differential expression analysis is performed to identify potential biomarkers, while the cancer subtypes identified by DCMC are validated for their responses to specific therapeutic drugs. The datasets and source code for DCMC are available online at https://github.com/LinMengX/DCMC. Mengxiang Lin, Rongqi Fan, Saisai Zhu, Quan Zou 0001, Zhen Tian 0004 |
PLoS Comput. Biol. | 5 |
| 2025 | PLNMFG: Pseudo-label guided non-negative matrix factorization model with graph constraint for single-cell multi-omics data clusteringabstractThe development of single-cell multi-omics sequencing technologies has enabled the simultaneous analysis of multi-omics data within the same cell. Accurate clustering of these cells is crucial for downstream analyses of complex biological functions. Despite significant advances in multi-omics integration approaches, current methodologies exhibit two major limitations. First, they inadequately incorporate prior biological knowledge from various omic layers. Second, these methods often conduct independent dimensionality reduction on individual omic datasets, thereby failing to capture the intrinsic complementary information and potentially overlooking crucial cross-platform interactions. Motivated by these, this study investigates a non-negative matrix factorization model called PLNMFG, which integrates the unified latent representation learning that retains the features between and within omics and the cluster structure learning that retains the intrinsic structure of the data into one joint framework. Specially, PLNMFG performs adaptive imputation to handle dropout events and uses prior pseudo-labels as constraints during the process of collective non-negative matrix factorization, as a result, a more robust latent representation that preserves the double similarity information is obtained. Graph Laplacian constraint is applied during clustering which further preserves structure characteristic of multi-omics data. In addition, the weight of each omic is adaptively learned based on the omic contribution. A series of experiments on 8 benchmark datasets show that our model performs well in terms of clustering accuracy and computational efficiency. Mingzhu Liu, Yushan Qiu, Wai-Ki Ching, Quan Zou 0001 |
PLoS Comput. Biol. | 5 |
| 2025 | ACP-ESM2: Enhancing Anticancer Peptide Prediction With Pre-Trained Protein Language ModelsabstractAnticancer peptide (ACP) are short peptides with anti-cancer properties that have generated increasing attention in recent years due to their low toxicity, minimal side effects, and their ability to precisely target and kill cancer cells. Traditionally, identifying ACP has relied on experimental methods, which are time-consuming and labor-intensive. While deep learning-based prediction methods have made significant progress, there is still room for improvement in achieving optimal performance. In this study, we present ACP-ESM2, a deep learning framework based on the Evolutionary Scale Modeling 2 (ESM2) pre-trained model, which captures rich evolutionary information from protein sequences. By combining ESM2 with convolutional neural network (CNN) that excels at detecting local patterns, ACP-ESM2 offers a highly accurate tool for ACP prediction. The experimental results indicate that ACP-ESM2 shows significant improvements over best-existing recognition techniques on the Test1 set, with enhancements of 2.3%, 7.2%, 12.6%, and 5% in ACC, SN, SP, and MCC, respectively. Notably, on the Test2 set, ACP-ESM2 achieves an accuracy of 97.6%, showcasing its exceptional robustness. This establishes ACP-ESM2 as an efficient and precise tool for predicting anticancer peptides. Shun Gao, Xingfeng Li 0008, Feifei Cui, Qingchen Zhang 0001, Quan Zou 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 6 |
| 2025 | Kernelized Fuzzy System for Predicting Therapeutic Peptides via Deep Stacked EncoderabstractTherapeutic peptides play a key role in regulating cellular functions and repairing damaged cells through targeted molecular interactions. Traditional wet-lab methods for identifying therapeutic peptides rely on time-consuming biochemical assays and low-throughput screening techniques, which struggle to capture complex sequence-stability relationships critical for drug development. To address these limitations, we innovatively integrate a pretrained protein language model with stacked bidirectional long short-term memory (BiLSTM) encoders. This hybrid architecture enables hierarchical extraction of both global contextual patterns (via the language model) and localized sequential dependencies (via BiLSTM), effectively modeling nonlinear correlations within peptide sequences. A key technical lies in the proposed kernelized Takagi-Sugeno-Kang fuzzy system (K-TSK-FS), which combines fuzzy logic with kernel methods to handle sequence ambiguity while maintaining interpretability. Unlike conventional classifiers, this system maps high-dimensional features into a reproducing kernel Hilbert space, enhancing discrimination between therapeutic and non-therapeutic peptides through nonlinear decision boundaries. To evaluate the model, six benchmark datasets are used to test our model. Experimental results show that our method achieves better prediction performance. Xiaoyi Guo, Yijie Ding, Quan Zou 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2025 | Quality Scores Compression of Genomic Sequencing Data: A Comprehensive Review and Performance EvaluationabstractAdvanced sequencing technologies have profoundly revolutionized biology and produced vast amounts of raw sequencing data during the past decades. The enormous amount of sequencing data proposed significant challenges of data storage and transmission. Compressing a big file into a small file is an encouraging method to tackle these challenges. Howver, it has been found that traditional text data compression algorithms are not well-suited for handing the vast sequencing datasets. Therefore, several algorithms are designed specifically for the efficient compression of sequencing data. Recently, considerable research has been devoted to compressing quality scores stored in the FASTQ format file, resulting in substantial advances in compression performance. Despite these advances, there has been no systematic review and evaluation of these algorithms or software. In this review, we aim to conduct a broad review of the existing quality score compression algorithms. We mainly discuss those algorithms from two categories, i.e., lossless and lossy compression. Additionally, we benchmark the compression performance of 12 tools using 14 real datasets. We anticipate that our review will provide practical guidance for others seeking to design an appropriate algorithm for compressing quality scores. Yuansheng Liu, Zexuan Zhu 0001, Xiangxiang Zeng, Quan Zou 0001, Keqin Li 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2025 | HodgeRankWeight: An Integration Algorithm for Feature Ranking Based on Weight QuantizationabstractThe identification of protein sequences depends on the effective selection of an optimized set of features. Traditional algorithms prioritize global feature importance, often overshadowing the significance of local metrics. Addressing this imbalance, we introduce an innovative algorithm that fuses feature ranking with an advanced weight quantization technique. This algorithm unfolds in two pivotal stages: initially, it generates a weighted directed graph based on normal distribution metrics; subsequently, it employs the HodgeRank algorithm to amalgamate these rankings. Specifically, the algorithm evaluates feature score normality by employing z-scores for skewness and kurtosis, resulting in a graph that quantitatively reflects both local and global feature contributions. HodgeRank then translates this graph into a Laplacian matrix, enabling the calculation of a comprehensive scoring function for each feature. We refine the initial rankings by incorporating weights during the integration phase, capturing a holistic view of feature significance. The proposed method, termed HodgeRankWeight, showcases superior performance, achieving accuracy rates of 87.02%, 92.84%, and 74.51% across different datasets. In head-to-head comparisons, HodgeRankWeight outstripped existing models, achieving an overall accuracy of 82.6923% and setting a new benchmark for precision in protein sequence identification. We also offer a complimentary web server for related research. Chaolu Meng, Yunyun Shi, Quan Zou 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2025 | A Multi-Omics Data Integration Framework for Gene Regulatory Network Inference Based on Contrastive LearningabstractThe Gene Regulatory Networks (GRNs) ensure the stability of cellular states, preserving specific phenotypes and functions throughout the differentiation process. However, current tools still need improvement to effectively integrate multi-omics data and infer GRNs for particular cell types. We introduce CLMOGRI, a multi-omics TF-gene regulatory network inference framework based on heterogeneous networks and contrastive learning, designed to integrate multi-omics data for GRN inference. Through random walk techniques, CLMOGRI embeds multi-omics data into a unified feature space and extracts similar features between nodes. It then measures node similarity and predicts node relationships by contrastive learning. Finally, it includes a regulatory network interpreter to identify critical nodes and modules in GRNs, offering an analytical method for understanding complex interactions within biological systems. CLMOGRI surpasses existing baseline methods in terms of Area Under the Precision-Recall Curve (AUPR) and F-Score metrics, indicating its efficacy in capturing multi-omics information for GRN inference. It also reveals vital nodes and modules within the gene regulatory network, improving the interpretability of CLMOGRI and the utility of GRNs. Yongqing Zhang 0001, Zhigan Zhou, Maocheng Wang, Zixuan Wang 0025, Quan Zou 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 6 |
| 2025 | AEGNN-M:A 3D Graph-Spatial Co-Representation Model for Molecular Property PredictionabstractImproving the drug development process can expedite the introduction of more novel drugs that cater to the demands of precision medicine. Accurately predicting molecular properties remains a fundamental challenge in drug discovery and development. Currently, a plethora of computer-aided drug discovery (CADD) methods have been widely employed in the field of molecular prediction. However, most of these methods primarily analyze molecules using low-dimensional representations such as SMILES notations, molecular fingerprints, and molecular graph-based descriptors. Only a few approaches have focused on incorporating and utilizing high-dimensional spatial structural representations of molecules. In light of the advancements in artificial intelligence, we introduce a 3D graph-spatial co-representation model called AEGNN-M, which combines two graph neural networks, GAT and EGNN. AEGNN-M enables learning of information from both molecular graphs representations and 3D spatial structural representations to predict molecular properties accurately. We conducted experiments on seven public datasets, three regression datasets and 14 breast cancer cell line phenotype screening datasets, comparing the performance of AEGNN-M with state-of-the-art deep learning methods. Extensive experimental results demonstrate the satisfactory performance of the AEGNN-M model. Furthermore, we analyzed the performance impact of different modules within AEGNN-M and the influence of spatial structural representations on the model's performance. The interpretability analysis also revealed the significance of specific atoms in determining particular molecular properties. Xiangzheng Fu, Linlin Zhuo, Quan Zou 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Multi-Modal Deep Representation Learning Accurately Identifies and Interprets Drug-Target InteractionsabstractDeep learning offers efficient solutions for drug-target interaction prediction, but current methods often fail to capture the full complexity of multi-modal data (i.e., sequence, graphs, and three-dimensional structures), limiting both performance and generalization. Here, we present UnitedDTA, a novel explainable deep learning framework capable of integrating multi-modal biomolecule data to improve the binding affinity prediction, especially for novel (unseen) drugs and targets. UnitedDTA enables automatic learning unified discriminative representations from multi-modality data via contrastive learning and cross-attention mechanisms for cross-modality alignment and integration. Comparative results on multiple benchmark datasets show that UnitedDTA significantly outperforms the state-of-the-art drug-target affinity prediction methods and exhibits better generalization ability in predicting unseen drug-target pairs. More importantly, unlike most "black-box" deep learning methods, our well-established model offers better interpretability which enables us to directly infer the important substructures of the drug-target complexes that influence the binding activity, thus providing the insights in unveiling the binding preferences. Moreover, by extending UnitedDTA to other downstream tasks (e.g., molecular property prediction), we showcase the proposed multi-modal representation learning is capable of capturing the latent molecular representations that are closely associated with the molecular property, demonstrating the broad application potential for advancing the drug discovery process. Jiayue Hu, Xiangxiang Zeng, Quan Zou 0001, Ran Su, Leyi Wei |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Trans-MoRFs: A Disordered Protein Predictor Based on the Transformer ArchitectureabstractIntrinsically disordered regions (IDRs) of proteins are crucial for a wide range of biological functions, with molecular recognition features (MoRFs) being of particular significance in protein interactions and cellular regulation. However, the identification of MoRFs has been a significant challenge in computational biology owing to their disorder-to-order transition properties. Currently, only a limited number of experimentally validated MoRFs are known, which has prompted the development of computational methods for predicting MoRFs from protein chains. Considering the limitations of existing MoRF predictors regarding prediction accuracy and adaptability to diverse protein sequence lengths, this study introduces Trans-MoRFs, a novel MoRF predictor based on the transformer architecture, for identifying MoRFs within IDRs of proteins. Trans-MoRFs employ the self-attention mechanism of the transformer to efficiently capture the interactions of distant residues in protein sequences. They demonstrate stability and high efficiency in dealing with protein sequences of different lengths and performs well on both short and long sequences. On multiple benchmark datasets, the model attained a mean area under the curve score of 0.94, which is higher than those of all existing models, and significantly outperformed existing combined and single MoRF prediction tools on multiple performance metrics. Trans-MoRFs have excellent accuracy and a wide range of applications for predicting MoRFs and other functionally important fragments in the disordered regions of proteins. They offer significant assistance in comprehending protein functions, precisely pinpointing functional segments within disordered protein regions and facilitating the discovery of novel drug targets. Chaolu Meng, Yunyun Shi, Xueliang Fu, Quan Zou 0001, Wu Han |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | DSANIB: Drug-Target Interaction Predictions With Dual-View Synergistic Attention Network and Information Bottleneck StrategyabstractPrediction of drug-target interactions (DTIs) is one of the crucial steps for drug repositioning. Identifying DTIs through bio-experimental manners is always expensive and time-consuming. Recently, deep learning-based approaches have shown promising advancements in DTI prediction, but they face two notable challenges: (i) how to explicitly capture local interactions between drug-target pairs and learn their higher-order substructure embeddings; (ii) How to filter out redundant information to obtain effective embeddings for drugs and targets. Results: In this study, we propose a novel approach, termed DSANIB, to infer potential interactions between drugs and targets. DSANIB comprises two primary components: (1) DSAN component: The Inter-view Attention Network Module explicitly learns the local interactions between drugs and targets, while the Intra-view Attention Network Module aggregates information from local interaction features to obtain their higher-order substructure embeddings. (2) Information Bottleneck (IB) component: DSANIB adopts the IB strategy, which could retain relevant information while minimizing the redundant features to obtain their discriminative representations. Extensive experimental results demonstrate that DSANIB outperforms other SOTA prediction models. In addition, visualization of drug and target embeddings learned through DSANIB could provide interpretable insights for the prediction results. Zhen Tian 0004, Wanning Zhou, Zhixia Teng, Quan Zou 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | GCNLA: Inferring Cell-Cell Interactions From Spatial Transcriptomics With Long Short-Term Memory and Graph Convolutional NetworksabstractSpatial transcriptomics analysis methods offer an opportunity to investigate highly diverse biological tissues. Cell-cell communication is fundamental for maintaining physiological homeostasis in organisms and coordinating complex biological processes. Identifying cell-cell interactions is critical for understanding cellular activities. The interaction of a cell with other cells depends on several factors, and most of the existing methods that consider only gene expression information of neighbouring cells and spatial location information are somewhat limited. In this paper, we propose a network architecture based on graph convolution network and long short-term memory attention module-GCNLA, which contains graph convolution layer, long short-term memory network, attention module, and residual connections. GCNLA not only learns the spatial structure of cells but also captures interaction information between distal cells, the attention module further extracting and enhancing features related to cell-cell interactions. Finally, the inner product decoding calculates the cosine similarity, which is used to infer cell-cell interactions. In addition, GCNLA is capable of reconstructing the complete cell-cell interaction network. The experimental results on seqFISH and MERFISH demonstrate that the GCNLA network structure has better robustness and noise immunity. The potential features learned by GCNLA enable other downstream analyses, including single-cell resolution cell clustering based on spatial information resolving cell heterogeneity. Xiuhao Fu, Zhenjie Luo, Leyi Wei, Jingbing Li, Feifei Cui, Quan Zou 0001, Qingchen Zhang 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | Guest Editorial: Large Language Models With Applications in Bioinformatics and Biomedicine
Quan Zou 0001, Limin Jiang, Leyi Wei |
IEEE J. Biomed. Health Informatics | 1 |
| 2025 | Clustering Categorical Data via Multiple Hypothesis TestingabstractCategorical data clustering is a fundamental data mining problem, which has been extensively studied during the past decades. To date, many effective clustering algorithms for categorical data are available in the literature. However, almost all existing categorical data clustering algorithms did not address the issue of the statistical significance of detected clusters. In particular, how to assess the statistical significance of a set of non-overlapping categorical clusters still remains unaddressed. In this article, we formulate the categorical data clustering problem as a multiple hypothesis testing problem, where the null hypothesis is that each attribute is independent of the given partition of clusters. Then, all individual \(p\) -values from different attributes are integrated to obtain a consensus \(p\) -value through statistical meta-analysis. Thereafter, a significance-based clustering algorithm is proposed in which the combined \(p\) -value is efficiently optimized in an indirectly and incremental manner. Experimental results on 25 real-world datasets demonstrate that our method is capable of achieving comparable performance to state-of-the-art categorical data clustering algorithms. Furthermore, our method has a good capability of determining whether there really exists a clustering structure and assessing whether a given set of clusters is statistically significant. Lianyu Hu 0001, Mudi Jiang, Yan Liu 0085, Quan Zou 0001, Zengyou He |
ACM Trans. Knowl. Discov. Data | 4 |
| 2025 | Multiview Deep Learning-Based Molecule Design and Structural Optimization Accelerates Inhibitor DiscoverabstractIn this work, we propose MEDICO, a multiview deep generative model for molecule generation, structural optimization, and the SARS-CoV-2 inhibitor discovery. To the best of our knowledge, MEDICO is the first-of-this-kind graph generative model that can generate molecular graphs similar to the structure of targeted molecules, with a multiview representation learning framework to sufficiently and adaptively learn comprehensive structural semantics from targeted molecular topology and geometry. We show that our MEDICO significantly outperforms the state-of-the-art methods in generating valid, novel, and unique molecules under benchmarking comparisons, particularly achieving $\tilde {8}5 \%$ improvement compared with the state-of-the-art methods in terms of validity. Importantly, we showcase that the multiview deep learning model enables us to generate not only the molecules structurally similar to the targeted molecules but also the molecules with desired chemical properties. Moreover, case study results on targeted molecule generation for the SARS-CoV-2 main protease (Mpro) show that we successfully generate new small molecules with desired drug-like properties for the Mpro by integrating molecular docking into our model as a chemical priori, potentially accelerating the de novo design of COVID-19 drugs. Furthermore, we apply MEDICO to the structural optimization of three well-known Mpro inhibitors (N3, 11a, and GC376) and achieve $\tilde {8}8 \%$ improvement compared with the origin inhibitors in their binding affinity to Mpro, demonstrating the application value of our model for the development of therapeutics for SARS-CoV-2 infection. Ruheng Wang, Quan Zou 0001, Xiangxiang Zeng, Ran Su, Leyi Wei |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | GRACE: Unveiling Gene Regulatory Networks With Causal Mechanistic Graph Neural Networks in Single-Cell RNA-Sequencing DataabstractReconstructing gene regulatory networks (GRNs) using single-cell RNA sequencing (scRNA-seq) data holds great promise for unraveling cellular fate development and heterogeneity. While numerous machine-learning methods have been proposed to infer GRNs from scRNA-seq gene expression data, many of them operate solely in a statistical or black box manner, limiting their capacity for making causal inferences between genes. In this study, we introduce GRN inference with Accuracy and Causal Explanation (GRACE), a novel graph-based causal autoencoder framework that combines a structural causal model (SCM) with graph neural networks (GNNs) to enable GRN inference and gene causal reasoning from scRNA-seq data. By explicitly modeling causal relationships between genes, GRACE facilitates the learning of regulatory context and gene embeddings. With the learned gene signals, our model successfully decoding the causal structures and alleviates the accurate determination of multiple attributes of gene regulation that is important to determine the regulatory levels. Through extensive evaluations on seven benchmarks, we demonstrate that GRACE outperforms 14 state-of-the-art GRN inference methods, with the incorporation of causal mechanisms significantly enhancing the accuracy of GRN and gene causality inference. Furthermore, the application to human peripheral blood mononuclear cell (PBMC) samples reveals cell type-specific regulators in monocyte phagocytosis and immune regulation, validated through network analysis and functional enrichment analysis. Jiacheng Wang 0009, Yaojia Chen, Quan Zou 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Cell-Specific Highly Correlated Network for Self-Supervised Distillation in Cell Type AnnotationabstractSelf-supervised learning has succeeded significantly in cell type annotation based on transcriptomic data. However, existing methods encode transcriptomic data as gene feature representations, limiting self-supervised learning to the gene level. This paper proposes a novel self-supervised distillation learning framework, scTCHCN, which designs a cell-specific highly correlated network. Firstly, this framework integrates the cell-specific highly correlated network with single-cell transcriptomic data to extract cell and gene features. ScTCHCN promotes interactions between the network and the transcriptomic data features during pre-training by constructing the cell-specific highly correlated network. Additionally, a self-supervised paradigm with distillation learning is proposed to enhance the global feature correlation between the two. Ultimately, the pre-trained student model aggregates deep feature constraints and complementary information at the single-cell and gene levels. It combines them with the prediction layer to construct the final cell type prediction model. Experimental results demonstrate that scTCHCN excels in cell type annotation and rare cell identification tasks, showcasing its potential for other applications. The source code and additional information for this study are available at https://github.com/ZhangLab312/scTCHCN. Yugui Xu, Zixuan Wang 0025, Zhigan Zhou, Yongqing Zhang 0001, Quan Zou 0001 |
BIBM | 8 |
| 2024 | RNASite: A one-stop tool website that integrates multiple RNA modification site databases and serversabstractRNASite is a comprehensive platform that integrates multiple RNA modification site databases and servers, focusing on RNA modification sites. These modification sites significantly affect RNA's structure, stability, and function, playing a crucial role in epigenetics and gene expression regulation. The website offers both datasets and online RNA modification site identification tools to enhance the recognition and understanding of these modification sites. With advancements in high-throughput sequencing technology and machine learning, RNASite leverages these innovations to provide efficient and cost-effective RNA modification site identification methods. The platform features 17 high-quality datasets covering 11 common RNA modification sites and includes 7 online identification tools. Notably, some of these tools exceed the accuracy of existing models. RNASite is designed to be a convenient and efficient resource for researchers in biochemistry and bioinformatics, facilitating progress in the study of RNA modification sites. The platform is available for free at http://www.bioai-lab.com/RNASite. Xingfeng Li 0001, Qingchen Zhang 0001, Quan Zou 0001, Feifei Cui |
BIBM | 5 |
| 2024 | iAFPs-Mv-BiTCN: Predicting antifungal peptides using self-attention transformer embedding and transform evolutionary based multi-view features with bidirectional temporal convolutional networks
Shahid Akbar, Quan Zou 0001, Fawaz Khaled Alarfaj |
Artif. Intell. Medicine | 2 |
| 2024 | Adversarial regularized autoencoder graph neural network for microbe-disease associations predictionabstractBACKGROUND: Microorganisms inhabit various regions of the human body and significantly contribute to numerous diseases. Predicting the associations between microbes and diseases is crucial for understanding pathogenic mechanisms and informing prevention and treatment strategies. Biological experiments to determine these associations are time-consuming and costly. Therefore, integrating deep learning with biological networks can efficiently identify potential microbe-disease associations on a large scale. METHODS: We propose an adversarial regularized autoencoder graph neural network algorithm, named Stacked Adversarial Regularization for Microbe-Disease Associations Prediction (SARMDA), for predicting associations between microbes and diseases. First, we integrate topological structural similarity and functional similarity metrics of microbes and diseases to construct a heterogeneous network. Then, utilizing an autoencoder based on GraphSAGE, we learn both the topological and attribute representations of nodes within the constructed network. Finally, we introduce an adversarial regularized autoencoder graph neural network embedding model to address the inherent limitations of traditional GraphSAGE autoencoders in capturing global information. RESULTS: Under the five-fold cross-validation on microbe-disease pairs, SARMDA was compared with eight advanced methods using the Human Microbe-Disease Association Database (HMDAD) and Disbiome databases. The best area under the ROC curve (AUC) achieved by SARMDA on HMDAD was 0.9891$\pm$0.0057, and the best area under the precision-recall curve (AUPR) was 0.9902$\pm$0.0128. On the Disbiome dataset, the AUC was 0.9328$\pm$0.0072, and the best AUPR was 0.9233$\pm$0.0089, outperforming the other eight MDAs prediction methods. Furthermore, the effectiveness of our model was demonstrated through a detailed analysis of asthma and inflammatory bowel disease cases. Limuxuan He, Quan Zou 0001, Shuang Cheng, Yansu Wang |
Briefings Bioinform. | 2 |
| 2024 | FEOpti-ACVP: identification of novel anti-coronavirus peptide sequences based on feature engineering and optimizationabstractAnti-coronavirus peptides (ACVPs) represent a relatively novel approach of inhibiting the adsorption and fusion of the virus with human cells. Several peptide-based inhibitors showed promise as potential therapeutic drug candidates. However, identifying such peptides in laboratory experiments is both costly and time consuming. Therefore, there is growing interest in using computational methods to predict ACVPs. Here, we describe a model for the prediction of ACVPs that is based on the combination of feature engineering (FE) optimization and deep representation learning. FEOpti-ACVP was pre-trained using two feature extraction frameworks. At the next step, several machine learning approaches were tested in to construct the final algorithm. The final version of FEOpti-ACVP outperformed existing methods used for ACVPs prediction and it has the potential to become a valuable tool in ACVP drug design. A user-friendly webserver of FEOpti-ACVP can be accessed at http://servers.aibiochem.net/soft/FEOpti-ACVP/. Jici Jiang, Hongdi Pei, Quan Zou 0001, Zhibin Lv |
Briefings Bioinform. | 5 |
| 2024 | scDFN: enhancing single-cell RNA-seq clustering with deep fusion networksabstractSingle-cell ribonucleic acid sequencing (scRNA-seq) technology can be used to perform high-resolution analysis of the transcriptomes of individual cells. Therefore, its application has gained popularity for accurately analyzing the ever-increasing content of heterogeneous single-cell datasets. Central to interpreting scRNA-seq data is the clustering of cells to decipher transcriptomic diversity and infer cell behavior patterns. However, its complexity necessitates the application of advanced methodologies capable of resolving the inherent heterogeneity and limited gene expression characteristics of single-cell data. Herein, we introduce a novel deep learning-based algorithm for single-cell clustering, designated scDFN, which can significantly enhance the clustering of scRNA-seq data through a fusion network strategy. The scDFN algorithm applies a dual mechanism involving an autoencoder to extract attribute information and an improved graph autoencoder to capture topological nuances, integrated via a cross-network information fusion mechanism complemented by a triple self-supervision strategy. This fusion is optimized through a holistic consideration of four distinct loss functions. A comparative analysis with five leading scRNA-seq clustering methodologies across multiple datasets revealed the superiority of scDFN, as determined by better the Normalized Mutual Information (NMI) and the Adjusted Rand Index (ARI) metrics. Additionally, scDFN demonstrated robust multi-cluster dataset performance and exceptional resilience to batch effects. Ablation studies highlighted the key roles of the autoencoder and the improved graph autoencoder components, along with the critical contribution of the four joint loss functions to the overall efficacy of the algorithm. Through these advancements, scDFN set a new benchmark in single-cell clustering and can be used as an effective tool for the nuanced analysis of single-cell transcriptomics. Cangzhi Jia, Yue Bi, Quan Zou 0001, Fuyi Li |
Briefings Bioinform. | 5 |
| 2024 | Identification, characterization and expression analysis of circRNA encoded by SARS-CoV-1 and SARS-CoV-2abstractVirus-encoded circular RNA (circRNA) participates in the immune response to viral infection, affects the human immune system, and can be used as a target for precision therapy and tumor biomarker. The coronaviruses SARS-CoV-1 and SARS-CoV-2 (SARS-CoV-1/2) that have emerged in recent years are highly contagious and have high mortality rates. In coronaviruses, little is known about the circRNA encoded by the SARS-CoV-1/2. Therefore, this study explores whether SARS-CoV-1/2 encodes circRNA and characteristics and functions of circRNA. Based on RNA-seq data of SARS-CoV-1 and SARS-CoV-2 infections, we used circRNA identification tools (circRNA_finder, find_circ and CIRI2) to identify circRNAs. The number of circRNAs encoded by SARS-CoV-1 and SARS-CoV-2 was identified as 151 and 470, respectively. It can be found that SARS-CoV-2 shows more prominent circRNA encoding ability than SARS-CoV-1. Expression analysis showed that only a few circRNAs encoded by SARS-CoV-1/2 showed high expression levels, and the positive strand produced more abundant circRNAs. Then, based on the identified SARS-CoV-1/2-encoded circRNAs, we performed circRNA identification and characterization using the previously developed CirRNAPL. Finally, target gene prediction and functional enrichment analysis were performed. It was found that viral circRNA is closely related to cancer and has a potential role in regulating host cell functions. This study studied the characteristics and functions of viral circRNA encoded by coronavirus SARS-CoV-1/2, providing a valuable resource for further research on the function and molecular mechanism of coronavirus circRNA. Mengting Niu, Chunyu Wang 0002, Yaojia Chen, Quan Zou 0001, Lei Xu 0047 |
Briefings Bioinform. | 4 |
| 2024 | scMNMF: a novel method for single-cell multi-omics clustering based on matrix factorizationabstractMOTIVATION: The technology for analyzing single-cell multi-omics data has advanced rapidly and has provided comprehensive and accurate cellular information by exploring cell heterogeneity in genomics, transcriptomics, epigenomics, metabolomics and proteomics data. However, because of the high-dimensional and sparse characteristics of single-cell multi-omics data, as well as the limitations of various analysis algorithms, the clustering performance is generally poor. Matrix factorization is an unsupervised, dimensionality reduction-based method that can cluster individuals and discover related omics variables from different blocks. Here, we present a novel algorithm that performs joint dimensionality reduction learning and cell clustering analysis on single-cell multi-omics data using non-negative matrix factorization that we named scMNMF. We formulate the objective function of joint learning as a constrained optimization problem and derive the corresponding iterative formulas through alternating iterative algorithms. The major advantage of the scMNMF algorithm remains its capability to explore hidden related features among omics data. Additionally, the feature selection for dimensionality reduction and cell clustering mutually influence each other iteratively, leading to a more effective discovery of cell types. We validated the performance of the scMNMF algorithm using two simulated and five real datasets. The results show that scMNMF outperformed seven other state-of-the-art algorithms in various measurements. AVAILABILITY AND IMPLEMENTATION: scMNMF code can be found at https://github.com/yushanqiu/scMNMF. Yushan Qiu, Quan Zou 0001 |
Briefings Bioinform. | 4 |
| 2024 | MS-BACL: enhancing metabolic stability prediction through bond graph augmentation and contrastive learningabstractMOTIVATION: Accurately predicting molecular metabolic stability is of great significance to drug research and development, ensuring drug safety and effectiveness. Existing deep learning methods, especially graph neural networks, can reveal the molecular structure of drugs and thus efficiently predict the metabolic stability of molecules. However, most of these methods focus on the message passing between adjacent atoms in the molecular graph, ignoring the relationship between bonds. This makes it difficult for these methods to estimate accurate molecular representations, thereby being limited in molecular metabolic stability prediction tasks. RESULTS: We propose the MS-BACL model based on bond graph augmentation technology and contrastive learning strategy, which can efficiently and reliably predict the metabolic stability of molecules. To our knowledge, this is the first time that bond-to-bond relationships in molecular graph structures have been considered in the task of metabolic stability prediction. We build a bond graph based on 'atom-bond-atom', and the model can simultaneously capture the information of atoms and bonds during the message propagation process. This enhances the model's ability to reveal the internal structure of the molecule, thereby improving the structural representation of the molecule. Furthermore, we perform contrastive learning training based on the molecular graph and its bond graph to learn the final molecular representation. Multiple sets of experimental results on public datasets show that the proposed MS-BACL model outperforms the state-of-the-art model. AVAILABILITY AND IMPLEMENTATION: The code and data are publicly available at https://github.com/taowang11/MS. Zhen Li 0015, Linlin Zhuo, Xiangzheng Fu, Quan Zou 0001 |
Briefings Bioinform. | 6 |
| 2024 | Diff-AMP: tailored designed antimicrobial peptide framework with all-in-one generation, identification, prediction and optimizationabstractAntimicrobial peptides (AMPs), short peptides with diverse functions, effectively target and combat various organisms. The widespread misuse of chemical antibiotics has led to increasing microbial resistance. Due to their low drug resistance and toxicity, AMPs are considered promising substitutes for traditional antibiotics. While existing deep learning technology enhances AMP generation, it also presents certain challenges. Firstly, AMP generation overlooks the complex interdependencies among amino acids. Secondly, current models fail to integrate crucial tasks like screening, attribute prediction and iterative optimization. Consequently, we develop a integrated deep learning framework, Diff-AMP, that automates AMP generation, identification, attribute prediction and iterative optimization. We innovatively integrate kinetic diffusion and attention mechanisms into the reinforcement learning framework for efficient AMP generation. Additionally, our prediction module incorporates pre-training and transfer learning strategies for precise AMP identification and screening. We employ a convolutional neural network for multi-attribute prediction and a reinforcement learning-based iterative optimization strategy to produce diverse AMPs. This framework automates molecule generation, screening, attribute prediction and optimization, thereby advancing AMP research. We have also deployed Diff-AMP on a web server, with code, data and server details available in the Data Availability section. Rui Wang 0168, Linlin Zhuo, Jinhang Wei, Xiangzheng Fu, Quan Zou 0001 |
Briefings Bioinform. | 6 |
| 2024 | A two-task predictor for discovering phase separation proteins and their undergoing mechanismabstractLiquid-liquid phase separation (LLPS) is one of the mechanisms mediating the compartmentalization of macromolecules (proteins and nucleic acids) in cells, forming biomolecular condensates or membraneless organelles. Consequently, the systematic identification of potential LLPS proteins is crucial for understanding the phase separation process and its biological mechanisms. A two-task predictor, Opt_PredLLPS, was developed to discover potential phase separation proteins and further evaluate their mechanism. The first task model of Opt_PredLLPS combines a convolutional neural network (CNN) and bidirectional long short-term memory (BiLSTM) through a fully connected layer, where the CNN utilizes evolutionary information features as input, and BiLSTM utilizes multimodal features as input. If a protein is predicted to be an LLPS protein, it is input into the second task model to predict whether this protein needs to interact with its partners to undergo LLPS. The second task model employs the XGBoost classification algorithm and 37 physicochemical properties following a three-step feature selection. The effectiveness of the model was validated on multiple benchmark datasets, and in silico saturation mutagenesis was used to identify regions that play a key role in phase separation. These findings may assist future research on the LLPS mechanism and the discovery of potential phase separation proteins. Yetong Zhou, Shengming Zhou, Yue Bi, Quan Zou 0001, Cangzhi Jia |
Briefings Bioinform. | 4 |
| 2024 | Joint deep autoencoder and subgraph augmentation for inferring microbial responses to drugsabstractExploring microbial stress responses to drugs is crucial for the advancement of new therapeutic methods. While current artificial intelligence methodologies have expedited our understanding of potential microbial responses to drugs, the models are constrained by the imprecise representation of microbes and drugs. To this end, we combine deep autoencoder and subgraph augmentation technology for the first time to propose a model called JDASA-MRD, which can identify the potential indistinguishable responses of microbes to drugs. In the JDASA-MRD model, we begin by feeding the established similarity matrices of microbe and drug into the deep autoencoder, enabling to extract robust initial features of both microbes and drugs. Subsequently, we employ the MinHash and HyperLogLog algorithms to account intersections and cardinality data between microbe and drug subgraphs, thus deeply extracting the multi-hop neighborhood information of nodes. Finally, by integrating the initial node features with subgraph topological information, we leverage graph neural network technology to predict the microbes' responses to drugs, offering a more effective solution to the 'over-smoothing' challenge. Comparative analyses on multiple public datasets confirm that the JDASA-MRD model's performance surpasses that of current state-of-the-art models. This research aims to offer a more profound insight into the adaptability of microbes to drugs and to furnish pivotal guidance for drug treatment strategies. Our data and code are publicly available at: https://github.com/ZZCrazy00/JDASA-MRD. Zhecheng Zhou, Linlin Zhuo, Xiangzheng Fu, Quan Zou 0001 |
Briefings Bioinform. | 4 |
| 2024 | HAlign 4: a new strategy for rapidly aligning millions of sequencesabstractMOTIVATION: HAlign is a high-performance multiple sequence alignment software based on the star alignment strategy, which is the preferred choice for rapidly aligning large numbers of sequences. HAlign3, implemented in Java, is the latest version capable of aligning an ultra-large number of similar DNA/RNA sequences. However, HAlign3 still struggles with long sequences and extremely large numbers of sequences. RESULTS: To address this issue, we have implemented HAlign4 in C++. In this version, we replaced the original suffix tree with Burrows-Wheeler Transform and introduced the wavefront alignment algorithm to further optimize both time and memory efficiency. Experiments show that HAlign4 significantly outperforms HAlign3 in runtime and memory usage in both single-threaded and multi-threaded configurations, while maintains high alignment accuracy comparable to MAFFT. HAlign4 can complete the alignment of 10 million coronavirus disease 2019 (COVID-19) sequences in about 12 min and 300 GB of memory using 96 threads, demonstrating its efficiency and practicality for large-scale alignment on standard workstations. AVAILABILITY AND IMPLEMENTATION: Source code is available at https://github.com/malabz/HAlign-4, dataset is available at https://zenodo.org/records/13934503. Tong Zhou 0016, Pinglu Zhang, Quan Zou 0001, Wu Han |
Bioinform. | 3 |
| 2024 | CircSI-SSL: circRNA-binding site identification based on self-supervised learningabstractMOTIVATION: In recent years, circular RNAs (circRNAs), the particular form of RNA with a closed-loop structure, have attracted widespread attention due to their physiological significance (they can directly bind proteins), leading to the development of numerous protein site identification algorithms. Unfortunately, these studies are supervised and require the vast majority of labeled samples in training to produce superior performance. But the acquisition of sample labels requires a large number of biological experiments and is difficult to obtain. RESULTS: To resolve this matter that a great deal of tags need to be trained in the circRNA-binding site prediction task, a self-supervised learning binding site identification algorithm named CircSI-SSL is proposed in this article. According to the survey, this is unprecedented in the research field. Specifically, CircSI-SSL initially combines multiple feature coding schemes and employs RNA_Transformer for cross-view sequence prediction (self-supervised task) to learn mutual information from the multi-view data, and then fine-tuning with only a few sample labels. Comprehensive experiments on six widely used circRNA datasets indicate that our CircSI-SSL algorithm achieves excellent performance in comparison to previous algorithms, even in the extreme case where the ratio of training data to test data is 1:9. In addition, the transplantation experiment of six linRNA datasets without network modification and hyperparameter adjustment shows that CircSI-SSL has good scalability. In summary, the prediction algorithm based on self-supervised learning proposed in this article is expected to replace previous supervised algorithms and has more extensive application value. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at https://github.com/cc646201081/CircSI-SSL. Chunyu Wang 0002, Shuhong Yang, Quan Zou 0001 |
Bioinform. | 4 |
| 2024 | Integrated convolution and self-attention for improving peptide toxicity predictionabstractMOTIVATION: Peptides are promising agents for the treatment of a variety of diseases due to their specificity and efficacy. However, the development of peptide-based drugs is often hindered by the potential toxicity of peptides, which poses a significant barrier to their clinical application. Traditional experimental methods for evaluating peptide toxicity are time-consuming and costly, making the development process inefficient. Therefore, there is an urgent need for computational tools specifically designed to predict peptide toxicity accurately and rapidly, facilitating the identification of safe peptide candidates for drug development. RESULTS: We provide here a novel computational approach, CAPTP, which leverages the power of convolutional and self-attention to enhance the prediction of peptide toxicity from amino acid sequences. CAPTP demonstrates outstanding performance, achieving a Matthews correlation coefficient of approximately 0.82 in both cross-validation settings and on independent test datasets. This performance surpasses that of existing state-of-the-art peptide toxicity predictors. Importantly, CAPTP maintains its robustness and generalizability even when dealing with data imbalances. Further analysis by CAPTP reveals that certain sequential patterns, particularly in the head and central regions of peptides, are crucial in determining their toxicity. This insight can significantly inform and guide the design of safer peptide drugs. AVAILABILITY AND IMPLEMENTATION: The source code for CAPTP is freely available at https://github.com/jiaoshihu/CAPTP. Shihu Jiao, Xiucai Ye, Tetsuya Sakurai, Quan Zou 0001 |
Bioinform. | 4 |
| 2024 | GraphADT: empowering interpretable predictions of acute dermal toxicity with multi-view graph pooling and structure remappingabstractMOTIVATION: Accurate prediction of acute dermal toxicity (ADT) is essential for the safe and effective development of contact drugs. Currently, graph neural networks, a form of deep learning technology, accurately model the structure of compound molecules, enhancing predictions of their ADT. However, many existing methods emphasize atom-level information transfer and overlook crucial data conveyed by molecular bonds and their interrelationships. Additionally, these methods often generate "equal" node representations across the entire graph, failing to accentuate "important" substructures like functional groups, pharmacophores, and toxicophores, thereby reducing interpretability. RESULTS: We introduce a novel model, GraphADT, utilizing structure remapping and multi-view graph pooling (MVPool) technologies to accurately predict compound ADT. Initially, our model applies structure remapping to better delineate bonds, transforming "bonds" into new nodes and "bond-atom-bond" interactions into new edges, thereby reconstructing the compound molecular graph. Subsequently, we use MVPool to amalgamate data from various perspectives, minimizing biases inherent to single-view analyses. Following this, the model generates a robust node ranking collaboratively, emphasizing critical nodes or substructures to enhance model interpretability. Lastly, we apply a graph comparison learning strategy to train both the original and structure remapped molecular graphs, deriving the final molecular representation. Experimental results on public datasets indicate that the GraphADT model outperforms existing state-of-the-art models. The GraphADT model has been demonstrated to effectively predict compound ADT, offering potential guidance for the development of contact drugs and related treatments. AVAILABILITY AND IMPLEMENTATION: Our code and data are accessible at: https://github.com/mxqmxqmxq/GraphADT.git. Xinqian Ma, Xiangzheng Fu, Linlin Zhuo, Quan Zou 0001 |
Bioinform. | 5 |
| 2024 | Identifying nucleotide-binding leucine-rich repeat receptor and pathogen effector pairing using transfer-learning and bilinear attention networkabstractMOTIVATION: Nucleotide-binding leucine-rich repeat (NLR) family is a class of immune receptors capable of detecting and defending against pathogen invasion. They have been widely used in crop breeding. Notably, the correspondence between NLRs and effectors (CNE) determines the applicability and effectiveness of NLRs. Unfortunately, CNE data is very scarce. In fact, we've found a substantial 91 291 NLRs confirmed via wet experiments and bioinformatics methods but only 387 CNEs are recognized, which greatly restricts the potential application of NLRs. RESULTS: We propose a deep learning algorithm called ProNEP to identify NLR-effector pairs in a high-throughput manner. Specifically, we conceptualized the CNE prediction task as a protein-protein interaction (PPI) prediction task. Then, ProNEP predicts the interaction between NLRs and effectors by combining the transfer learning with a bilinear attention network. ProNEP achieves superior performance against state-of-the-art models designed for PPI predictions. Based on ProNEP, we conduct extensive identification of potential CNEs for 91 291 NLRs. With the rapid accumulation of genomic data, we expect that this tool will be widely used to predict CNEs in new species, advancing biology, immunology, and breeding. AVAILABILITY AND IMPLEMENTATION: The ProNEP is available at http://nerrd.cn/#/prediction. The project code is available at https://github.com/QiaoYJYJ/ProNEP. Baixue Qiao, Shuda Wang, Mingjun Hou, Haodi Chen, Zhengwenyang Zhou, Xueying Xie, Shaozi Pang, Chunxue Yang, Fenglong Yang, Quan Zou 0001, Shanwen Sun |
Bioinform. | 10 |
| 2024 | scTPC: a novel semisupervised deep clustering model for scRNA-seq dataabstractMOTIVATION: Continuous advancements in single-cell RNA sequencing (scRNA-seq) technology have enabled researchers to further explore the study of cell heterogeneity, trajectory inference, identification of rare cell types, and neurology. Accurate scRNA-seq data clustering is crucial in single-cell sequencing data analysis. However, the high dimensionality, sparsity, and presence of "false" zero values in the data can pose challenges to clustering. Furthermore, current unsupervised clustering algorithms have not effectively leveraged prior biological knowledge, making cell clustering even more challenging. RESULTS: This study investigates a semisupervised clustering model called scTPC, which integrates the triplet constraint, pairwise constraint, and cross-entropy constraint based on deep learning. Specifically, the model begins by pretraining a denoising autoencoder based on a zero-inflated negative binomial distribution. Deep clustering is then performed in the learned latent feature space using triplet constraints and pairwise constraints generated from partial labeled cells. Finally, to address imbalanced cell-type datasets, a weighted cross-entropy loss is introduced to optimize the model. A series of experimental results on 10 real scRNA-seq datasets and five simulated datasets demonstrate that scTPC achieves accurate clustering with a well-designed framework. AVAILABILITY AND IMPLEMENTATION: scTPC is a Python-based algorithm, and the code is available from https://github.com/LF-Yang/Code or https://zenodo.org/records/10951780. Yushan Qiu, Lingfei Yang, Hao Jiang 0009, Quan Zou 0001 |
Bioinform. | 4 |
| 2024 | Drug-target interaction predictions with multi-view similarity network fusion strategy and deep interactive attention mechanismabstractMOTIVATION: Accurately identifying the drug-target interactions (DTIs) is one of the crucial steps in the drug discovery and drug repositioning process. Currently, many computational-based models have already been proposed for DTI prediction and achieved some significant improvement. However, these approaches pay little attention to fuse the multi-view similarity networks related to drugs and targets in an appropriate way. Besides, how to fully incorporate the known interaction relationships to accurately represent drugs and targets is not well investigated. Therefore, there is still a need to improve the accuracy of DTI prediction models. RESULTS: In this study, we propose a novel approach that employs Multi-view similarity network fusion strategy and deep Interactive attention mechanism to predict Drug-Target Interactions (MIDTI). First, MIDTI constructs multi-view similarity networks of drugs and targets with their diverse information and integrates these similarity networks effectively in an unsupervised manner. Then, MIDTI obtains the embeddings of drugs and targets from multi-type networks simultaneously. After that, MIDTI adopts the deep interactive attention mechanism to further learn their discriminative embeddings comprehensively with the known DTI relationships. Finally, we feed the learned representations of drugs and targets to the multilayer perceptron model and predict the underlying interactions. Extensive results indicate that MIDTI significantly outperforms other baseline methods on the DTI prediction task. The results of the ablation experiments also confirm the effectiveness of the attention mechanism in the multi-view similarity network fusion strategy and the deep interactive attention mechanism. AVAILABILITY AND IMPLEMENTATION: https://github.com/XuLew/MIDTI. Lewen Xu, Chenguang Han, Zhen Tian 0004, Quan Zou 0001 |
Bioinform. | 5 |
| 2024 | DeepAVP-TPPred: identification of antiviral peptides using transformed image-based localized descriptors and binary tree growth algorithmabstractMOTIVATION: Despite the extensive manufacturing of antiviral drugs and vaccination, viral infections continue to be a major human ailment. Antiviral peptides (AVPs) have emerged as potential candidates in the pursuit of novel antiviral drugs. These peptides show vigorous antiviral activity against a diverse range of viruses by targeting different phases of the viral life cycle. Therefore, the accurate prediction of AVPs is an essential yet challenging task. Lately, many machine learning-based approaches have developed for this purpose; however, their limited capabilities in terms of feature engineering, accuracy, and generalization make these methods restricted. RESULTS: In the present study, we aim to develop an efficient machine learning-based approach for the identification of AVPs, referred to as DeepAVP-TPPred, to address the aforementioned problems. First, we extract two new transformed feature sets using our designed image-based feature extraction algorithms and integrate them with an evolutionary information-based feature. Next, these feature sets were optimized using a novel feature selection approach called binary tree growth Algorithm. Finally, the optimal feature space from the training dataset was fed to the deep neural network to build the final classification model. The proposed model DeepAVP-TPPred was tested using stringent 5-fold cross-validation and two independent dataset testing methods, which achieved the maximum performance and showed enhanced efficiency over existing predictors in terms of both accuracy and generalization capabilities. AVAILABILITY AND IMPLEMENTATION: https://github.com/MateeullahKhan/DeepAVP-TPPred. Matee Ullah, Shahid Akbar, Quan Zou 0001 |
Bioinform. | 4 |
| 2024 | FMAlign2: a novel fast multiple nucleotide sequence alignment method for ultralong datasetsabstractMOTIVATION: In bioinformatics, multiple sequence alignment (MSA) is a crucial task. However, conventional methods often struggle with aligning ultralong sequences. To address this issue, researchers have designed MSA methods rooted in a vertical division strategy, which segments sequence data for parallel alignment. A prime example of this approach is FMAlign, which utilizes the FM-index to extract common seeds and segment the sequences accordingly. RESULTS: FMAlign2 leverages the suffix array to identify maximal exact matches, redefining the approach of FMAlign from searching for global chains to partial chains. By using a vertical division strategy, large-scale problem is deconstructed into manageable tasks, enabling parallel execution of subMSA. Furthermore, sequence-profile alignment and refinement are incorporated to concatenate subsets, yielding the final result seamlessly. Compared to FMAlign, FMAlign2 markedly augments the segmentation of sequences and significantly reduces the time while maintaining accuracy, especially on ultralong datasets. Importantly, FMAlign2 enhances existing MSA methods by conferring the capability to handle sequences reaching billions in length within an acceptable time frame. AVAILABILITY AND IMPLEMENTATION: Source code and datasets are available at https://github.com/malabz/FMAlign2 and https://zenodo.org/records/10435770. Pinglu Zhang, Huan Liu 0024, Yanming Wei, Yixiao Zhai, Qinzhong Tian, Quan Zou 0001 |
Bioinform. | 6 |
| 2024 | Revisiting drug-protein interaction prediction: a novel global-local perspectiveabstractMOTIVATION: Accurate inference of potential drug-protein interactions (DPIs) aids in understanding drug mechanisms and developing novel treatments. Existing deep learning models, however, struggle with accurate node representation in DPI prediction, limiting their performance. RESULTS: We propose a new computational framework that integrates global and local features of nodes in the drug-protein bipartite graph for efficient DPI inference. Initially, we employ pre-trained models to acquire fundamental knowledge of drugs and proteins and to determine their initial features. Subsequently, the MinHash and HyperLogLog algorithms are utilized to estimate the similarity and set cardinality between drug and protein subgraphs, serving as their local features. Then, an energy-constrained diffusion mechanism is integrated into the transformer architecture, capturing interdependencies between nodes in the drug-protein bipartite graph and extracting their global features. Finally, we fuse the local and global features of nodes and employ multilayer perceptrons to predict the likelihood of potential DPIs. A comprehensive and precise node representation guarantees efficient prediction of unknown DPIs by the model. Various experiments validate the accuracy and reliability of our model, with molecular docking results revealing its capability to identify potential DPIs not present in existing databases. This approach is expected to offer valuable insights for furthering drug repurposing and personalized medicine research. AVAILABILITY AND IMPLEMENTATION: Our code and data are accessible at: https://github.com/ZZCrazy00/DPI. Zhecheng Zhou, Qingquan Liao, Jinhang Wei, Linlin Zhuo, Xiaonan Wu, Xiangzheng Fu, Quan Zou 0001 |
Bioinform. | 7 |
| 2024 | Deepstacked-AVPs: predicting antiviral peptides using tri-segment evolutionary profile and word embedding based multi-perspective features with deep stacking modelabstractBACKGROUND: Viral infections have been the main health issue in the last decade. Antiviral peptides (AVPs) are a subclass of antimicrobial peptides (AMPs) with substantial potential to protect the human body against various viral diseases. However, there has been significant production of antiviral vaccines and medications. Recently, the development of AVPs as an antiviral agent suggests an effective way to treat virus-affected cells. Recently, the involvement of intelligent machine learning techniques for developing peptide-based therapeutic agents is becoming an increasing interest due to its significant outcomes. The existing wet-laboratory-based drugs are expensive, time-consuming, and cannot effectively perform in screening and predicting the targeted motif of antiviral peptides. METHODS: In this paper, we proposed a novel computational model called Deepstacked-AVPs to discriminate AVPs accurately. The training sequences are numerically encoded using a novel Tri-segmentation-based position-specific scoring matrix (PSSM-TS) and word2vec-based semantic features. Composition/Transition/Distribution-Transition (CTDT) is also employed to represent the physiochemical properties based on structural features. Apart from these, the fused vector is formed using PSSM-TS features, semantic information, and CTDT descriptors to compensate for the limitations of single encoding methods. Information gain (IG) is applied to choose the optimal feature set. The selected features are trained using a stacked-ensemble classifier. RESULTS: The proposed Deepstacked-AVPs model achieved a predictive accuracy of 96.60%%, an area under the curve (AUC) of 0.98, and a precision-recall (PR) value of 0.97 using training samples. In the case of the independent samples, our model obtained an accuracy of 95.15%, an AUC of 0.97, and a PR value of 0.97. CONCLUSION: Our Deepstacked-AVPs model outperformed existing models with a ~ 4% and ~ 2% higher accuracy using training and independent samples, respectively. The reliability and efficacy of the proposed Deepstacked-AVPs model make it a valuable tool for scientists and may perform a beneficial role in pharmaceutical design and research academia. Shahid Akbar, Quan Zou 0001 |
BMC Bioinform. | 3 |
| 2024 | StackedEnC-AOP: prediction of antioxidant proteins using transform evolutionary and sequential features based multi-scale vector with stacked ensemble learningabstractBACKGROUND: Antioxidant proteins are involved in several biological processes and can protect DNA and cells from the damage of free radicals. These proteins regulate the body's oxidative stress and perform a significant role in many antioxidant-based drugs. The current invitro-based medications are costly, time-consuming, and unable to efficiently screen and identify the targeted motif of antioxidant proteins. METHODS: In this model, we proposed an accurate prediction method to discriminate antioxidant proteins namely StackedEnC-AOP. The training sequences are formulation encoded via incorporating a discrete wavelet transform (DWT) into the evolutionary matrix to decompose the PSSM-based images via two levels of DWT to form a Pseudo position-specific scoring matrix (PsePSSM-DWT) based embedded vector. Additionally, the Evolutionary difference formula and composite physiochemical properties methods are also employed to collect the structural and sequential descriptors. Then the combined vector of sequential features, evolutionary descriptors, and physiochemical properties is produced to cover the flaws of individual encoding schemes. To reduce the computational cost of the combined features vector, the optimal features are chosen using Minimum redundancy and maximum relevance (mRMR). The optimal feature vector is trained using a stacking-based ensemble meta-model. RESULTS: Our developed StackedEnC-AOP method reported a prediction accuracy of 98.40% and an AUC of 0.99 via training sequences. To evaluate model validation, the StackedEnC-AOP training model using an independent set achieved an accuracy of 96.92% and an AUC of 0.98. CONCLUSION: Our proposed StackedEnC-AOP strategy performed significantly better than current computational models with a ~ 5% and ~ 3% improved accuracy via training and independent sets, respectively. The efficacy and consistency of our proposed StackedEnC-AOP make it a valuable tool for data scientists and can execute a key role in research academia and drug design. Gul Rukh, Shahid Akbar, Gauhar Rehman, Fawaz Khaled Alarfaj, Quan Zou 0001 |
BMC Bioinform. | 5 |
| 2024 | SBSM-Pro: support bio-sequence machine for proteins
Yizheng Wang, Yixiao Zhai, Yijie Ding, Quan Zou 0001 |
Sci. China Inf. Sci. | 4 |
| 2024 | Identification of human microRNA-disease association via low-rank approximation-based link propagation and multiple kernel learning
Yizheng Wang, Xin Zhang 0103, Ying Ju 0002, Quan Zou 0001, Yazhou Zhang 0001, Yijie Ding, Ying Zhang 0060 |
Frontiers Comput. Sci. | 5 |
| 2024 | Multi-source data integration for explainable miRNA-driven drug discovery
Zhen Li 0015, Qingquan Liao, Peng Xu 0004, Linlin Zhuo, Xiangzheng Fu, Quan Zou 0001 |
Future Gener. Comput. Syst. | 7 |
| 2024 | Random subsequence forests
Zengyou He, Mudi Jiang, Lianyu Hu 0001, Quan Zou 0001 |
Inf. Sci. | 5 |
| 2024 | AMDGT: Attention aware multi-modal fusion using a dual graph transformer for drug-disease associations predictionabstractIdentification of new indications for existing drugs is crucial through the various stages of drug discovery. Computational methods are valuable in establishing meaningful associations between drugs and diseases. However, most methods predict the drug-disease associations based solely on similarity data, neglecting valuable biological and chemical information. These methods often use basic concatenation to integrate information from different modalities, limiting their ability to capture features from a comprehensive and in-depth perspective. Therefore, a novel multimodal framework called AMDGT was proposed to predict new drug associations based on dual-graph transformer modules. By combining similarity data and complex biochemical information, AMDGT understands the multimodal feature fusion of drugs and diseases effectively and comprehensively with an attention-aware modality interaction architecture. Extensive experimental results indicate that AMDGT surpasses state-of-the-art methods in real-world datasets. Moreover, case and molecular docking studies demonstrated that AMDGT is an effective tool for drug repositioning. Our code is available at GitHub: https://github.com/JK-Liu7/AMDGT. Quan Zou 0001, Hongjie Wu, Prayag Tiwari, Yijie Ding |
Knowl. Based Syst. | 3 |
| 2024 | Sequence homology score-based deep fuzzy network for identifying therapeutic peptidesabstractThe detection of therapeutic peptides is a topic of immense interest in the biomedical field. Conventional biochemical experiment-based detection techniques are tedious and time-consuming. Computational biology has become a useful tool for improving the detection efficiency of therapeutic peptides. Most computational methods do not consider the deviation caused by noise. To improve the generalization performance of therapeutic peptide prediction methods, this work presents a sequence homology score-based deep fuzzy echo-state network with maximizing mixture correntropy (SHS-DFESN-MMC) model. Our method is compared with the existing methods on eight types of therapeutic peptide datasets. The model parameters are determined by 10 fold cross-validation on their training sets and verified by independent test sets. Across the 8 datasets, the average area under the receiver operating characteristic curve (AUC) values of SHS-DFESN-MMC are the highest on both the training (0.926) and independent sets (0.923). Xiaoyi Guo, Ziyu Zheng, Kang Hao Cheong, Quan Zou 0001, Prayag Tiwari, Yijie Ding |
Neural Networks | 4 |
| 2024 | AttentionMGT-DTA: A multi-modal drug-target affinity prediction using graph transformer and attention mechanismabstractThe accurate prediction of drug-target affinity (DTA) is a crucial step in drug discovery and design. Traditional experiments are very expensive and time-consuming. Recently, deep learning methods have achieved notable performance improvements in DTA prediction. However, one challenge for deep learning-based models is appropriate and accurate representations of drugs and targets, especially the lack of effective exploration of target representations. Another challenge is how to comprehensively capture the interaction information between different instances, which is also important for predicting DTA. In this study, we propose AttentionMGT-DTA, a multi-modal attention-based model for DTA prediction. AttentionMGT-DTA represents drugs and targets by a molecular graph and binding pocket graph, respectively. Two attention mechanisms are adopted to integrate and interact information between different protein modalities and drug-target pairs. The experimental results showed that our proposed model outperformed state-of-the-art baselines on two benchmark datasets. In addition, AttentionMGT-DTA also had high interpretability by modeling the interaction strength between drug atoms and protein residues. Our code is available at https://github.com/JK-Liu7/AttentionMGT-DTA. Hongjie Wu, Tengsheng Jiang, Quan Zou 0001, Shujie Qi, Zhiming Cui 0002, Prayag Tiwari, Yijie Ding |
Neural Networks | 4 |
| 2024 | AutoEdge-CCP: A novel approach for predicting cancer-associated circRNAs and drugs based on automated edge embeddingabstractThe unique expression patterns of circRNAs linked to the advancement and prognosis of cancer underscore their considerable potential as valuable biomarkers. Repurposing existing drugs for new indications can significantly reduce the cost of cancer treatment. Computational prediction of circRNA-cancer and drug-cancer relationships is crucial for precise cancer therapy. However, prior computational methods fail to analyze the interaction between circRNAs, drugs, and cancer at the systematic level. It is essential to propose a method that uncover more valuable information for achieving cancer-centered multi-association prediction. In this paper, we present a novel computational method, AutoEdge-CCP, to unveil cancer-associated circRNAs and drugs. We abstract the complex relationships between circRNAs, drugs, and cancer into a multi-source heterogeneous network. In this network, each molecule is represented by two types information, one is the intrinsic attribute information of molecular features, and the other is the link information explicitly modeled by autoGNN, which searches information from both intra-layer and inter-layer of message passing neural network. The significant performance on multi-scenario applications and case studies establishes AutoEdge-CCP as a potent and promising association prediction tool. Yaojia Chen, Jiacheng Wang 0009, Chunyu Wang 0002, Quan Zou 0001 |
PLoS Comput. Biol. | 4 |
| 2024 | MVST: Identifying spatial domains of spatial transcriptomes from multiple views using multi-view graph convolutional networksabstractSpatial transcriptome technology can parse transcriptomic data at the spatial level to detect high-throughput gene expression and preserve information regarding the spatial structure of tissues. Identifying spatial domains, that is identifying regions with similarities in gene expression and histology, is the most basic and critical aspect of spatial transcriptome data analysis. Most current methods identify spatial domains only through a single view, which may obscure certain important information and thus fail to make full use of the information embedded in spatial transcriptome data. Therefore, we propose an unsupervised clustering framework based on multiview graph convolutional networks (MVST) to achieve accurate spatial domain recognition by the learning graph embedding features of neighborhood graphs constructed from gene expression information, spatial location information, and histopathological image information through multiview graph convolutional networks. By exploring spatial transcriptomes from multiple views, MVST enables data from all parts of the spatial transcriptome to be comprehensively and fully utilized to obtain more accurate spatial expression patterns. We verified the effectiveness of MVST on real spatial transcriptome datasets, the robustness of MVST on some simulated datasets, and the reasonableness of the framework structure of MVST in ablation experiments, and from the experimental results, it is clear that MVST can achieve a more accurate spatial domain identification compared with the current more advanced methods. In conclusion, MVST is a powerful tool for spatial transcriptome research with improved spatial domain recognition. Qingchen Zhang 0001, Feifei Cui, Quan Zou 0001 |
PLoS Comput. Biol. | 4 |
| 2024 | scRNMF: An imputation method for single-cell RNA-seq data by robust and non-negative matrix factorizationabstractSingle-cell RNA sequencing (scRNA-seq) has emerged as a powerful tool in genomics research, enabling the analysis of gene expression at the individual cell level. However, scRNA-seq data often suffer from a high rate of dropouts, where certain genes fail to be detected in specific cells due to technical limitations. This missing data can introduce biases and hinder downstream analysis. To overcome this challenge, the development of effective imputation methods has become crucial in the field of scRNA-seq data analysis. Here, we propose an imputation method based on robust and non-negative matrix factorization (scRNMF). Instead of other matrix factorization algorithms, scRNMF integrates two loss functions: L2 loss and C-loss. The L2 loss function is highly sensitive to outliers, which can introduce substantial errors. We utilize the C-loss function when dealing with zero values in the raw data. The primary advantage of the C-loss function is that it imposes a smaller punishment for larger errors, which results in more robust factorization when handling outliers. Various datasets of different sizes and zero rates are used to evaluate the performance of scRNMF against other state-of-the-art methods. Our method demonstrates its power and stability as a tool for imputation of scRNA-seq data. Yuqing Qian, Quan Zou 0001, Yi Liu 0112, Fei Guo 0001, Yijie Ding |
PLoS Comput. Biol. | 2 |
| 2024 | MFPSP: Identification of fungal species-specific phosphorylation site using offspring competition-based genetic algorithmabstractProtein phosphorylation is essential in various signal transduction and cellular processes. To date, most tools are designed for model organisms, but only a handful of methods are suitable for predicting task in fungal species, and their performance still leaves much to be desired. In this study, a novel tool called MFPSP is developed for phosphorylation site prediction in multi-fungal species. The amino acids sequence features were derived from physicochemical and distributed information, and an offspring competition-based genetic algorithm was applied for choosing the most effective feature subset. The comparison results shown that MFPSP achieves a more advanced and balanced performance to several state-of-the-art available toolkits. Feature contribution and interaction exploration indicating the proposed model is efficient in uncovering concealed patterns within sequence. We anticipate MFPSP to serve as a valuable bioinformatics tool and benefiting practical experiments by pre-screening potential phosphorylation sites and enhancing our functional understanding of phosphorylation modifications in fungi. The source code and datasets are accessible at https://github.com/AI4HKB/MFPSP/. Chao Wang 0043, Quan Zou 0001 |
PLoS Comput. Biol. | 2 |
| 2024 | ECD-CDGI: An efficient energy-constrained diffusion model for cancer driver gene identificationabstractThe identification of cancer driver genes (CDGs) poses challenges due to the intricate interdependencies among genes and the influence of measurement errors and noise. We propose a novel energy-constrained diffusion (ECD)-based model for identifying CDGs, termed ECD-CDGI. This model is the first to design an ECD-Attention encoder by combining the ECD technique with an attention mechanism. ECD-Attention encoder excels at generating robust gene representations that reveal the complex interdependencies among genes while reducing the impact of data noise. We concatenate topological embedding extracted from gene-gene networks through graph transformers to these gene representations. We conduct extensive experiments across three testing scenarios. Extensive experiments show that the ECD-CDGI model possesses the ability to not only be proficient in identifying known CDGs but also efficiently uncover unknown potential CDGs. Furthermore, compared to the GNN-based approach, the ECD-CDGI model exhibits fewer constraints by existing gene-gene networks, thereby enhancing its capability to identify CDGs. Additionally, ECD-CDGI is open-source and freely available. We have also launched the model as a complimentary online tool specifically crafted to expedite research efforts focused on CDGs identification. Linlin Zhuo, Xiangzheng Fu, Xiangxiang Zeng, Quan Zou 0001 |
PLoS Comput. Biol. | 6 |
| 2024 | Overcoming CRISPR-Cas9 off-target prediction hurdles: A novel approach with ESB rebalancing strategy and CRISPR-MCA modelabstractThe off-target activities within the CRISPR-Cas9 system remains a formidable barrier to its broader application and development. Recent advancements have highlighted the potential of deep learning models in predicting these off-target effects, yet they encounter significant hurdles including imbalances within datasets and the intricacies associated with encoding schemes and model architectures. To surmount these challenges, our study innovatively introduces an Efficiency and Specificity-Based (ESB) class rebalancing strategy, specifically devised for datasets featuring mismatches-only off-target instances, marking a pioneering approach in this realm. Furthermore, through a meticulous evaluation of various One-hot encoding schemes alongside numerous hybrid neural network models, we discern that encoding and models of moderate complexity ideally balance performance and efficiency. On this foundation, we advance a novel hybrid model, the CRISPR-MCA, which capitalizes on multi-feature extraction to enhance predictive accuracy. The empirical results affirm that the ESB class rebalancing strategy surpasses five conventional methods in addressing extreme dataset imbalances, demonstrating superior efficacy and broader applicability across diverse models. Notably, the CRISPR-MCA model excels in off-target effect prediction across four distinct mismatches-only datasets and significantly outperforms contemporary state-of-the-art models in datasets comprising both mismatches and indels. In summation, the CRISPR-MCA model, coupled with the ESB rebalancing strategy, offers profound insights and a robust framework for future explorations in this field. Yanpeng Yang, Yanyi Zheng, Quan Zou 0001, Jian Li 0032, Hailin Feng |
PLoS Comput. Biol. | 3 |
| 2024 | TPMA: A two pointers meta-alignment tool to ensemble different multiple nucleic acid sequence alignmentsabstractAccurate multiple sequence alignment (MSA) is imperative for the comprehensive analysis of biological sequences. However, a notable challenge arises as no single MSA tool consistently outperforms its counterparts across diverse datasets. Users often have to try multiple MSA tools to achieve optimal alignment results, which can be time-consuming and memory-intensive. While the overall accuracy of certain MSA results may be lower, there could be local regions with the highest alignment scores, prompting researchers to seek a tool capable of merging these locally optimal results from multiple initial alignments into a globally optimal alignment. In this study, we introduce Two Pointers Meta-Alignment (TPMA), a novel tool designed for the integration of nucleic acid sequence alignments. TPMA employs two pointers to partition the initial alignments into blocks containing identical sequence fragments. It selects blocks with the high sum of pairs (SP) scores to concatenate them into an alignment with an overall SP score superior to that of the initial alignments. Through tests on simulated and real datasets, the experimental results consistently demonstrate that TPMA outperforms M-Coffee in terms of aSP, Q, and total column (TC) scores across most datasets. Even in cases where TPMA's scores are comparable to M-Coffee, TPMA exhibits significantly lower running time and memory consumption. Furthermore, we comprehensively assessed all the MSA tools used in the experiments, considering accuracy, time, and memory consumption. We propose accurate and fast combination strategies for small and large datasets, which streamline the user tool selection process and facilitate large-scale dataset integration. The dataset and source code of TPMA are available on GitHub (https://github.com/malabz/TPMA). Yixiao Zhai, Jiannan Chao, Yizheng Wang, Pinglu Zhang, Furong Tang, Quan Zou 0001 |
PLoS Comput. Biol. | 6 |
| 2024 | PseU-KeMRF: A Novel Method for Identifying RNA Pseudouridine SitesabstractPseudouridine is a type of abundant RNA modification that is seen in many different animals and is crucial for a variety of biological functions. Accurately identifying pseudouridine sites within the RNA sequence is vital for the subsequent study of various biological mechanisms of pseudouridine. However, the use of traditional experimental methods faces certain challenges. The development of fast and convenient computational methods is necessary to accurately identify pseudouridine sites from RNA sequence information. To address this, we introduce a novel pseudouridine site prediction model called PseU-KeMRF, which can identify pseudouridine sites in three species, H. sapiens, S. cerevisiae, and M. musculus. Through comprehensive analysis, we selected four RNA coding schemes, including binary feature, position-specific trinucleotide propensity based on single strand (PSTNPss), nucleotide chemical property (NCP) and pseudo k-tuple composition (PseKNC). Then the support vector machine-recursive feature elimination (SVM-RFE) method was used for feature selection and the feature subset was optimized. Finally, the best feature subsets are input into the kernel based on multinomial random forests (KeMRF) classifier for cross-validation and independent testing. As a new classification method, compared with the traditional random forest, KeMRF not only improves the node splitting process of decision tree construction based on multinomial distribution, but also combines the easy to interpret kernel method for prediction, which makes the classification performance better. Our results indicate superior predictive performance of PseU-KeMRF over other existing models, which can prove that PseU-KeMRF is a highly competitive predictive model that can successfully identify pseudouridine sites in RNA sequences. Mingshuai Chen, Quan Zou 0001, Ren Qi, Yijie Ding |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2024 | Prediction of Potential miRNA-Disease Associations Based on a Masked Graph AutoencoderabstractBiomedical evidence has demonstrated the relevance of microRNA (miRNA) dysregulation in complex human diseases, and determining the relationship between miRNAs and diseases can aid in the early detection and prevention of diseases. Traditional biological experimental methods have the disadvantages of high cost and low efficiency, which are well compensated by computational methods. However, many computational methods have the challenge of excessively focusing on the neighbor relationship, ignoring the structural information of the graph, and belittling the redundant information of the graph structure. This study proposed a computational model based on a graph-masking autoencoder named MGAEMDA. MGAEMDA is an asymmetric framework in which the encoder maps partially observed graphs into latent representations. The decoder reconstructs the masked structural information based on the edge and node levels and combines it with linear matrices to obtain the result. The empirical results on the two datasets reveal that the MGAEMDA model performs better than its counterparts. We also demonstrated the predictive performance of MGAEMDA using a case study of four diseases, and all the top 30 predicted miRNAs were validated in the database, providing further evidence of the excellent performance of the model. Hailin Feng, Chenchen Ke, Quan Zou 0001, Zhechen Zhu, Tongcun Liu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2024 | Hyb_SEnc: An Antituberculosis Peptide Predictor Based on a Hybrid Feature Vector and Stacked Ensemble LearningabstractTuberculosis has plagued mankind since ancient times, and the struggle between humans and tuberculosis continues. Mycobacterium tuberculosis is the leading cause of tuberculosis, infecting nearly one-third of the world's population. The rise of peptide drugs has created a new direction in the treatment of tuberculosis. Therefore, for the treatment of tuberculosis, the prediction of anti-tuberculosis peptides is crucial. This paper proposes an anti-tuberculosis peptide prediction method based on hybrid features and stacked ensemble learning. First, a random forest (RF) and extremely randomized tree (ERT) are selected as first-level learning of stacked ensembles. Then, the five best-performing feature encoding methods are selected to obtain the hybrid feature vector, and then the decision tree and recursive feature elimination (DT-RFE) are used to refine the hybrid feature vector. After selection, the optimal feature subset is used as the input of the stacked ensemble model. At the same time, logistic regression (LR) is used as a stacked ensemble secondary learner to build the final stacked ensemble model Hyb_SEnc. The prediction accuracy of Hyb_SEnc achieved 94.68% and 95.74% on the independent test sets of AntiTb_MD and AntiTb_RD, respectively. Xiuhao Fu, Xiaofeng Zang, Xingfeng Li 0001, Qingchen Zhang 0001, Quan Zou 0001, Feifei Cui |
IEEE ACM Trans. Comput. Biol. Bioinform. | 8 |
| 2024 | AGML: Adaptive Graph-Based Multi-Label Learning for Prediction of RBP and as Event Associations During EMTabstractIncreasing evidence has indicated that RNA-binding proteins (RBPs) play an essential role in mediating alternative splicing (AS) events during epithelial-mesenchymal transition (EMT). However, due to the substantial cost and complexity of biological experiments, how AS events are regulated and influenced remains largely unknown. Thus, it is important to construct effective models for inferring hidden RBP-AS event associations during EMT process. In this paper, a novel and efficient model was developed to identify AS event-related candidate RBPs based on Adaptive Graph-based Multi-Label learning (AGML). In particular, we propose to adaptively learn a new affinity graph to capture the intrinsic structure of data for both RBPs and AS events. Multi-view similarity matrices are employed for maintaining the intrinsic structure and guiding the adaptive graph learning. We then simultaneously update the RBP and AS event associations that are predicted from both spaces by applying multi-label learning. The experimental results have shown that our AGML achieved AUC values of 0.9521 and 0.9873 by 5-fold and leave-one-out cross-validations, respectively, indicating the superiority and effectiveness of our proposed model. Furthermore, AGML can serve as an efficient and reliable tool for uncovering novel AS events-associated RBPs and is applicable for predicting the associations between other biological entities. Yushan Qiu, Wai-Ki Ching, Hongmin Cai, Hao Jiang 0009, Quan Zou 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2024 | Fuzzy Neural Tangent Kernel Model for Identifying DNA N4-Methylcytosine SitesabstractDNA N4-methylcytosine (4mC) site identification is a crucial field in bioinformatics, where machine learning methods have been effectively utilized. Due to the presence of noise, the existing deep learning methods for detecting 4mC have consistently low recognition rates in positive samples. With fuzzy rules and membership functions, fuzzy systems can achieve good results in processing noisy signals. In contrast to traditional fuzzy systems that lack deep feature representation and sample measurement, we introduce novel techniques to enhance generalization and feature representation. By incorporating the neural tangent kernel (NTK) and kernel learning algorithm into the fuzzy system, we propose the fuzzy NTK (FNTK) model and the radius-based FNTK (R-FNTK) model to predict DNA 4mC sites. To achieve better generalization performance than traditional kernel functions, we first train the NTK for feature representation learning and sample measurement. Based on the membership function and NTK matrix, different fuzzy kernel matrices are constructed for each fuzzy subset of the fuzzy system. Finally, we utilize two types of iterative kernel optimization algorithms to effectively fuse multiple NTK-based fuzzy kernels and obtain the final prediction model. Rigorous testing using six benchmark datasets demonstrates the superiority of our approach, yielding significant improvements in the experiment's performance. Yijie Ding, Prayag Tiwari, Fei Guo 0001, Quan Zou 0001, Weiping Ding 0001 |
IEEE Trans. Fuzzy Syst. | 4 |
| 2024 | iProps: A Comprehensive Software Tool for Protein Classification and Analysis With Automatic Machine Learning Capabilities and Model Interpretation CapabilitiesabstractProtein classification is a crucial field in bioinformatics. The development of a comprehensive tool that can perform feature evaluation, visualization, automated machine learning, and model interpretation would significantly advance research in protein classification. However, there is a significant gap in the literature regarding tools that integrate all these essential functionalities. This paper presents iProps, a novel Python-based software package, meticulously crafted to fulfill these multifaceted requirements. iProps is distinguished by its proficiency in feature extraction, evaluation, automated machine learning, and interpretation of classification models. Firstly, iProps fully leverages evolutionary information and amino acid reduction information to propose or extend several numerical protein features that are independent of sequence length, including SC-PSSM, ORDip, TRC, CTDC-E, CKSAAGP-E, and so forth; at the same time, it also implements the calculation of 17 other numerical features within the software. iProps also provides feature combination operations for the aforementioned features to generate more hybrid features, and has added data balancing sampling processing as well as built-in classifier settings, among other functionalities. Thus, It can discern the most effective protein class recognition feature from a multitude of candidates, utilizing three automated machine learning algorithms to identify the most optimal classifiers and parameter settings. Furthermore, iProps generates a detailed explanatory report that includes 23 informative graphs derived from three interpretable models. To assess the performance of iProps, a series of numerical experiments were conducted using two well-established datasets. The results demonstrated that our software achieved superior recognition performance in every case. Beyond its contributions to bioinformatics, iProps broadens its applicability by offering robust data analysis tools that are beneficial across various disciplines, capitalizing on its automated machine learning and model interpretation capabilities. As an open-source platform, iProps is readily accessible and features an intuitive user interface, ensuring ease of use for individuals, even those without a background in programming. Changli Feng, Haiyan Wei, Chugui Xu, Bin Feng 0002, Xiaorong Zhu, Jing Liu 0068, Quan Zou 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2024 | MAMLCDA: A Meta-Learning Model for Predicting circRNA-Disease Association Based on MAML Combined With CNNabstractCircular RNAs (circRNAs) exist in vivo and are a class of noncoding RNA molecules. They have a single-stranded, closed, annular structure. Many studies have shown that circRNAs and diseases are linked. Therefore, it is critical to build a reliable and accurate predictor to find the circRNA-disease association. In this paper, we presented a meta-learning model named MAMLCDA to identify the circRNA-disease association, which is based on model-agnostic meta-learning (MAML) combined with CNN classification. Specifically, similarities between diseases and circRNAs are extracted and integrated to characterize their relationships, and k-means is used to cluster majority samples and select a certain number of samples from each cluster to obtain the same number of negative samples as the positive samples. To further reduce the dimension of the features and save operation time, we applied probabilistic principal component analysis (PPCA) to compact the integrated circRNA and disease similarity network feature vectors. The feature vectors are converted into images. At this time, the prediction problem is transformed into the 2-way 1-shot problem of the image and input into the model with MAML as the meta-learner and CNN as the base-learner. Comparison results of five-fold cross-validation on two benchmark datasets illustrate that MAMLCDA outperforms several state-of-the-art approaches with the best accuracies of 95.33% and 98%. Therefore, MAMLCDA can help to understand the pathogenesis of complex diseases at the circRNA level. Yuanyi Tian, Quan Zou 0001, Chunyu Wang 0002, Cangzhi Jia |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Exploring Parameter-Efficient Fine-Tuning of a Large-Scale Pre-Trained Model for scRNA-seq Cell Type AnnotationabstractAccurate identification of cell types is a pivotal and intricate task in scRNA-seq data analysis. Recently, significant strides have been made in cell type annotation of scRNA-seq data using pre-trained language models (PLMs). This method has surmounted the constraints of conventional approaches regarding precision, robustness, and generalization. However, the fine-tuning process of large-scale pre-trained models incurs substantial computational expenses. To tackle this issue, a promising avenue of research has emerged, proposing parameter-efficient fine-tuning techniques for PLMs. These techniques concentrate on fine-tuning only a small portion of the model parameters while attaining comparable performance. In this study, we extensively research parameter-efficient fine-tuning methods for scRNA-seq cell type annotation, employing scBERT as the backbone. We scrutinize the performance and compatibility of various parameter-efficient fine-tuning methodologies across multiple datasets. Through comprehensive analysis, we demonstrate the remarkable performance of parameter-efficient fine-tuning methods in cell type annotation. Hopefully, this study can inspire new thinking in analyzing scRNA-seq data. Zixuan Wang 0025, Guiquan Zhu, Yongqing Zhang 0001, Quan Zou 0001 |
BIBM | 6 |
| 2023 | Human-Spa: An Online Platform Based on Spatial Transcriptome Data for Diseases of Human SystemsabstractSpatial transcriptomics has become a major method for high-throughput analysis of gene expression at the current level of cells, which can directly study gene expression changes in disease cells, reveal the occurrence and development mechanism of diseases, identify potential therapeutic targets related to diseases, and provide new clues for disease diagnosis and treatment. Although there have been major breakthroughs in the analysis and acquisition of transcriptome data, there are still many challenges in how to effectively manage, share and utilize these valuable data resources in the study of human systemic diseases. To solve this problem, we are committed to building a spatial transcriptome database website related to human systemic diseases, aiming to provide reliable datasets related to multiple diseases, and provide certain data information and analysis results to accelerate the progress of human systemic diseases research. This paper will introduce the construction process of the database website in detail, and discuss its application prospects in disease research, with the aim of promoting further development and innovation in the field of human health. Here, Human-spa mainly includes 12 Human systems, 38 disease types, 55 datasets, and Human-Spa provides a very friendly web interface for visualization and dataset parsing. In conclusion, the construction of Human-Spa will provide a powerful tool and resource platform for Human disease research. Human-Spa is available for free at http://www.human-spa.cn/ Yunyun Su, Feifei Cui, Shiyu Yan, Quan Zou 0001, Chen Cao 0002 |
BIBM | 4 |
| 2023 | KDProg: A Knowledge distillation graph neural network for cancer prognosis prediction and analysisabstractAccurately predicting cancer prognosis remains challenging, owing to the combination of computational and practical challenges. This study proposes KDProg, a knowledge distillation-based graph learning framework for predicting cancer prognosis and exploring downstream tasks. The framework includes a novel feature distillation paradigm that compresses a multi-layer complex teacher model to a single-layer simple student by using the teacher model’s middle-layer feature representations and outputs as supervision information to improve the student model’s performance. In addition, instead of introducing a unified temperature hyperparameter, KDProg adopts a novel strategy to parameterize the distillation temperature and combine it with the Cox partial log-likelihood function. So the model can learn the appropriate temperature. Furthermore, considering multi-omics data of patients are often complex to obtain in practical cancer prognosis, this paper uses different input data for the teacher and student models, respectively. The input data for the teacher model are multi-omics data (mRNA, CNV, and DNA methylation), clinical data, and KEGG pathways. The input data for the student model are mRNA, clinical data, and KEGG pathways. Extensive experiments on 15 real-world datasets from TCGA demonstrated the effectiveness and efficiency of the proposed method in predicting cancer prognosis. The results suggest that the proposed model can guide clinical decision-making. Shuwen Xiong, Zixuan Wang 0025, Yongqing Zhang 0001, Quan Zou 0001 |
BIBM | 6 |
| 2023 | HGTDG: An Interpretable Heterogeneous Graph Transformer Framework for Cancer Driver Gene PredictionabstractAccurately predicting cancer driver genes remains challenging due to the increasing size and complexity of cancer genomic data. In this study, HGTDG is proposed, a heterogeneous graph transformer framework for predicting cancer driver genes and exploring downstream tasks. The framework includes a heterogeneous graph construction module that constructs a gene-protein heterogeneous network based on KEGG pathways and the protein-protein interactions from the STRING database. In addition, the framework introduces a novel heterogeneous graph transformer module that uses multi-head attention mechanisms for gene node embedding. The transformer module can capture dedicated representations for genes and edges. Finally, the generated gene embeddings are fed into the classification module to classify genes into driver and non-driver genes. The experiment results show that HGTDG outperforms the state-of-the-art methods regarding the area under the receiver operating characteristic curves (AUROC) and the area under the precision-recall curves (AUPRC). Shuwen Xiong, Zixuan Wang 0025, Guiquan Zhu, Yongqing Zhang 0001, Quan Zou 0001 |
BIBM | 6 |
| 2023 | WMSA 2: a multiple DNA/RNA sequence alignment tool implemented with accurate progressive mode and a fast win-win mode combining the center star and progressive strategiesabstractMultiple sequence alignment is widely used for sequence analysis, such as identifying important sites and phylogenetic analysis. Traditional methods, such as progressive alignment, are time-consuming. To address this issue, we introduce StarTree, a novel method to fast construct a guide tree by combining sequence clustering and hierarchical clustering. Furthermore, we develop a new heuristic similar region detection algorithm using the FM-index and apply the k-banded dynamic program to the profile alignment. We also introduce a win-win alignment algorithm that applies the central star strategy within the clusters to fast the alignment process, then uses the progressive strategy to align the central-aligned profiles, guaranteeing the final alignment's accuracy. We present WMSA 2 based on these improvements and compare the speed and accuracy with other popular methods. The results show that the guide tree made by the StarTree clustering method can lead to better accuracy than that of PartTree while consuming less time and memory than that of UPGMA and mBed methods on datasets with thousands of sequences. During the alignment of simulated data sets, WMSA 2 can consume less time and memory while ranking at the top of Q and TC scores. The WMSA 2 is still better at the time, and memory efficiency on the real datasets and ranks at the top on the average sum of pairs score. For the alignment of 1 million SARS-CoV-2 genomes, the win-win mode of WMSA 2 significantly decreased the consumption time than the former version. The source code and data are available at https://github.com/malabz/WMSA2. Jiannan Chao, Huan Liu 0024, Fenglong Yang, Quan Zou 0001, Furong Tang |
Briefings Bioinform. | 5 |
| 2023 | Matrix reconstruction with reliable neighbors for predicting potential MiRNA-disease associationsabstractNumerous experimental studies have indicated that alteration and dysregulation in mircroRNAs (miRNAs) are associated with serious diseases. Identifying disease-related miRNAs is therefore an essential and challenging task in bioinformatics research. Computational methods are an efficient and economical alternative to conventional biomedical studies and can reveal underlying miRNA-disease associations for subsequent experimental confirmation with reasonable confidence. Despite the success of existing computational approaches, most of them only rely on the known miRNA-disease associations to predict associations without adding other data to increase the prediction accuracy, and they are affected by issues of data sparsity. In this paper, we present MRRN, a model that combines matrix reconstruction with node reliability to predict probable miRNA-disease associations. In MRRN, the most reliable neighbors of miRNA and disease are used to update the original miRNA-disease association matrix, which significantly reduces data sparsity. Unknown miRNA-disease associations are reconstructed by aggregating the most reliable first-order neighbors to increase prediction accuracy by representing the local and global structure of the heterogeneous network. Five-fold cross-validation of MRRN produced an area under the curve (AUC) of 0.9355 and area under the precision-recall curve (AUPR) of 0.2646, values that were greater than those produced by comparable models. Two different types of case studies using three diseases were conducted to demonstrate the accuracy of MRRN, and all top 30 predicted miRNAs were verified. Hailin Feng, Dongdong Jin, Jian Li 0032, Yane Li, Quan Zou 0001, Tongcun Liu |
Briefings Bioinform. | 5 |
| 2023 | Dimensionality reduction and visualization of single-cell RNA-seq data with an improved deep variational autoencoderabstractSingle-cell RNA sequencing (scRNA-seq) is a revolutionary breakthrough that determines the precise gene expressions on individual cells and deciphers cell heterogeneity and subpopulations. However, scRNA-seq data are much noisier than traditional high-throughput RNA-seq data because of technical limitations, leading to many scRNA-seq data studies about dimensionality reduction and visualization remaining at the basic data-stacking stage. In this study, we propose an improved variational autoencoder model (termed DREAM) for dimensionality reduction and a visual analysis of scRNA-seq data. Here, DREAM combines the variational autoencoder and Gaussian mixture model for cell type identification, meanwhile explicitly solving 'dropout' events by introducing the zero-inflated layer to obtain the low-dimensional representation that describes the changes in the original scRNA-seq dataset. Benchmarking comparisons across nine scRNA-seq datasets show that DREAM outperforms four state-of-the-art methods on average. Moreover, we prove that DREAM can accurately capture the expression dynamics of human preimplantation embryonic development. DREAM is implemented in Python, freely available via the GitHub website, https://github.com/Crystal-JJ/DREAM. Junlin Xu, Yuansheng Liu, Bosheng Song, Xiulan Guo, Xiangxiang Zeng, Quan Zou 0001 |
Briefings Bioinform. | 7 |
| 2023 | CoraL: interpretable contrastive meta-learning for the prediction of cancer-associated ncRNA-encoded small peptidesabstractNcRNA-encoded small peptides (ncPEPs) have recently emerged as promising targets and biomarkers for cancer immunotherapy. Therefore, identifying cancer-associated ncPEPs is crucial for cancer research. In this work, we propose CoraL, a novel supervised contrastive meta-learning framework for predicting cancer-associated ncPEPs. Specifically, the proposed meta-learning strategy enables our model to learn meta-knowledge from different types of peptides and train a promising predictive model even with few labeled samples. The results show that our model is capable of making high-confidence predictions on unseen cancer biomarkers with only five samples, potentially accelerating the discovery of novel cancer biomarkers for immunotherapy. Moreover, our approach remarkably outperforms existing deep learning models on 15 cancer-associated ncPEPs datasets, demonstrating its effectiveness and robustness. Interestingly, our model exhibits outstanding performance when extended for the identification of short open reading frames derived from ncPEPs, demonstrating the strong prediction ability of CoraL at the transcriptome level. Importantly, our feature interpretation analysis discovers unique sequential patterns as the fingerprint for each cancer-associated ncPEPs, revealing the relationship among certain cancer biomarkers that are validated by relevant literature and motif comparison. Overall, we expect CoraL to be a useful tool to decipher the pathogenesis of cancer and provide valuable information for cancer research. The dataset and source code of our proposed method can be found at https://github.com/Johnsunnn/CoraL. Zhongshen Li, Junru Jin, Wentao Long, Haoqing Yu, Xin Gao 0001, Kenta Nakai, Quan Zou 0001, Leyi Wei |
Briefings Bioinform. | 8 |
| 2023 | SSNMDI: a novel joint learning model of semi-supervised non-negative matrix factorization and data imputation for clustering of single-cell RNA-seq dataabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) technology attracts extensive attention in the biomedical field. It can be used to measure gene expression and analyze the transcriptome at the single-cell level, enabling the identification of cell types based on unsupervised clustering. Data imputation and dimension reduction are conducted before clustering because scRNA-seq has a high 'dropout' rate, noise and linear inseparability. However, independence of dimension reduction, imputation and clustering cannot fully characterize the pattern of the scRNA-seq data, resulting in poor clustering performance. Herein, we propose a novel and accurate algorithm, SSNMDI, that utilizes a joint learning approach to simultaneously perform imputation, dimensionality reduction and cell clustering in a non-negative matrix factorization (NMF) framework. In addition, we integrate the cell annotation as prior information, then transform the joint learning into a semi-supervised NMF model. Through experiments on 14 datasets, we demonstrate that SSNMDI has a faster convergence speed, better dimensionality reduction performance and a more accurate cell clustering performance than previous methods, providing an accurate and robust strategy for analyzing scRNA-seq data. Biological analysis are also conducted to validate the biological significance of our method, including pseudotime analysis, gene ontology and survival analysis. We believe that we are among the first to introduce imputation, partial label information, dimension reduction and clustering to the single-cell field. AVAILABILITY AND IMPLEMENTATION: The source code for SSNMDI is available at https://github.com/yushanqiu/SSNMDI. Yushan Qiu, Chang Yan, Quan Zou 0001 |
Briefings Bioinform. | 4 |
| 2023 | MMiKG: a knowledge graph-based platform for path mining of microbiota-mental diseases interactionsabstractThe microbiota-gut-brain axis denotes a two-way system of interactions between the gut and the brain, comprising three key components: (1) gut microbiota, (2) intermediates and (3) mental ailments. These constituents communicate with one another to induce changes in the host's mood, cognition and demeanor. Knowledge concerning the regulation of the host central nervous system by gut microbiota is fragmented and mostly confined to disorganized or semi-structured unrestricted texts. Such a format hinders the exploration and comprehension of unknown territories or the further advancement of artificial intelligence systems. Hence, we collated crucial information by scrutinizing an extensive body of literature, amalgamated the extant knowledge of the microbiota-gut-brain axis and depicted it in the form of a knowledge graph named MMiKG, which can be visualized on the GraphXR platform and the Neo4j database, correspondingly. By merging various associated resources and deducing prospective connections between gut microbiota and the central nervous system through MMiKG, users can acquire a more comprehensive perception of the pathogenesis of mental disorders and generate novel insights for advancing therapeutic measures. As a free and open-source platform, MMiKG can be accessed at http://yangbiolab.cn:8501/ with no login requirement. Zhaoqi Song, Qiuming Chen, Furong Tang, Lijun Dou, Quan Zou 0001, Fenglong Yang |
Briefings Bioinform. | 7 |
| 2023 | Adaptive learning embedding features to improve the predictive performance of SARS-CoV-2 phosphorylation sitesabstractMOTIVATION: The rapid and extensive transmission of the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) has led to an unprecedented global health emergency, affecting millions of people and causing an immense socioeconomic impact. The identification of SARS-CoV-2 phosphorylation sites plays an important role in unraveling the complex molecular mechanisms behind infection and the resulting alterations in host cell pathways. However, currently available prediction tools for identifying these sites lack accuracy and efficiency. RESULTS: In this study, we presented a comprehensive biological function analysis of SARS-CoV-2 infection in a clonal human lung epithelial A549 cell, revealing dramatic changes in protein phosphorylation pathways in host cells. Moreover, a novel deep learning predictor called PSPred-ALE is specifically designed to identify phosphorylation sites in human host cells that are infected with SARS-CoV-2. The key idea of PSPred-ALE lies in the use of a self-adaptive learning embedding algorithm, which enables the automatic extraction of context sequential features from protein sequences. In addition, the tool uses multihead attention module that enables the capturing of global information, further improving the accuracy of predictions. Comparative analysis of features demonstrated that the self-adaptive learning embedding features are superior to hand-crafted statistical features in capturing discriminative sequence information. Benchmarking comparison shows that PSPred-ALE outperforms the state-of-the-art prediction tools and achieves robust performance. Therefore, the proposed model can effectively identify phosphorylation sites assistant the biomedical scientists in understanding the mechanism of phosphorylation in SARS-CoV-2 infection. AVAILABILITY AND IMPLEMENTATION: PSPred-ALE is available at https://github.com/jiaoshihu/PSPred-ALE and Zenodo (https://doi.org/10.5281/zenodo.8330277). Shihu Jiao, Xiucai Ye, Chunyan Ao, Tetsuya Sakurai, Quan Zou 0001, Lei Xu 0047 |
Bioinform. | 5 |
| 2023 | Optimization of drug-target affinity prediction methods through feature processing schemesabstractMOTIVATION: Numerous high-accuracy drug-target affinity (DTA) prediction models, whose performance is heavily reliant on the drug and target feature information, are developed at the expense of complexity and interpretability. Feature extraction and optimization constitute a critical step that significantly influences the enhancement of model performance, robustness, and interpretability. Many existing studies aim to comprehensively characterize drugs and targets by extracting features from multiple perspectives; however, this approach has drawbacks: (i) an abundance of redundant or noisy features; and (ii) the feature sets often suffer from high dimensionality. RESULTS: In this study, to obtain a model with high accuracy and strong interpretability, we utilize various traditional and cutting-edge feature selection and dimensionality reduction techniques to process self-associated features and adjacent associated features. These optimized features are then fed into learning to rank to achieve efficient DTA prediction. Extensive experimental results on two commonly used datasets indicate that, among various feature optimization methods, the regression tree-based feature selection method is most beneficial for constructing models with good performance and strong robustness. Then, by utilizing Shapley Additive Explanations values and the incremental feature selection approach, we obtain that the high-quality feature subset consists of the top 150D features and the top 20D features have a breakthrough impact on the DTA prediction. In conclusion, our study thoroughly validates the importance of feature optimization in DTA prediction and serves as inspiration for constructing high-performance and high-interpretable models. AVAILABILITY AND IMPLEMENTATION: https://github.com/RUXIAOQING964914140/FS_DTA. Xiaoqing Ru, Quan Zou 0001, Chen Lin 0001 |
Bioinform. | 2 |
| 2023 | Predicting active enhancers with DNA methylation and histone modificationabstractBACKGROUND: Enhancers play a crucial role in gene regulation, and some active enhancers produce noncoding RNAs known as enhancer RNAs (eRNAs) bi-directionally. The most commonly used method for detecting eRNAs is CAGE-seq, but the instability of eRNAs in vivo leads to data noise in sequencing results. Unfortunately, there is currently a lack of research focused on the noise inherent in CAGE-seq data, and few approaches have been developed for predicting eRNAs. Bridging this gap and developing widely applicable eRNA prediction models is of utmost importance. RESULTS: In this study, we proposed a method to reduce false positives in the identification of eRNAs by adjusting the statistical distribution of expression levels. We also developed eRNA prediction models using joint gene expressions, DNA methylation, and histone modification. These models achieved impressive performance with an AUC value of approximately 0.95 for intra-cell prediction and 0.9 for cross-cell prediction. CONCLUSIONS: Our method effectively attenuates the noise generated by stochastic RNA production, resulting in more accurate detection of eRNAs. Furthermore, our eRNA prediction model exhibited significant accuracy in both intra-cell and cross-cell validation, highlighting its robustness and potential application in various cellular contexts. Ximei Luo, Yan Liu 0085, Quan Zou 0001, Ying Zhang 0060, Lei Xu 0047 |
BMC Bioinform. | 5 |
| 2023 | Subspace projection-based weighted echo state networks for predicting therapeutic peptidesabstractDetection of therapeutic peptide is a major research direction in the current biopharmaceutical field. However, traditional biochemical experimental detection methods take a lot of time. As supplementary methods for biochemical experiments, the computational methods can improve the efficiency of therapeutic peptide detection. Currently, most machine learning-based therapeutic peptide identification algorithms do not consider the processing of noisy samples. We propose a therapeutic peptide classifier, called weighted echo state networks based on subspace projection (WESN-SP), which reduces the bias caused by high-dimensional noisy features and noisy samples. WESN-SP is trained by sparse Bayesian learning algorithm (SBL) and introduces a weight coefficient for each sample by kernel dependence maximization-based subspace projection. The experimental results show that WESN-SP has better performance than other existing methods. Xiaoyi Guo, Prayag Tiwari, Quan Zou 0001, Yijie Ding |
Knowl. Based Syst. | 3 |
| 2023 | Recall DNA methylation levels at low coverage sites using a CNN model in WGBSabstractDNA methylation is an important regulator of gene transcription. WGBS is the gold-standard approach for base-pair resolution quantitative of DNA methylation. It requires high sequencing depth. Many CpG sites with insufficient coverage in the WGBS data, resulting in inaccurate DNA methylation levels of individual sites. Many state-of-arts computation methods were proposed to predict the missing value. However, many methods required either other omics datasets or other cross-sample data. And most of them only predicted the state of DNA methylation. In this study, we proposed the RcWGBS, which can impute the missing (or low coverage) values from the DNA methylation levels on the adjacent sides. Deep learning techniques were employed for the accurate prediction. The WGBS datasets of H1-hESC and GM12878 were down-sampled. The average difference between the DNA methylation level at 12× depth predicted by RcWGBS and that at >50× depth in the H1-hESC and GM2878 cells are less than 0.03 and 0.01, respectively. RcWGBS performed better than METHimpute even though the sequencing depth was as low as 12×. Our work would help to process methylation data of low sequencing depth. It is beneficial for researchers to save sequencing costs and improve data utilization through computational methods. Ximei Luo, Yansu Wang, Quan Zou 0001, Lei Xu 0047 |
PLoS Comput. Biol. | 3 |
| 2023 | PCB: A pseudotemporal causality-based Bayesian approach to identify EMT-associated regulatory relationships of AS events and RBPs during breast cancer progressionabstractDuring breast cancer metastasis, the developmental process epithelial-mesenchymal (EM) transition is abnormally activated. Transcriptional regulatory networks controlling EM transition are well-studied; however, alternative RNA splicing also plays a critical regulatory role during this process. Alternative splicing was proved to control the EM transition process, and RNA-binding proteins were determined to regulate alternative splicing. A comprehensive understanding of alternative splicing and the RNA-binding proteins that regulate it during EM transition and their dynamic impact on breast cancer remains largely unknown. To accurately study the dynamic regulatory relationships, time-series data of the EM transition process are essential. However, only cross-sectional data of epithelial and mesenchymal specimens are available. Therefore, we developed a pseudotemporal causality-based Bayesian (PCB) approach to infer the dynamic regulatory relationships between alternative splicing events and RNA-binding proteins. Our study sheds light on facilitating the regulatory network-based approach to identify key RNA-binding proteins or target alternative splicing events for the diagnosis or treatment of cancers. The data and code for PCB are available at: http://hkumath.hku.hk/~wkc/PCB(data+code).zip. Liangjie Sun, Yushan Qiu, Wai-Ki Ching, Quan Zou 0001 |
PLoS Comput. Biol. | 5 |
| 2023 | Laplacian Regularized Sparse Representation Based Classifier for Identifying DNA N4-Methylcytosine Sites via $L_{2,1/2}$L2,1/2-Matrix NormabstractN4-methylcytosine (4mC) is one of important epigenetic modifications in DNA sequences. Detecting 4mC sites is time-consuming. The computational method based on machine learning has provided effective help for identifying 4mC. To further improve the performance of prediction, we propose a Laplacian Regularized Sparse Representation based Classifier with L2,1/2-matrix norm (LapRSRC). We also utilize kernal trick to derive the kernel LapRSRC for nonlinear modeling. Matrix factorization technology is employed to solve the sparse representation coefficients of all test samples in the training set. And an efficient iterative algorithm is proposed to solve the objective function. We implement our model on six benchmark datasets of 4mC and eight UCI datasets to test evaluate performance. The results show that the performance of our method is better or comparable. Yijie Ding, Wenying He, Jijun Tang, Quan Zou 0001, Fei Guo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2023 | Protein-DNA Binding Residues Prediction Using a Deep Learning Model With Hierarchical Feature ExtractionabstractBiologically important effects occur when proteins bind to other substances, of which binding to DNA is a crucial one. Therefore, accurate identification of protein-DNA binding residues is important for further understanding of the protein-DNA interaction mechanism. Although wet-lab methods can accurately obtain the location of bound residues, it requires significant human, financial and time costs. There is thus an urgent need to develop efficient computational-based methods. Most current state-of-the-art methods are two-step approaches: the first step uses a sliding window technique to extract residue features; the second step uses each residue as an input to the model for prediction. This has a negative impact on the efficiency of prediction and ease of use. In this study, we propose a sequence-to-sequence (seq2seq) model that can input the entire protein sequence of variable length and use two modules, Transformer Encoder Block and Feature Extracting Block, for hierarchical feature extraction, where Transformer Encoder Block is used to extract global features, and then Feature Extracting Block is used to extract local features to further improve the recognition capability of the model. The comparison results on two benchmark datasets, namely PDNA-543 and PDNA-41, prove the effectiveness of our method in identifying protein-DNA binding residues. Quan Zou 0001, Hongjie Wu, Yijie Ding |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2023 | Multi-View Kernel Sparse Representation for Identification of Membrane Protein TypesabstractMembrane proteins are the main undertaker of biomembrane functions and play a vital role in many biological activities of organisms. Prediction of membrane protein types has a great help in determining the function of proteins and understanding the interactions of membrane proteins. However, the biochemical experiment is expensive and not suitable for the large-scale identification of membrane protein types. Therefore, computational methods were used to improve the efficiency of biological experiments. Most existing computational methods only use a single feature of protein, or use multiple features but do not integrate these well. In our study, the protein sequence is described via three different views (features), including amino acid composition, evolutionary information and physicochemical properties of amino acids. To exploit information among all views (features), we introduce a coupling strategy for Kernel Sparse Representation based Classification (KSRC) and construct a new model called Multi-view KSRC (MvKSRC). We implement our method on 4 benchmark data sets of membrane proteins. The comparison results indicate that our method is much superior to all existing methods. Yuqing Qian, Yijie Ding, Quan Zou 0001, Fei Guo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2023 | Enhancer-FRL: Improved and Robust Identification of Enhancers and Their Activities Using Feature Representation LearningabstractEnhancers are crucial for precise regulation of gene expression, while enhancer identification and strength prediction are challenging because of their free distribution and tremendous number of similar fractions in the genome. Although several bioinformatics tools have been developed, shortfalls in these models remain, and their performances need further improvement. In the present study, a two-layer predictor called Enhancer-FRL was proposed for identifying enhancers (enhancers or nonenhancers) and their activities (strong and weak). More specifically, to build an efficient model, the feature representation learning scheme was applied to generate a 50D probabilistic vector based on 10 feature encodings and five machine learning algorithms. Subsequently, the multiview probabilistic features were integrated to construct the final prediction model. Compared with the single feature-based model, Enhancer-FRL showed significant performance improvement and model robustness. Performance assessment on the independent test dataset indicated that the proposed model outperformed state-of-the-art available toolkits. The webserver Enhancer-FRL is freely accessible at http://lab.malab.cn/∼wangchao/softwares/Enhancer-FRL/, The code and datasets can be downloaded at the webserver page or at the Github https://github.com/wangchao-malab/Enhancer-FRL/. Chao Wang 0043, Quan Zou 0001, Ying Ju 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2023 | Decision Tree for SequencesabstractCurrent decision trees such as C4.5 and CART are widely used in different fields due to their simplicity, accuracy and intuitive interpretation. Similar to other popular classifiers, these tree-based classification algorithms are developed for fixed-length vector data and suffer from intrinsic limitations in handling complex data such as sequences. To tackle the discrete sequence classification task, the dominant strategy is to adopt a two-step procedure: first transform the sequential dataset into a vector dataset and then apply existing tree-based classifiers on the new vector data. However, such methods are highly dependent on the feature generation procedure and some features that are critical to the tree construction may be missed. To alleviate these issues, we present a new tree-based sequence classification method, which is able to construct a concise decision tree from the feature space that is composed of all subsequences present in the training sequences. Experimental results on fourteen real datasets show that our method can achieve better performance than those state-of-the-art sequence classification algorithms. The source codes of our method are available at: https://github.com/ZiyaoWu/SeqDT. Zengyou He, Ziyao Wu, Guangyao Xu, Yan Liu 0085, Quan Zou 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2022 | Kernel Risk Sensitive Loss-based Echo State Networks for Predicting Therapeutic Peptides with Sparse LearningabstractThe detection of therapeutic peptides is usually a biochemical experimental method, which is time-consuming and labor-intensive. Lots of computational biology methods had been proposed to solve the problem of therapeutic peptide prediction. However, the existing methods did not consider the processing of noisy samples. We propose a kernel risk-sensitive mean p-power error-based echo state network with sparse learning (KRP-ESN-SL). An efficient iterative optimization algorithm is used to train the model. The KRP-ESN-SL has better performance than other methods. Xiaoyi Guo, Yuqing Qian, Prayag Tiwari, Quan Zou 0001, Yijie Ding |
BIBM | 4 |
| 2022 | Single-cell TF-DNA binding prediction and analysis based on transfer learning frameworkabstractCell type-specific gene expressions during development or in disease are regulated by interactions between transcription factors (TFs) and their binding sites. Recently, many deep learning approaches have been developed to characterize TF-DNA binding within a population of cells. However, determining TF binding sites (TFBSs) in single cells remains challenging due to the sparsity of data. Here, we propose a multi-stage transfer learning framework called STAPLE for single-cell TF-DNA binding prediction and analysis. Specifically, we design the Cell Type Learning to capture the relationship between different TF-DNA binding events in the same cell type. Meanwhile, we present Individual Learning to extract common motif and chromatin accessibility features of a particular binding event in a cellular population. In addition, we leverage Single-cell Learning to annotate TFBSs in each cell without any supervised label. Extensive experiments based on 570 single-cell datasets validate the effectiveness of our framework for considering cellular heterogeneity, outperforming current methods. This work can provide new insight into the relationship between TF-DNA binding and cellular heterogeneity. The source code of STAPLE can be found at https://github.com/ZhangLab312/STAPLE. Zixuan Wang 0025, Yongqing Zhang 0001, Maocheng Wang, Quan Zou 0001 |
BIBM | 6 |
| 2022 | Predicting cell type-specific effects of variants on TF-DNA binding by meta-learningabstractInterpreting the regulatory code of gene expression and further understanding the functionality of noncoding variants on transcriptional effect is a crucial challenge. However, this remains difficult due to the complex association between SNPs and chromatin state. Here, we develop a meta-learning-based framework, U-TransNet, that can accurately predict the TF-DNA binding based on multiple chromatin features. Motivated by ab initio, our proposed framework contain two steps. (i) Meta-learning strategy is applied to predict chromatin profiles from DNA sequence. (ii) DNA sequence and all of these predicted chromatin features are used to predict TF-DNA binding affinity. Experiments demonstrate that U-TransNet has excellent performance, achieving significant improvements over existing methods in predicting base-resolution TF-DNA binding signals, TF binding sites, and motifs. We also demonstrate that integrating the more extended TFBS flank regions is a potential path to better understanding gene transcription. In addition, U-TransNet is applied to infer the effects of variants on TF-DNA binding affinity via in silico mutagenesis, and further to identify cell type-specific functional variants via comparing different cells. To the best of the authors’ knowledge, U-TransNet provides an efficient end-to-end computational framework for deciphering cis-regulator evolution. Yongqing Zhang 0001, Zixuan Wang 0025, Maocheng Wang, Shuwen Xiong, Quan Zou 0001 |
BIBM | 6 |
| 2022 | GATSDCD: Prediction of circRNA-Disease Associations Based on Singular Value Decomposition and Graph Attention Network
Mengting Niu, Abd El-Latif Hesham, Quan Zou 0001 |
ICIC (2) | 3 |
| 2022 | NmRF: identification of multispecies RNA 2'-O-methylation modification sites from RNA sequencesabstract2'-O-methylation (Nm) is a post-transcriptional modification of RNA that is catalyzed by 2'-O-methyltransferase and involves replacing the H on the 2'-hydroxyl group with a methyl group. The 2'-O-methylation modification site is detected in a variety of RNA types (miRNA, tRNA, mRNA, etc.), plays an important role in biological processes and is associated with different diseases. There are few functional mechanisms developed at present, and traditional high-throughput experiments are time-consuming and expensive to explore functional mechanisms. For a deeper understanding of relevant biological mechanisms, it is necessary to develop efficient and accurate recognition tools based on machine learning. Based on this, we constructed a predictor called NmRF based on optimal mixed features and random forest classifier to identify 2'-O-methylation modification sites. The predictor can identify modification sites of multiple species at the same time. To obtain a better prediction model, a two-step strategy is adopted; that is, the optimal hybrid feature set is obtained by combining the light gradient boosting algorithm and incremental feature selection strategy. In 10-fold cross-validation, the accuracies of Homo sapiens and Saccharomyces cerevisiae were 89.069 and 93.885%, and the AUC were 0.9498 and 0.9832, respectively. The rigorous 10-fold cross-validation and independent tests confirm that the proposed method is significantly better than existing tools. A user-friendly web server is accessible at http://lab.malab.cn/∼acy/NmRF. Chunyan Ao, Quan Zou 0001, Liang Yu 0002 |
Briefings Bioinform. | 2 |
| 2022 | Deep learning models for disease-associated circRNA prediction: a reviewabstractEmerging evidence indicates that circular RNAs (circRNAs) can provide new insights and potential therapeutic targets for disease diagnosis and treatment. However, traditional biological experiments are expensive and time-consuming. Recently, deep learning with a more powerful ability for representation learning enables it to be a promising technology for predicting disease-associated circRNAs. In this review, we mainly introduce the most popular databases related to circRNA, and summarize three types of deep learning-based circRNA-disease associations prediction methods: feature-generation-based, type-discrimination and hybrid-based methods. We further evaluate seven representative models on benchmark with ground truth for both balance and imbalance classification tasks. In addition, we discuss the advantages and limitations of each type of method and highlight suggested applications for future research. Yaojia Chen, Jiacheng Wang 0009, Quan Zou 0001 |
Briefings Bioinform. | 5 |
| 2022 | Identification of drug-target interactions via multiple kernel-based triple collaborative matrix factorizationabstractTargeted drugs have been applied to the treatment of cancer on a large scale, and some patients have certain therapeutic effects. It is a time-consuming task to detect drug-target interactions (DTIs) through biochemical experiments. At present, machine learning (ML) has been widely applied in large-scale drug screening. However, there are few methods for multiple information fusion. We propose a multiple kernel-based triple collaborative matrix factorization (MK-TCMF) method to predict DTIs. The multiple kernel matrices (contain chemical, biological and clinical information) are integrated via multi-kernel learning (MKL) algorithm. And the original adjacency matrix of DTIs could be decomposed into three matrices, including the latent feature matrix of the drug space, latent feature matrix of the target space and the bi-projection matrix (used to join the two feature spaces). To obtain better prediction performance, MKL algorithm can regulate the weight of each kernel matrix according to the prediction error. The weights of drug side-effects and target sequence are the highest. Compared with other computational methods, our model has better performance on four test data sets. Yijie Ding, Jijun Tang, Fei Guo 0001, Quan Zou 0001 |
Briefings Bioinform. | 4 |
| 2022 | Structured Sparse Regularized TSK Fuzzy System for predicting therapeutic peptidesabstractTherapeutic peptides act on the skeletal system, digestive system and blood system, have antibacterial properties and help relieve inflammation. In order to reduce the resource consumption of wet experiments for the identification of therapeutic peptides, many computational-based methods have been developed to solve the identification of therapeutic peptides. Due to the insufficiency of traditional machine learning methods in dealing with feature noise. We propose a novel therapeutic peptide identification method called Structured Sparse Regularized Takagi-Sugeno-Kang Fuzzy System on Within-Class Scatter (SSR-TSK-FS-WCS). Our method achieves good performance on multiple therapeutic peptides and UCI datasets. Xiaoyi Guo, Yizhang Jiang, Quan Zou 0001 |
Briefings Bioinform. | 3 |
| 2022 | DeepCap-Kcr: accurate identification and investigation of protein lysine crotonylation sites based on capsule networkabstractLysine crotonylation (Kcr) is a posttranslational modification widely detected in histone and nonhistone proteins. It plays a vital role in human disease progression and various cellular processes, including cell cycle, cell organization, chromatin remodeling and a key mechanism to increase proteomic diversity. Thus, accurate information on such sites is beneficial for both drug development and basic research. Existing computational methods can be improved to more effectively identify Kcr sites in proteins. In this study, we proposed a deep learning model, DeepCap-Kcr, a capsule network (CapsNet) based on a convolutional neural network (CNN) and long short-term memory (LSTM) for robust prediction of Kcr sites on histone and nonhistone proteins (mammals). The proposed model outperformed the existing CNN architecture Deep-Kcr and other well-established tools in most cases and provided promising outcomes for practical use; in particular, the proposed model characterized the internal hierarchical representation as well as the important features from multiple levels of abstraction automatically learned from a small number of samples. The trained model was well generalized in other species (papaya). Moreover, we showed the features and properties generated by the internal capsule layer that can explore the internal data distribution related to biological significance (as a motif detector). The source code and data are freely available at https://github.com/Jhabindra-bioinfo/DeepCap-Kcr. Jhabindra Khanal, Hilal Tayara, Quan Zou 0001, Kil To Chong 0001 |
Briefings Bioinform. | 3 |
| 2022 | A novel fast multiple nucleotide sequence alignment method based on FM-indexabstractMultiple sequence alignment (MSA) is fundamental to many biological applications. But most classical MSA algorithms are difficult to handle large-scale multiple sequences, especially long sequences. Therefore, some recent aligners adopt an efficient divide-and-conquer strategy to divide long sequences into several short sub-sequences. Selecting the common segments (i.e. anchors) for division of sequences is very critical as it directly affects the accuracy and time cost. So, we proposed a novel algorithm, FMAlign, to improve the performance of multiple nucleotide sequence alignment. We use FM-index to extract long common segments at a low cost rather than using a space-consuming hash table. Moreover, after finding the longer optimal common segments, the sequences are divided by the longer common segments. FMAlign has been tested on virus and bacteria genome and human mitochondrial genome datasets, and compared with existing MSA methods such as MAFFT, HAlign and FAME. The experiments show that our method outperforms the existing methods in terms of running time, and has a high accuracy on long sequence sets. All the results demonstrate that our method is applicable to the large-scale nucleotide sequences in terms of sequence length and sequence number. The source code and related data are accessible in https://github.com/iliuh/FMAlign. Huan Liu 0024, Quan Zou 0001 |
Briefings Bioinform. | 2 |
| 2022 | Characterizing viral circRNAs and their application in identifying circRNAs in virusesabstractCircular RNAs (circRNAs) are non-coding RNAs with a special circular structure produced formed by the reverse splicing mechanism, which play an important role in a variety of biological activities. Viruses can encode circRNA, and viral circRNAs have been found in multiple single-stranded and double-stranded viruses. However, the characteristics and functions of viral circRNAs remain unknown. Sequence alignment showed that viral circRNAs are less conserved than circRNAs in animal, indicating that the viral circRNAs may evolve rapidly. Through the analysis of the sequence characteristics of viral circRNAs and circRNAs in animal, it was found that viral circRNAs and animals circRNAs are similar in nucleic acid composition, but have obvious differences in secondary structure and autocorrelation characteristics. Based on these characteristics of viral circRNAs, machine learning algorithms were employed to construct a prediction model to identify viral circRNA. Additionally, analysis of the interaction between viral circRNA and miRNAs showed that viral circRNA is expected to interact with 518 human miRNAs, and preliminary analysis of the role of viral circRNA. And it has been also found that viral circRNAs may be involved in many KEGG pathways related to nervous system and cancer. We curated an online server, and the data and code are available: http://server.malab.cn/viral-CircRNA/. Mengting Niu, Ying Ju 0002, Chen Lin 0001, Quan Zou 0001 |
Briefings Bioinform. | 4 |
| 2022 | Identification of drug-side effect association via restricted Boltzmann machines with penalized termabstractIn the entire life cycle of drug development, the side effect is one of the major failure factors. Severe side effects of drugs that go undetected until the post-marketing stage leads to around two million patient morbidities every year in the United States. Therefore, there is an urgent need for a method to predict side effects of approved drugs and new drugs. Following this need, we present a new predictor for finding side effects of drugs. Firstly, multiple similarity matrices are constructed based on the association profile feature and drug chemical structure information. Secondly, these similarity matrices are integrated by Centered Kernel Alignment-based Multiple Kernel Learning algorithm. Then, Weighted K nearest known neighbors is utilized to complement the adjacency matrix. Next, we construct Restricted Boltzmann machines (RBM) in drug space and side effect space, respectively, and apply a penalized maximum likelihood approach to train model. At last, the average decision rule was adopted to integrate predictions from RBMs. Comparison results and case studies demonstrate, with four benchmark datasets, that our method can give a more accurate and reliable prediction result. Yuqing Qian, Yijie Ding, Quan Zou 0001, Fei Guo 0001 |
Briefings Bioinform. | 3 |
| 2022 | Distant metastasis identification based on optimized graph representation of gene interaction patternsabstractMetastasis is a major cause of cancer morbidity and mortality, and most cancer deaths are caused by cancer metastasis rather than by the primary tumor. The prediction of metastasis based on computational methods has not been explored much in the previous research. In this study, we proposed a graph convolutional network embedded with a graph learning (GL) module, named glmGCN, to predict the distant metastasis of cancer. Both the mRNA and lncRNA expressions were used to provide more genetic information than using the mRNA alone and we used them to construct gene interaction graph representation to consider the effect of genetic interaction. Then, the prediction of the cancer metastasis was performed under a GCN framework, which extracted informative and advanced features from the built non-regular graph structures. Particularly, a GL module was embedded in the proposed glmGCN to learn an optimal graph representation of the gene interaction. We firstly constructed the protein-protein interaction network to represent the initial gene(node) relationship graph. Then, through the GL module, a new graph representation was built which optimally learned the gene interaction strength. Finally, the GCN was adopted to identify the distant metastasis cases. It is worth mentioning that the proposed method pays more attentions on the gene-gene relation than the previous GCN-based method, so more accurate prediction performance can be obtained. The glmGCN was trained based on two types of cancer and was further validated using two other cancer types. A series of experiments have shown that the effectiveness of the proposed method. The implementation for the proposed method is available at https://github.com/RanSuLab/Metastasis-glmGCN. Ran Su, Quan Zou 0001, Leyi Wei |
Briefings Bioinform. | 3 |
| 2022 | A comparison of deep learning-based pre-processing and clustering approaches for single-cell RNA sequencing dataabstractThe emergence of single cell RNA sequencing has facilitated the studied of genomes, transcriptomes and proteomes. As available single-cell RNA-seq datasets are released continuously, one of the major challenges facing traditional RNA analysis tools is the high-dimensional, high-sparsity, high-noise and large-scale characteristics of single-cell RNA-seq data. Deep learning technologies match the characteristics of single-cell RNA-seq data perfectly and offer unprecedented promise. Here, we give a systematic review for most popular single-cell RNA-seq analysis methods and tools based on deep learning models, involving the procedures of data preprocessing (quality control, normalization, data correction, dimensionality reduction and data visualization) and clustering task for downstream analysis. We further evaluate the deep model-based analysis methods of data correction and clustering quantitatively on 11 gold standard datasets. Moreover, we discuss the data preferences of these methods and their limitations, and give some suggestions and guidance for users to select appropriate methods and tools. Jiacheng Wang 0009, Quan Zou 0001, Chen Lin 0001 |
Briefings Bioinform. | 2 |
| 2022 | MDICC: novel method for multi-omics data integration and cancer subtype identificationabstractEach type of cancer usually has several subtypes with distinct clinical implications, and therefore the discovery of cancer subtypes is an important and urgent task in disease diagnosis and therapy. Using single-omics data to predict cancer subtypes is difficult because genomes are dysregulated and complicated by multiple molecular mechanisms, and therefore linking cancer genomes to cancer phenotypes is not an easy task. Using multi-omics data to effectively predict cancer subtypes is an area of much interest; however, integrating multi-omics data is challenging. Here, we propose a novel method of multi-omics data integration for clustering to identify cancer subtypes (MDICC) that integrates new affinity matrix and network fusion methods. Our experimental results show the effectiveness and generalization of the proposed MDICC model in identifying cancer subtypes, and its performance was better than those of currently available state-of-the-art clustering methods. Furthermore, the survival analysis demonstrates that MDICC delivered comparable or even better results than many typical integrative methods. Sha Tian, Yushan Qiu, Quan Zou 0001 |
Briefings Bioinform. | 5 |
| 2022 | Critical assessment of computational tools for prokaryotic and eukaryotic promoter predictionabstractPromoters are crucial regulatory DNA regions for gene transcriptional activation. Rapid advances in next-generation sequencing technologies have accelerated the accumulation of genome sequences, providing increased training data to inform computational approaches for both prokaryotic and eukaryotic promoter prediction. However, it remains a significant challenge to accurately identify species-specific promoter sequences using computational approaches. To advance computational support for promoter prediction, in this study, we curated 58 comprehensive, up-to-date, benchmark datasets for 7 different species (i.e. Escherichia coli, Bacillus subtilis, Homo sapiens, Mus musculus, Arabidopsis thaliana, Zea mays and Drosophila melanogaster) to assist the research community to assess the relative functionality of alternative approaches and support future research on both prokaryotic and eukaryotic promoters. We revisited 106 predictors published since 2000 for promoter identification (40 for prokaryotic promoter, 61 for eukaryotic promoter, and 5 for both). We systematically evaluated their training datasets, computational methodologies, calculated features, performance and software usability. On the basis of these benchmark datasets, we benchmarked 19 predictors with functioning webservers/local tools and assessed their prediction performance. We found that deep learning and traditional machine learning-based approaches generally outperformed scoring function-based approaches. Taken together, the curated benchmark dataset repository and the benchmarking analysis in this study serve to inform the design and implementation of computational approaches for promoter prediction and facilitate more rigorous comparison of new techniques in the future. Meng Zhang 0046, Cangzhi Jia, Fuyi Li, Chen Li 0021, Yan Zhu 0006, Tatsuya Akutsu, Geoffrey I. Webb, Quan Zou 0001, Lachlan James M. Coin, Jiangning Song |
Briefings Bioinform. | 8 |
| 2022 | A novel convolution attention model for predicting transcription factor binding sites by combination of sequence and shapeabstractThe discovery of putative transcription factor binding sites (TFBSs) is important for understanding the underlying binding mechanism and cellular functions. Recently, many computational methods have been proposed to jointly account for DNA sequence and shape properties in TFBSs prediction. However, these methods fail to fully utilize the latent features derived from both sequence and shape profiles and have limitation in interpretability and knowledge discovery. To this end, we present a novel Deep Convolution Attention network combining Sequence and Shape, dubbed as D-SSCA, for precisely predicting putative TFBSs. Experiments conducted on 165 ENCODE ChIP-seq datasets reveal that D-SSCA significantly outperforms several state-of-the-art methods in predicting TFBSs, and justify the utility of channel attention module for feature refinements. Besides, the thorough analysis about the contribution of five shapes to TFBSs prediction demonstrates that shape features can improve the predictive power for transcription factors-DNA binding. Furthermore, D-SSCA can realize the cross-cell line prediction of TFBSs, indicating the occupancy of common interplay patterns concerning both sequence and shape across various cell lines. The source code of D-SSCA can be found at https://github.com/MoonLord0525/. Yongqing Zhang 0001, Zixuan Wang 0025, Yuanqi Zeng, Shuwen Xiong, Maocheng Wang, Jiliu Zhou, Quan Zou 0001 |
Briefings Bioinform. | 8 |
| 2022 | A survey on the algorithm and development of multiple sequence alignmentabstractMultiple sequence alignment (MSA) is an essential cornerstone in bioinformatics, which can reveal the potential information in biological sequences, such as function, evolution and structure. MSA is widely used in many bioinformatics scenarios, such as phylogenetic analysis, protein analysis and genomic analysis. However, MSA faces new challenges with the gradual increase in sequence scale and the increasing demand for alignment accuracy. Therefore, developing an efficient and accurate strategy for MSA has become one of the research hotspots in bioinformatics. In this work, we mainly summarize the algorithms for MSA and its applications in bioinformatics. To provide a structured and clear perspective, we systematically introduce MSA's knowledge, including background, database, metric and benchmark. Besides, we list the most common applications of MSA in the field of bioinformatics, including database searching, phylogenetic analysis, genomic analysis, metagenomic analysis and protein analysis. Furthermore, we categorize and analyze classical and state-of-the-art algorithms, divided into progressive alignment, iterative algorithm, heuristics, machine learning and divide-and-conquer. Moreover, we also discuss the challenges and opportunities of MSA in bioinformatics. Our work provides a comprehensive survey of MSA applications and their relevant algorithms. It could bring valuable insights for researchers to contribute their knowledge to MSA and relevant studies. Yongqing Zhang 0001, Jiliu Zhou, Quan Zou 0001 |
Briefings Bioinform. | 4 |
| 2022 | A hybrid deep learning framework for gene regulatory network inference from single-cell transcriptomic dataabstractInferring gene regulatory networks (GRNs) based on gene expression profiles is able to provide an insight into a number of cellular phenotypes from the genomic level and reveal the essential laws underlying various life phenomena. Different from the bulk expression data, single-cell transcriptomic data embody cell-to-cell variance and diverse biological information, such as tissue characteristics, transformation of cell types, etc. Inferring GRNs based on such data offers unprecedented advantages for making a profound study of cell phenotypes, revealing gene functions and exploring potential interactions. However, the high sparsity, noise and dropout events of single-cell transcriptomic data pose new challenges for regulation identification. We develop a hybrid deep learning framework for GRN inference from single-cell transcriptomic data, DGRNS, which encodes the raw data and fuses recurrent neural network and convolutional neural network (CNN) to train a model capable of distinguishing related gene pairs from unrelated gene pairs. To overcome the limitations of such datasets, it applies sliding windows to extract valuable features while preserving the direction of regulation. DGRNS is constructed as a deep learning model containing gated recurrent unit network for exploring time-dependent information and CNN for learning spatially related information. Our comprehensive and detailed comparative analysis on the dataset of mouse hematopoietic stem cells illustrates that DGRNS outperforms state-of-the-art methods. The networks inferred by DGRNS are about 16% higher than the area under the receiver operating characteristic curve of other unsupervised methods and 10% higher than the area under the precision recall curve of other supervised methods. Experiments on human datasets show the strong robustness and excellent generalization of DGRNS. By comparing the predictions with standard network, we discover a series of novel interactions which are proved to be true in some specific cell types. Importantly, DGRNS identifies a series of regulatory relationships with high confidence and functional consistency, which have not yet been experimentally confirmed and merit further research. Wenying He, Jijun Tang, Quan Zou 0001, Fei Guo 0001 |
Briefings Bioinform. | 4 |
| 2022 | GMNN2CD: identification of circRNA-disease associations based on variational inference and graph Markov neural networksabstractMOTIVATION: With the analysis of the characteristic and function of circular RNAs (circRNAs), people have realized that they play a critical role in the diseases. Exploring the relationship between circRNAs and diseases is of far-reaching significance for searching the etiopathogenesis and treatment of diseases. Nevertheless, it is inefficient to learn new associations only through biotechnology. RESULTS: Consequently, we present a computational method, GMNN2CD, which employs a graph Markov neural network (GMNN) algorithm to predict unknown circRNA-disease associations. First, used verified associations, we calculate semantic similarity and Gaussian interactive profile kernel similarity (GIPs) of the disease and the GIPs of circRNA and then merge them to form a unified descriptor. After that, GMNN2CD uses a fusion feature variational map autoencoder to learn deep features and uses a label propagation map autoencoder to propagate tags based on known associations. Based on variational inference, GMNN alternate training enhances the ability of GMNN2CD to obtain high-efficiency high-dimensional features from low-dimensional representations. Finally, 5-fold cross-validation of five benchmark datasets shows that GMNN2CD is superior to the state-of-the-art methods. Furthermore, case studies have shown that GMNN2CD can detect potential associations. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at https://github.com/nmt315320/GMNN2CD.git. Mengting Niu, Quan Zou 0001, Chunyu Wang 0002 |
Bioinform. | 2 |
| 2022 | i6mA-Caps: a CapsuleNet-based framework for identifying DNA N6-methyladenine sitesabstractMOTIVATION: DNA N6-methyladenine (6mA) has been demonstrated to have an essential function in epigenetic modification in eukaryotic species in recent research. 6mA has been linked to various biological processes. It's critical to create a new algorithm that can rapidly and reliably detect 6mA sites in genomes to investigate their biological roles. The identification of 6mA marks in the genome is the first and most important step in understanding the underlying molecular processes, as well as their regulatory functions. RESULTS: In this article, we proposed a novel computational tool called i6mA-Caps which CapsuleNet based a framework for identifying the DNA N6-methyladenine sites. The proposed framework uses a single encoding scheme for numerical representation of the DNA sequence. The numerical data is then used by the set of convolution layers to extract low-level features. These features are then used by the capsule network to extract intermediate-level and later high-level features to classify the 6mA sites. The proposed network is evaluated on three datasets belonging to three genomes which are Rosaceae, Rice and Arabidopsis thaliana. Proposed method has attained an accuracy of 96.71%, 94% and 86.83% for independent Rosaceae dataset, Rice dataset and A.thaliana dataset respectively. The proposed framework has exhibited improved results when compared with the existing top-of-the-line methods. AVAILABILITY AND IMPLEMENTATION: A user-friendly web-server is made available for the biological experts which can be accessed at: http://nsclbio.jbnu.ac.kr/tools/i6mA-Caps/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mobeen Ur Rehman, Hilal Tayara, Quan Zou 0001, Kil To Chong 0001 |
Bioinform. | 3 |
| 2022 | NerLTR-DTA: drug-target binding affinity prediction based on neighbor relationship and learning to rankabstractMOTIVATION: Drug-target interaction prediction plays an important role in new drug discovery and drug repurposing. Binding affinity indicates the strength of drug-target interactions. Predicting drug-target binding affinity is expected to provide promising candidates for biologists, which can effectively reduce the workload of wet laboratory experiments and speed up the entire process of drug research. Given that, numerous new proteins are sequenced and compounds are synthesized, several improved computational methods have been proposed for such predictions, but there are still some challenges. (i) Many methods only discuss and implement one application scenario, they focus on drug repurposing and ignore the discovery of new drugs and targets. (ii) Many methods do not consider the priority order of proteins (or drugs) related to each target drug (or protein). Therefore, it is necessary to develop a comprehensive method that can be used in multiple scenarios and focuses on candidate order. RESULTS: In this study, we propose a method called NerLTR-DTA that uses the neighbor relationship of similarity and sharing to extract features, and applies a ranking framework with regression attributes to predict affinity values and priority order of query drug (or query target) and its related proteins (or compounds). It is worth noting that using the characteristics of learning to rank to set different queries can smartly realize the multi-scenario application of the method, including the discovery of new drugs and new targets. Experimental results on two commonly used datasets show that NerLTR-DTA outperforms some state-of-the-art competing methods. NerLTR-DTA achieves excellent performance in all application scenarios mentioned in this study, and the rm(test)2 values guarantee such excellent performance is not obtained by chance. Moreover, it can be concluded that NerLTR-DTA can provide accurate ranking lists for the relevant results of most queries through the statistics of the association relationship of each query drug (or query protein). In general, NerLTR-DTA is a powerful tool for predicting drug-target associations and can contribute to new drug discovery and drug repurposing. AVAILABILITY AND IMPLEMENTATION: The proposed method is implemented in Python and Java. Source codes and datasets are available at https://github.com/RUXIAOQING964914140/NerLTR-DTA. Xiaoqing Ru, Xiucai Ye, Tetsuya Sakurai, Quan Zou 0001 |
Bioinform. | 4 |
| 2022 | Predicting protein-peptide binding residues via interpretable deep learningabstractSUMMARY: Identifying the protein-peptide binding residues is fundamentally important to understand the mechanisms of protein functions and explore drug discovery. Although several computational methods have been developed, most of them highly rely on third-party tools or complex data preprocessing for feature design, easily resulting in low computational efficacy and suffering from low predictive performance. To address the limitations, we propose PepBCL, a novel BERT (Bidirectional Encoder Representation from Transformers) -based contrastive learning framework to predict the protein-peptide binding residues based on protein sequences only. PepBCL is an end-to-end predictive model that is independent of feature engineering. Specifically, we introduce a well pre-trained protein language model that can automatically extract and learn high-latent representations of protein sequences relevant for protein structures and functions. Further, we design a novel contrastive learning module to optimize the feature representations of binding residues underlying the imbalanced dataset. We demonstrate that our proposed method significantly outperforms the state-of-the-art methods under benchmarking comparison, and achieves more robust performance. Moreover, we found that we further improve the performance via the integration of traditional features and our learnt features. Interestingly, the interpretable analysis of our model highlights the flexibility and adaptability of deep learning-based protein language model to capture both conserved and non-conserved sequential characteristics of peptide-binding residues. Finally, to facilitate the use of our method, we establish an online predictive platform as the implementation of the proposed PepBCL, which is now available at http://server.wei-group.net/PepBCL/. AVAILABILITY AND IMPLEMENTATION: https://github.com/Ruheng-W/PepBCL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ruheng Wang, Junru Jin, Quan Zou 0001, Kenta Nakai, Leyi Wei |
Bioinform. | 3 |
| 2022 | Effector-GAN: prediction of fungal effector proteins based on pretrained deep representation learning methods and generative adversarial networksabstractMOTIVATION: Phytopathogenic fungi secrete effector proteins to subvert host defenses and facilitate infection. Systematic analysis and prediction of candidate fungal effector proteins are crucial for experimental validation and biological control of plant disease. However, two problems are still considered intractable to be solved in fungal effector prediction: one is the high-level diversity in effector sequences that increases the difficulty of protein feature learning, and the other is the class imbalance between effector and non-effector samples in the training dataset. RESULTS: In our study, pretrained deep representation learning methods are presented to represent multiple characteristics of sequences for predicting fungal effectors and generative adversarial networks are adapted to create synthetic feature samples to address the data imbalance problem. Compared with the state-of-the-art fungal effector prediction methods, Effector-GAN shows an overall improvement in accuracy in the independent test set. AVAILABILITY AND IMPLEMENTATION: Effector-GAN offers a user-friendly interface to inspect potential fungal effector proteins (http://lab.malab.cn/~wys/webserver/Effector-GAN). The Python script can be downloaded from http://lab.malab.cn/~wys/gitlab/effector-gan. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yansu Wang, Ximei Luo, Quan Zou 0001 |
Bioinform. | 3 |
| 2022 | WMSA: a novel method for multiple sequence alignment of DNA sequencesabstractMOTIVATION: Multiple sequence alignment (MSA) is a fundamental problem in bioinformatics. The quality of alignment will affect downstream analysis. MAFFT has adopted the Fast Fourier Transform method for searching the homologous segments and using them as anchors to divide the sequences, then making alignment only on segments, which can save time and memory without overly reducing the sequence alignment quality. MAFFT becomes slow when the dataset is large. RESULTS: We made a software, WMSA, which uses the divide-and-conquer method to split the sequences into clusters, aligns those clusters into profiles with the center star strategy and then makes a progressive profile-profile alignment. The alignment is conducted by the compiled algorithms of MAFFT, K-Band with multithread parallelism. Our method can balance time, space and quality and performs better than MAFFT in test experiments on highly conserved datasets. AVAILABILITY AND IMPLEMENTATION: Source code is freely available at https://github.com/malabz/WMSA/, which is implemented in C/C++ and supported on Linux, and datasets are available at https://github.com/malabz/WMSA-dataset. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yanming Wei, Quan Zou 0001, Furong Tang, Liang Yu 0002 |
Bioinform. | 2 |
| 2022 | webSCST: an interactive web application for single-cell RNA-sequencing data and spatial transcriptomic data integrationabstractSUMMARY: Integrative analysis of single-cell RNA-sequencing (scRNA-seq) data with spatial data for the same species and organ would provide each cell sample with a predictive spatial location, which would facilitate biological study. However, publicly available spatial sequencing datasets for specific species and organs are rare and are often displayed in different formats. In this study, we introduce a new web-based scRNA-seq analysis tool, webSCST, that integrates well-organized spatial transcriptome sequencing datasets categorized by species and organs, provides a user-friendly interface for raw single-cell processing with popular integration methods and allows users to submit their raw scRNA-seq data once to obtain predicted spatial locations for each cell type. AVAILABILITY AND IMPLEMENTATION: webSCST implemented in shiny with all major browsers supported is available at http://www.webscst.com. webSCST is also freely available as an R package at https://github.com/swsoyee/webSCST. Feifei Cui, Lijun Dou, Chen Cao 0002, Quan Zou 0001 |
Bioinform. | 7 |
| 2022 | DeepMPM: a mortality risk prediction model using longitudinal EHR dataabstractBACKGROUND: Accurate precision approaches have far not been developed for modeling mortality risk in intensive care unit (ICU) patients. Conventional mortality risk prediction methods can hardly extract the information in longitudinal electronic medical records (EHRs) effectively, since they simply aggregate the heterogeneous variables in EHRs, ignoring the complex relationship and interactions between variables and the time dependence in longitudinal records. Recently deep learning approaches have been widely used in modeling longitudinal EHR data. However, most existing deep learning-based risk prediction approaches only use the information of a single disease, neglecting the interactions between multiple diseases and different conditions. RESULTS: In this paper, we address this unmet need by leveraging disease and treatment information in EHRs to develop a mortality risk prediction model based on deep learning (DeepMPM). DeepMPM utilizes a two-level attention mechanism, i.e. visit-level and variable-level attention, to derive the representation of patient risk status from patient's multiple longitudinal medical records. Benefiting from using EHR of patients with multiple diseases and different conditions, DeepMPM can achieve state-of-the-art performances in mortality risk prediction. CONCLUSIONS: Experiment results on MIMIC III database demonstrates that with the disease and treatment information DeepMPM can achieve a good performance in terms of Area Under ROC Curve (0.85). Moreover, DeepMPM can successfully model the complex interactions between diseases to achieve better representation learning of disease and treatment than other deep learning approaches, so as to improve the accuracy of mortality prediction. A case study also shows that DeepMPM offers the potential to provide users with insights into feature correlation in data as well as model behavior for each prediction. Fan Yang 0010, Yongxuan Lai, Quan Zou 0001 |
BMC Bioinform. | 6 |
| 2022 | DeepM6ASeq-EL: prediction of human N6-methyladenosine (m6A) sites with LSTM and ensemble learning
Quan Zou 0001 |
Frontiers Comput. Sci. | 2 |
| 2022 | Identification and classification of promoters using the attention mechanism based on long short-term memory
Lei Xu 0047, Quan Zou 0001, Jin Wu 0002 |
Frontiers Comput. Sci. | 4 |
| 2022 | String kernels construction and fusion: a survey with bioinformatics application
Ren Qi, Fei Guo 0001, Quan Zou 0001 |
Frontiers Comput. Sci. | 3 |
| 2022 | MLapSVM-LBS: Predicting DNA-binding proteins via a multiple Laplacian regularized support vector machine with local behavior similarityabstractDNA-binding proteins (DBPs) are of great significance in many basic cellular processes. Experiment-based methods for identifying DBPs are costly and time-consuming. To deal with large-scale DBP identification tasks, a variety of computation-based methods have been developed. Inspired by previous work, we propose a multiple Laplacian regularized support vector machine with local behavior similarity (MLapSVM-LBS) to predict DBP. We serially combine three features that are extracted from protein sequences (including PsePSSM, GE, NMBAC) and feed them into MLapSVM-LBS. Based on human behavior learning theory, MLapSVM-LBS can better represent the relationship between samples through local behavior similarity. We introduce a new edge weight calculation method that takes label information into consideration. In addition, a local distribution parameter reflecting the underlying probability distribution of a sample’s neighborhood is also employed. To further improve the robustness of the model, we utilize multiple Laplacian regularization to build a multigraph model in which five Laplacian graphs are constructed with local behavior similarity by changing the neighborhood size. To appraise the performance of our model, MLapSVM-LBS is trained and tested on the PDB186, PDB1075, PDB2272 and PDB14189 datasets. On two independent testing sets (PDB186 and PDB2272), our method reaches the accuracies of 0.887 and 0.712, respectively. The good results on both datasets demonstrate the reliable performance of our model. Mengwei Sun, Prayag Tiwari, Yuqing Qian, Yijie Ding, Quan Zou 0001 |
Knowl. Based Syst. | 5 |
| 2022 | Shared subspace-based radial basis function neural network for identifying ncRNAs subcellular localizationabstractNon-coding RNAs (ncRNAs) play an important role in revealing the mechanism of human disease for anti-tumor and anti-virus substances. Detecting subcellular locations of ncRNAs is a necessary way to study ncRNA. Traditional biochemical methods are time-consuming and labor-intensive, and computational-based methods can help detect the location of ncRNAs on a large scale. However, many models did not consider the correlation information among multiple subcellular localizations of ncRNAs. This study proposes a radial basis function neural network based on shared subspace learning (RBFNN-SSL), which extract shared structures in multi-labels. To evaluate performance, our classifier is tested on three ncRNA datasets. Our model achieves better performance in experimental results. Yijie Ding, Prayag Tiwari, Fei Guo 0001, Quan Zou 0001 |
Neural Networks | 4 |
| 2022 | CRBPDL: Identification of circRNA-RBP interaction sites using an ensemble neural network approachabstractCircular RNAs (circRNAs) are non-coding RNAs with a special circular structure produced formed by the reverse splicing mechanism. Increasing evidence shows that circular RNAs can directly bind to RNA-binding proteins (RBP) and play an important role in a variety of biological activities. The interactions between circRNAs and RBPs are key to comprehending the mechanism of posttranscriptional regulation. Accurately identifying binding sites is very useful for analyzing interactions. In past research, some predictors on the basis of machine learning (ML) have been presented, but prediction accuracy still needs to be ameliorated. Therefore, we present a novel calculation model, CRBPDL, which uses an Adaboost integrated deep hierarchical network to identify the binding sites of circular RNA-RBP. CRBPDL combines five different feature encoding schemes to encode the original RNA sequence, uses deep multiscale residual networks (MSRN) and bidirectional gating recurrent units (BiGRUs) to effectively learn high-level feature representations, it is sufficient to extract local and global context information at the same time. Additionally, a self-attention mechanism is employed to train the robustness of the CRBPDL. Ultimately, the Adaboost algorithm is applied to integrate deep learning (DL) model to improve prediction performance and reliability of the model. To verify the usefulness of CRBPDL, we compared the efficiency with state-of-the-art methods on 37 circular RNA data sets and 31 linear RNA data sets. Moreover, results display that CRBPDL is capable of performing universal, reliable, and robust. The code and data sets are obtainable at https://github.com/nmt315320/CRBPDL.git. Mengting Niu, Quan Zou 0001, Chen Lin 0001 |
PLoS Comput. Biol. | 2 |
| 2022 | A multi-label learning model for predicting drug-induced pathology in multi-organ based on toxicogenomics dataabstractDrug-induced toxicity damages the health and is one of the key factors causing drug withdrawal from the market. It is of great significance to identify drug-induced target-organ toxicity, especially the detailed pathological findings, which are crucial for toxicity assessment, in the early stage of drug development process. A large variety of studies have devoted to identify drug toxicity. However, most of them are limited to single organ or only binary toxicity. Here we proposed a novel multi-label learning model named Att-RethinkNet, for predicting drug-induced pathological findings targeted on liver and kidney based on toxicogenomics data. The Att-RethinkNet is equipped with a memory structure and can effectively use the label association information. Besides, attention mechanism is embedded to focus on the important features and obtain better feature presentation. Our Att-RethinkNet is applicable in multiple organs and takes account the compound type, dose, and administration time, so it is more comprehensive and generalized. And more importantly, it predicts multiple pathological findings at the same time, instead of predicting each pathology separately as the previous model did. To demonstrate the effectiveness of the proposed model, we compared the proposed method with a series of state-of-the-arts methods. Our model shows competitive performance and can predict potential hepatotoxicity and nephrotoxicity in a more accurate and reliable way. The implementation of the proposed method is available at https://github.com/RanSuLab/Drug-Toxicity-Prediction-MultiLabel. Ran Su, Haitang Yang, Leyi Wei, Siqi Chen 0001, Quan Zou 0001 |
PLoS Comput. Biol. | 5 |
| 2022 | CRCF: A Method of Identifying Secretory Proteins of Malaria ParasitesabstractMalaria is a mosquito-borne disease that results in millions of cases and deaths annually. The development of a fast computational method that identifies secretory proteins of the malaria parasite is important for research on antimalarial drugs and vaccines. Thus, a method was developed to identify the secretory proteins of malaria parasites. In this method, a reduced alphabet was selected to recode the original protein sequence. A feature synthesis method was used to synthesise three different types of feature information. Finally, the random forest method was used as a classifier to identify the secretory proteins. In addition, a web server was developed to share the proposed algorithm. Experiments using the benchmark dataset demonstrated that the overall accuracy achieved by the proposed method was greater than 97.8 percent using the 10-fold cross-validation method. Furthermore, the reduced schemes and characteristic performance analyses are discussed. Changli Feng, Jin Wu 0002, Haiyan Wei, Lei Xu 0047, Quan Zou 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2022 | Significance-Based Essential Protein DiscoveryabstractThe identification of essential proteins is an important problem in bioinformatics. During the past decades, many centrality measures and algorithms have been proposed to address this issue. However, existing methods still deserve the following drawbacks: (1) the lack of a context-free and readily interpretable quantification of their centrality values; (2) the difficulty of specifying a proper threshold for their centrality values; (3) the incapability of controlling the quality of reported essential proteins in a statistically sound manner. To overcome the limitations of existing solutions, we tackle the essential protein discovery problem from a significance testing perspective. More precisely, the essential protein discovery problem is formulated as a multiple hypothesis testing problem, where the null hypothesis is that each protein is not an essential protein. To quantify the statistical significance of each protein, we present a p-value calculation method in which both the degree and the local clustering coefficient are used as the test statistic and the Erdös-Rényi model is employed as the random graph model. After calculating the p-value for each protein, the false discovery rate is used as the error rate in the multiple testing correction procedure. Our significance-based essential protein discovery method is named as SigEP, which is tested on both simulated networks and real PPI networks. The experimental results show that our method is able to achieve better performance than those competing algorithms. Yan Liu 0085, Quan Zou 0001, Zengyou He |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | SgRNA-RF: Identification of SgRNA On-Target Activity With Imbalanced DatasetsabstractSingle-guide RNA is a guide RNA (gRNA), which guides the insertion or deletion of uridine residues into kinetoplastid during RNA editing. It is a small non-coding RNA that can be combined with pre -mRNA pairing. SgRNA is a critical component of the CRISPR/Cas9 gene knockout system and play an important role in gene editing and gene regulation. It is important to accurately and quickly identify highly on-target activity sgRNAs. Due to its importance, several computational predictors have been proposed to predict sgRNAs on-target activity. All these methods have clearly contributed to the development of this very important field. However, they also have certain limitations. In the paper, we developed a new classifier SgRNA-RF, which extracts the features of nucleic acid composition and structure of on-target activity sgRNA sequence and identified by random forest algorithm. In addition to solving an imbalanced dataset, this paper proposed a new method called CS-Smote. We compared sgRNA-RF with state-of-the-art predictors on the five datasets, and found SgRNA-RF significantly improved the identification accuracy, with accuracies of 0.8636,0.9161,0.894,0.938,0.965,0.77,0.979,0.973, respectively. The user-friendly web server that implements sgRNA-RF is freely available at http://server.malab.cn/sgRNA-RF/. Mengting Niu, Quan Zou 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | C-Loss Based Higher Order Fuzzy Inference Systems for Identifying DNA N4-Methylcytosine SitesabstractDNA methylation is an epigenetic marker that plays an important role in the biological processes of regulating gene expression, maintaining chromatin structure, imprinting genes, inactivating X chromosomes, and developing embryos. The traditional detection method is time-consuming. Currently, researchers have used effective computational methods to improve the efficiency of methylation detection. This study proposes a fuzzy model with correntropy induced loss (C-loss) function to identify DNA N4-methylcytosine (4 mC) sites. To improve the robustness and performance of the model, we use kernel method and the C-loss function to build a higher order fuzzy inference systems. To test performance, our model is implemented on six 4 mC and eight University of California Irvine (UCI) datasets. The experimental results show that our model achieves better prediction performance. Yijie Ding, Prayag Tiwari, Quan Zou 0001, Fei Guo 0001, Hari Mohan Pandey |
IEEE Trans. Fuzzy Syst. | 3 |
| 2022 | MDADP: A Webserver Integrating Database and Prediction Tools for Microbe-Disease AssociationsabstractMore and more evidence has demonstrated that microbiota play important roles in the life processes of the human body. In recent years, various computational methods have been proposed for identifying potentially disease-associated microbes to save costs in traditional biological experiments. However, prediction performances of these methods are generally limited by outdated and incomplete datasets. And moreover, until now, there are limited studies that can provide visual predictive tools for inferring possible microbe-disease associations (MDAs) as well. Hence, in this manuscript, a novel webserver called MDADP will be proposed to identify latent MDAs, in which, a new MDA database together with interactive prediction tools for MDAs studies will be designed simultaneously. Especially, in the newly constructed MDA database, 2019 known MDAs between 58 diseases and 703 microbes have been manually collected first. And then, through adopting the average ranking method and the co-confidence method respectively, eight representative computational models have been integrated together to identify potential disease-related microbes. As a result, MDADP can provide not only interactive features for users to access and capture MDAs entities, but alsoeffective tools for users to identify candidate microbes for different diseases. To our knowledge, MDADP is the first online platform that incorporates a new MDA database with comprehensive MDA prediction tools. Therefore, we believe that it will be a valuable source of information for researches in microbiology and disease-related fields. MDADP can be accessed at http://mdadp.leelab2997.cn. Lei Wang 0069, Yuqi Wang 0006, Yihong Tan, Tingrui Pei, Quan Zou 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2021 | Membrane Protein Identification via Multi-view Graph Regularized k-Local Hyperplane Distance Nearest Neighbor ModelabstractX-ray diffraction and nuclear magnetic resonance spectroscopy are the main methods for measuring membrane proteins. The traditional methods are time-consuming and labor-intensive. To large-scale prediction and screening of membrane proteins, a graph regularized k-local hyperplane distance nearest neighbor model (GHKNN) is proposed to identify of membrane protein types. For effectively integrating features, multi-view learning (MVL) is employed to estimate the weight of each graph. We test GHKNN on 2 data sets of membrane protein. Compared with other methods, the accuracy of GHKNN is better or comparable. Mengwei Sun, Yuqing Qian, Yijie Ding, Jijun Tang, Quan Zou 0001 |
BIBM | 5 |
| 2021 | By hybrid neural networks for prediction and interpretation of transcription factor binding sites based on multi-omicsabstractTranscription factors (TFs) binding sites prediction and analysis are vital for comprehending cis-regulatory mechanisms. Recently, several deep learning-based methods have shown outstanding performance on TFs binding sites (TFBSs) recognition by leveraging solely base-pair arrangement of regulatory sequences. Except for the aforementioned genomic features, the epigenomics signature represented by the histone modification is also a critical factor related to TFs-DNA binding. We present a multi-omics based hybrid neural network, dubbed as BHSite, for TFBSs prediction by adaptively integrating base-pair arrangements and histone modification signatures. Experiments over 196 ChIP-seq datasets demonstrate that BHSite significantly outperforms several state-of-the-art methods in TFBSs prediction. Besides, studies of the relative importance of histone modification signatures prove that diverse signatures complement each other. Furthermore, visualization analysis of Squeeze-and-Excitation Network reveals the contribution of multi-omics latent features concerning different cell types to TFBS prediction. Thus, BHSite improves both performance and interpretability by combining the multi-omic features into deep learning architecture. Yongqing Zhang 0001, Zixuan Wang 0025, Libo Lu, Xiaoyao Tan, Quan Zou 0001 |
BIBM | 6 |
| 2021 | News Popularity Prediction with Local-Global Long-Short-Term Embedding
Shuai Fan 0007, Chen Lin 0001, Hui Li 0057, Quan Zou 0001 |
WISE (2) | 4 |
| 2021 | ITP-Pred: an interpretable method for predicting, therapeutic peptides with fused features low-dimension representationabstractThe peptide therapeutics market is providing new opportunities for the biotechnology and pharmaceutical industries. Therefore, identifying therapeutic peptides and exploring their properties are important. Although several studies have proposed different machine learning methods to predict peptides as being therapeutic peptides, most do not explain the decision factors of model in detail. In this work, an Interpretable Therapeutic Peptide Prediction (ITP-Pred) model based on efficient feature fusion was developed. First, we proposed three kinds of feature descriptors based on sequence and physicochemical property encoded, namely amino acid composition (AAC), group AAC and coding autocorrelation, and concatenated them to obtain the feature representation of therapeutic peptide. Then, we input it into the CNN-Bi-directional Long Short-Term Memory (BiLSTM) model to automatically learn recognition of therapeutic peptides. The cross-validation and independent verification experiments results indicated that ITP-Pred has a higher prediction performance on the benchmark dataset than other comparison methods. Finally, we analyzed the output of the model from two aspects: sequence order and physical and chemical properties, mining important features as guidance for the design of better models that can complement existing methods. Li Wang 0145, Xiangzheng Fu, Chenxing Xia, Xiangxiang Zeng, Quan Zou 0001 |
Briefings Bioinform. | 6 |
| 2021 | Molecular design in drug discovery: a comprehensive review of deep generative modelsabstractDeep generative models have been an upsurge in the deep learning community since they were proposed. These models are designed for generating new synthetic data including images, videos and texts by fitting the data approximate distributions. In the last few years, deep generative models have shown superior performance in drug discovery especially de novo molecular design. In this study, deep generative models are reviewed to witness the recent advances of de novo molecular design for drug discovery. In addition, we divide those models into two categories based on molecular representations in silico. Then these two classical types of models are reported in detail and discussed about both pros and cons. We also indicate the current challenges in deep generative models for de novo molecular design. De novo molecular design automatically is promising but a long road to be explored. Yongshun Gong, Yuansheng Liu, Bosheng Song, Quan Zou 0001 |
Briefings Bioinform. | 5 |
| 2021 | A comprehensive review of the imbalance classification of protein post-translational modificationsabstractPost-translational modifications (PTMs) play significant roles in regulating protein structure, activity and function, and they are closely involved in various pathologies. Therefore, the identification of associated PTMs is the foundation of in-depth research on related biological mechanisms, disease treatments and drug design. Due to the high cost and time consumption of high-throughput sequencing techniques, developing machine learning-based predictors has been considered an effective approach to rapidly recognize potential modified sites. However, the imbalanced distribution of true and false PTM sites, namely, the data imbalance problem, largely effects the reliability and application of prediction tools. In this article, we conduct a systematic survey of the research progress in the imbalanced PTMs classification. First, we describe the modeling process in detail and outline useful data imbalance solutions. Then, we summarize the recently proposed bioinformatics tools based on imbalanced PTM data and simultaneously build a convenient website, ImClassi_PTMs (available at lab.malab.cn/∼dlj/ImbClassi_PTMs/), to facilitate the researchers to view. Moreover, we analyze the challenges of current computational predictors and propose some suggestions to improve the efficiency of imbalance learning. We hope that this work will provide comprehensive knowledge of imbalanced PTM recognition and contribute to advanced predictors in the future. Lijun Dou, Fenglong Yang, Lei Xu 0047, Quan Zou 0001 |
Briefings Bioinform. | 4 |
| 2021 | MMFGRN: a multi-source multi-model fusion method for gene regulatory network reconstructionabstractLots of biological processes are controlled by gene regulatory networks (GRNs), such as growth and differentiation of cells, occurrence and development of the diseases. Therefore, it is important to persistently concentrate on the research of GRN. The determination of the gene-gene relationships from gene expression data is a complex issue. Since it is difficult to efficiently obtain the regularity behind the gene-gene relationship by only relying on biochemical experimental methods, thus various computational methods have been used to construct GRNs, and some achievements have been made. In this paper, we propose a novel method MMFGRN (for "Multi-source Multi-model Fusion for Gene Regulatory Network reconstruction") to reconstruct the GRN. In order to make full use of the limited datasets and explore the potential regulatory relationships contained in different data types, we construct the MMFGRN model from three perspectives: single time series data model, single steady-data model and time series and steady-data joint model. And, we utilize the weighted fusion strategy to get the final global regulatory link ranking. Finally, MMFGRN model yields the best performance on the DREAM4 InSilico_Size10 data, outperforming other popular inference algorithms, with an overall area under receiver operating characteristic score of 0.909 and area under precision-recall (AUPR) curves score of 0.770 on the 10-gene network. Additionally, as the network scale increases, our method also has certain advantages with an overall AUPR score of 0.335 on the DREAM4 InSilico_Size100 data. These results demonstrate the good robustness of MMFGRN on different scales of networks. At the same time, the integration strategy proposed in this paper provides a new idea for the reconstruction of the biological network model without prior knowledge, which can help researchers to decipher the elusive mechanism of life. Wenying He, Jijun Tang, Quan Zou 0001, Fei Guo 0001 |
Briefings Bioinform. | 3 |
| 2021 | DeepATT: a hybrid category attention neural network for identifying functional effects of DNA sequencesabstractQuantifying DNA properties is a challenging task in the broad field of human genomics. Since the vast majority of non-coding DNA is still poorly understood in terms of function, this task is particularly important to have enormous benefit for biology research. Various DNA sequences should have a great variety of representations, and specific functions may focus on corresponding features in the front part of learning model. Currently, however, for multi-class prediction of non-coding DNA regulatory functions, most powerful predictive models do not have appropriate feature extraction and selection approaches for specific functional effects, so that it is difficult to gain a better insight into their internal correlations. Hence, we design a category attention layer and category dense layer in order to select efficient features and distinguish different DNA functions. In this study, we propose a hybrid deep neural network method, called DeepATT, for identifying $919$ regulatory functions on nearly $5$ million DNA sequences. Our model has four built-in neural network constructions: convolution layer captures regulatory motifs, recurrent layer captures a regulatory grammar, category attention layer selects corresponding valid features for different functions and category dense layer classifies predictive labels with selected features of regulatory functions. Importantly, we compare our novel method, DeepATT, with existing outstanding prediction tools, DeepSEA and DanQ. DeepATT performs significantly better than other existing tools for identifying DNA functions, at least increasing $1.6\%$ area under precision recall. Furthermore, we can mine the important correlation among different DNA functions according to the category attention module. Moreover, our novel model can greatly reduce the number of parameters by the mechanism of attention and locally connected, on the basis of ensuring accuracy. Jiawei Li 0018, Yuqian Pu, Jijun Tang, Quan Zou 0001, Fei Guo 0001 |
Briefings Bioinform. | 4 |
| 2021 | EP3: an ensemble predictor that accurately identifies type III secreted effectorsabstractType III secretion systems (T3SS) can be found in many pathogenic bacteria, such as Dysentery bacillus, Salmonella typhimurium, Vibrio cholera and pathogenic Escherichia coli. The routes of infection of these bacteria include the T3SS transferring a large number of type III secreted effectors (T3SE) into host cells, thereby blocking or adjusting the communication channels of the host cells. Therefore, the accurate identification of T3SEs is the precondition for the further study of pathogenic bacteria. In this article, a new T3SEs ensemble predictor was developed, which can accurately distinguish T3SEs from any unknown protein. In the course of the experiment, methods and models are strictly trained and tested. Compared with other methods, EP3 demonstrates better performance, including the absence of overfitting, strong robustness and powerful predictive ability. EP3 (an ensemble predictor that accurately identifies T3SEs) is designed to simplify the user's (especially nonprofessional users) access to T3SEs for further investigation, which will have a significant impact on understanding the progression of pathogenic bacterial infections. Based on the integrated model that we proposed, a web server had been established to distinguish T3SEs from non-T3SEs, where have EP3_1 and EP3_2. The users can choose the model according to the species of the samples to be tested. Our related tools and data can be accessed through the link http://lab.malab.cn/∼lijing/EP3.html. Leyi Wei, Fei Guo 0001, Quan Zou 0001 |
Briefings Bioinform. | 4 |
| 2021 | SubLocEP: a novel ensemble predictor of subcellular localization of eukaryotic mRNA based on machine learningabstractMOTIVATION: mRNA location corresponds to the location of protein translation and contributes to precise spatial and temporal management of the protein function. However, current assignment of subcellular localization of eukaryotic mRNA reveals important limitations: (1) turning multiple classifications into multiple dichotomies makes the training process tedious; (2) the majority of the models trained by classical algorithm are based on the extraction of single sequence information; (3) the existing state-of-the-art models have not reached an ideal level in terms of prediction and generalization ability. To achieve better assignment of subcellular localization of eukaryotic mRNA, a better and more comprehensive model must be developed. RESULTS: In this paper, SubLocEP is proposed as a two-layer integrated prediction model for accurate prediction of the location of sequence samples. Unlike the existing models based on limited features, SubLocEP comprehensively considers additional feature attributes and is combined with LightGBM to generated single feature classifiers. The initial integration model (single-layer model) is generated according to the categories of a feature. Subsequently, two single-layer integration models are weighted (sequence-based: physicochemical properties = 3:2) to produce the final two-layer model. The performance of SubLocEP on independent datasets is sufficient to indicate that SubLocEP is an accurate and stable prediction model with strong generalization ability. Additionally, an online tool has been developed that contains experimental data and can maximize the user convenience for estimation of subcellular localization of eukaryotic mRNA. Shida He, Fei Guo 0001, Quan Zou 0001 |
Briefings Bioinform. | 5 |
| 2021 | Anticancer peptides prediction with deep representation learning featuresabstractAnticancer peptides constitute one of the most promising therapeutic agents for combating common human cancers. Using wet experiments to verify whether a peptide displays anticancer characteristics is time-consuming and costly. Hence, in this study, we proposed a computational method named identify anticancer peptides via deep representation learning features (iACP-DRLF) using light gradient boosting machine algorithm and deep representation learning features. Two kinds of sequence embedding technologies were used, namely soft symmetric alignment embedding and unified representation (UniRep) embedding, both of which involved deep neural network models based on long short-term memory networks and their derived networks. The results showed that the use of deep representation learning features greatly improved the capability of the models to discriminate anticancer peptides from other peptides. Also, UMAP (uniform manifold approximation and projection for dimension reduction) and SHAP (shapley additive explanations) analysis proved that UniRep have an advantage over other features for anticancer peptide identification. The python script and pretrained models could be downloaded from https://github.com/zhibinlv/iACP-DRLF or from http://public.aibiochem.net/iACP-DRLF/. Zhibin Lv, Feifei Cui, Quan Zou 0001, Lei Xu 0047 |
Briefings Bioinform. | 3 |
| 2021 | A spectral clustering with self-weighted multiple kernel learning method for single-cell RNA-seq dataabstractSingle-cell RNA-sequencing (scRNA-seq) data widely exist in bioinformatics. It is crucial to devise a distance metric for scRNA-seq data. Almost all existing clustering methods based on spectral clustering algorithms work in three separate steps: similarity graph construction; continuous labels learning; discretization of the learned labels by k-means clustering. However, this common practice has potential flaws that may lead to severe information loss and degradation of performance. Furthermore, the performance of a kernel method is largely determined by the selected kernel; a self-weighted multiple kernel learning model can help choose the most suitable kernel for scRNA-seq data. To this end, we propose to automatically learn similarity information from data. We present a new clustering method in the form of a multiple kernel combination that can directly discover groupings in scRNA-seq data. The main proposition is that automatically learned similarity information from scRNA-seq data is used to transform the candidate solution into a new solution that better approximates the discrete one. The proposed model can be efficiently solved by the standard support vector machine (SVM) solvers. Experiments on benchmark scRNA-Seq data validate the superior performance of the proposed model. Spectral clustering with multiple kernels is implemented in Matlab, licensed under Massachusetts Institute of Technology (MIT) and freely available from the Github website, https://github.com/Cuteu/SMSC/. Ren Qi, Jin Wu 0002, Fei Guo 0001, Lei Xu 0047, Quan Zou 0001 |
Briefings Bioinform. | 5 |
| 2021 | Matrix factorization-based data fusion for the prediction of RNA-binding proteins and alternative splicing event associations during epithelial-mesenchymal transitionabstractMOTIVATION: The epithelial-mesenchymal transition (EMT) is a cellular-developmental process activated during tumor metastasis. Transcriptional regulatory networks controlling EMT are well studied; however, alternative RNA splicing also plays a critical regulatory role during this process. Unfortunately, a comprehensive understanding of alternative splicing (AS) and the RNA-binding proteins (RBPs) that regulate it during EMT remains largely unknown. Therefore, a great need exists to develop effective computational methods for predicting associations of RBPs and AS events. Dramatically increasing data sources that have direct and indirect information associated with RBPs and AS events have provided an ideal platform for inferring these associations. RESULTS: In this study, we propose a novel method for RBP-AS target prediction based on weighted data fusion with sparse matrix tri-factorization (WDFSMF in short) that simultaneously decomposes heterogeneous data source matrices into low-rank matrices to reveal hidden associations. WDFSMF can select and integrate data sources by assigning different weights to those sources, and these weights can be assigned automatically. In addition, WDFSMF can identify significant RBP complexes regulating AS events and eliminate noise and outliers from the data. Our proposed method achieves an area under the receiver operating characteristic curve (AUC) of $90.78\%$, which shows that WDFSMF can effectively predict RBP-AS event associations with higher accuracy compared with previous methods. Furthermore, this study identifies significant RBPs as complexes for AS events during EMT and provides solid ground for further investigation into RNA regulation during EMT and metastasis. WDFSMF is a general data fusion framework, and as such it can also be adapted to predict associations between other biological entities. Yushan Qiu, Wai-Ki Ching, Quan Zou 0001 |
Briefings Bioinform. | 3 |
| 2021 | Prediction of RNA-binding protein and alternative splicing event associations during epithelial-mesenchymal transition based on inductive matrix completionabstractMOTIVATION: The developmental process of epithelial-mesenchymal transition (EMT) is abnormally activated during breast cancer metastasis. Transcriptional regulatory networks that control EMT have been well studied; however, alternative RNA splicing plays a vital regulatory role during this process and the regulating mechanism needs further exploration. Because of the huge cost and complexity of biological experiments, the underlying mechanisms of alternative splicing (AS) and associated RNA-binding proteins (RBPs) that regulate the EMT process remain largely unknown. Thus, there is an urgent need to develop computational methods for predicting potential RBP-AS event associations during EMT. RESULTS: We developed a novel model for RBP-AS target prediction during EMT that is based on inductive matrix completion (RAIMC). Integrated RBP similarities were calculated based on RBP regulating similarity, and RBP Gaussian interaction profile (GIP) kernel similarity, while integrated AS event similarities were computed based on AS event module similarity and AS event GIP kernel similarity. Our primary objective was to complete missing or unknown RBP-AS event associations based on known associations and on integrated RBP and AS event similarities. In this paper, we identify significant RBPs for AS events during EMT and discuss potential regulating mechanisms. Our computational results confirm the effectiveness and superiority of our model over other state-of-the-art methods. Our RAIMC model achieved AUC values of 0.9587 and 0.9765 based on leave-one-out cross-validation (CV) and 5-fold CV, respectively, which are larger than the AUC values from the previous models. RAIMC is a general matrix completion framework that can be adopted to predict associations between other biological entities. We further validated the prediction performance of RAIMC on the genes CD44 and MAP3K7. RAIMC can identify the related regulating RBPs for isoforms of these two genes. AVAILABILITY AND IMPLEMENTATION: The source code for RAIMC is available at https://github.com/yushanqiu/RAIMC. CONTACT: [email protected] online. Yushan Qiu, Wai-Ki Ching, Quan Zou 0001 |
Briefings Bioinform. | 3 |
| 2021 | Application of learning to rank in bioinformatics tasksabstractOver the past decades, learning to rank (LTR) algorithms have been gradually applied to bioinformatics. Such methods have shown significant advantages in multiple research tasks in this field. Therefore, it is necessary to summarize and discuss the application of these algorithms so that these algorithms are convenient and contribute to bioinformatics. In this paper, the characteristics of LTR algorithms and their strengths over other types of algorithms are analyzed based on the application of multiple perspectives in bioinformatics. Finally, the paper further discusses the shortcomings of the LTR algorithms, the methods and means to better use the algorithms and some open problems that currently exist. Xiaoqing Ru, Xiucai Ye, Tetsuya Sakurai, Quan Zou 0001 |
Briefings Bioinform. | 4 |
| 2021 | Revisiting genome-wide association studies from statistical modelling to machine learningabstractOver the last decade, genome-wide association studies (GWAS) have discovered thousands of genetic variants underlying complex human diseases and agriculturally important traits. These findings have been utilized to dissect the biological basis of diseases, to develop new drugs, to advance precision medicine and to boost breeding. However, the potential of GWAS is still underexploited due to methodological limitations. Many challenges have emerged, including detecting epistasis and single-nucleotide polymorphisms (SNPs) with small effects and distinguishing causal variants from other SNPs associated through linkage disequilibrium. These issues have motivated advancements in GWAS analyses in two contrasting cultures-statistical modelling and machine learning. In this review, we systematically present the basic concepts and the benefits and limitations in both methods. We further discuss recent efforts to mitigate their weaknesses. Additionally, we summarize the state-of-the-art tools for detecting the missed signals, ultrarare mutations and gene-gene interactions and for prioritizing SNPs. Our work can offer both theoretical and practical guidelines for performing GWAS analyses and for developing further new robust methods to fully exploit the potential of GWAS. Shanwen Sun, BenZhi Dong, Quan Zou 0001 |
Briefings Bioinform. | 3 |
| 2021 | The accurate prediction and characterization of cancerlectin by a combined machine learning and GO analysisabstractCancerlectins, lectins linked to tumor progression, have become the focus of cancer therapy research for their carbohydrate-binding specificity. However, the specific characterization for cancerlectins involved in tumor progression is still unclear. By taking advantage of the g-gap tripeptide and tetrapeptide composition feature descriptors, we increased the accuracy of the classification model of cancerlectin and lectin to 98.54% and 95.38%, respectively. About 36 cancerlectin and 135 lectin features were selected for functional characterization by P/N feature ranking method, which particularly selects the features in positive samples. The specific protein domains of cancerlectins are found to be p-GalNAc-T, crystal and annexin by comparing with lectins through the exclusion method. Moreover, the combined GO analysis showed that the conserved cation binding sites of cancerlectin specific domains are covered by selected feature peptides, suggesting that the capability of cation binding, critical for enzyme activity and stability, could be the key characteristic of cancerlectins in tumor progression. These results will help to identify potential cancerlectin and provide clues for mechanism study of cancerlectin in tumor progression. Furong Tang, Lei Xu 0047, Quan Zou 0001, Hailin Feng |
Briefings Bioinform. | 4 |
| 2021 | Machine learning for phytopathology: from the molecular scale towards the network scaleabstractWith the increasing volume of high-throughput sequencing data from a variety of omics techniques in the field of plant-pathogen interactions, sorting, retrieving, processing and visualizing biological information have become a great challenge. Within the explosion of data, machine learning offers powerful tools to process these complex omics data by various algorithms, such as Bayesian reasoning, support vector machine and random forest. Here, we introduce the basic frameworks of machine learning in dissecting plant-pathogen interactions and discuss the applications and advances of machine learning in plant-pathogen interactions from molecular to network biology, including the prediction of pathogen effectors, plant disease resistance protein monitoring and the discovery of protein-protein networks. The aim of this review is to provide a summary of advances in plant defense and pathogen infection and to indicate the important developments of machine learning in phytopathology. Yansu Wang, Murong Zhou, Quan Zou 0001, Lei Xu 0047 |
Briefings Bioinform. | 3 |
| 2021 | VPTMdb: a viral posttranslational modification databaseabstractIn viruses, posttranslational modifications (PTMs) are essential for their life cycle. Recognizing viral PTMs is very important for a better understanding of the mechanism of viral infections and finding potential drug targets. However, few studies have investigated the roles of viral PTMs in virus-human interactions using comprehensive viral PTM datasets. To fill this gap, we developed the first comprehensive viral posttranslational modification database (VPTMdb) for collecting systematic information of PTMs in human viruses and infected host cells. The VPTMdb contains 1240 unique viral PTM sites with 8 modification types from 43 viruses (818 experimentally verified PTM sites manually extracted from 150 publications and 422 PTMs extracted from SwissProt) as well as 13 650 infected cells' PTMs extracted from seven global proteomics experiments in six human viruses. The investigation of viral PTM sequences motifs showed that most viral PTMs have the consensus motifs with human proteins in phosphorylation and five cellular kinase families phosphorylate more than 10 viral species. The analysis of protein disordered regions presented that more than 50% glycosylation sites of double-strand DNA viruses are in the disordered regions, whereas single-strand RNA and retroviruses prefer ordered regions. Domain-domain interaction analysis indicating potential roles of viral PTMs play in infections. The findings should make an important contribution to the field of virus-human interaction. Moreover, we created a novel sequence-based classifier named VPTMpre to help users predict viral protein phosphorylation sites. VPTMdb online web server (http://vptmdb.com:8787/VPTMdb/) was implemented for users to download viral PTM data and predict phosphorylation sites of interest. Yujia Xiang, Quan Zou 0001, Lilin Zhao |
Briefings Bioinform. | 2 |
| 2021 | An in silico approach to identification, categorization and prediction of nucleic acid binding proteinsabstractThe interaction between proteins and nucleic acid plays an important role in many processes, such as transcription, translation and DNA repair. The mechanisms of related biological events can be understood by exploring the function of proteins in these interactions. The number of known protein sequences has increased rapidly in recent years, but the databases for describing the structure and function of protein have unfortunately grown quite slowly. Thus, improving such databases is meaningful for predicting protein-nucleic acid interactions. Furthermore, the mechanism of related biological events, such as viral infection or designing novel drug targets, can be further understood by understanding the function of proteins in these interactions. The information for each sequence, including its function and interaction sites, were collected and identified, and a database called PNIDB was built. The proteins in PNIDB were grouped into 27 classes, such as transcription, immune system, and structural protein, etc. The function of each protein was then predicted using a machine learning method. Using our method, the predictor was trained on labeled sequences, and then the function of a protein was predicted based on the trained classifier. The prediction accuracy achieved a score of 77.43% by 10-fold cross validation. Lei Xu 0047, Jin Wu 0002, Quan Zou 0001 |
Briefings Bioinform. | 4 |
| 2021 | DisBalance: a platform to automatically build balance-based disease prediction models and discover microbial biomarkers from microbiome dataabstractHow best to utilize the microbial taxonomic abundances in regard to the prediction and explanation of human diseases remains appealing and challenging, and the relative nature of microbiome data necessitates a proper feature selection method to resolve the compositional problem. In this study, we developed an all-in-one platform to address a series of issues in microbiome-based human disease prediction and taxonomic biomarkers discovery. We prioritize the interpretation, runtime and classification accuracy of the distal discriminative balances analysis (DBA-distal) method in selecting a set of distal discriminative balances, and develop DisBalance, a comprehensive platform, to integrate and streamline the workflows of disease model building, disease risk prediction and disease-related biomarker discovery for microbiome-based binary classifications. DisBalance allows the de novo model-building and disease risk prediction in a very fast and convenient way. To facilitate the model-driven and knowledge-driven discoveries, DisBalance dedicates multiple strategies for the mining of microbial biomarkers. The independent validation of the models constructed by the DisBalance pipeline is performed on seven microbiome datasets from the original article of DBA-distal. The implementation of the DisBalance platform is demonstrated by a complete analysis of a shotgun metagenomic dataset of Ulcerative Colitis (UC). As a free and open-source, DisBlance can be accessed at http://lab.malab.cn/soft/DisBalance. The source code and demo data for Disbalance are available at https://github.com/yangfenglong/DisBalance. Fenglong Yang, Quan Zou 0001 |
Briefings Bioinform. | 2 |
| 2021 | GutBalance: a server for the human gut microbiome-based disease prediction and biomarker discovery with compositionality addressedabstractThe compositionality of the microbiome data is well-known but often neglected. The compositional transformation pertains to the supervised learning of microbiome data and is a critical step that decides the performance and reliability of the disease classifiers. We value the excellent performance of the distal discriminative balance analysis (DBA) method, which selects distal balances of pairs and trios of bacteria, in addressing the classification of high-dimensional microbiome data. By applying this method to the species-level abundances of all the disease phenotypes in the GMrepo database, we build a balance-based model repository for the classification of human gut microbiome-related diseases. The model repository supports the prediction of disease risks for new sample(s). More importantly, we highlight the concept of balance-disease associations rather than the conventional microbe-disease associations and develop the human Gut Balance-Disease Association Database (GBDAD). Each predictable balance for each disease model indicates a potential biomarker-disease relationship and can be interpreted as a bacteria ratio positively or negatively correlated with the disease. Furthermore, by linking the balance-disease associations to the evidenced microbe-disease associations in MicroPhenoDB, we surprisingly found that most species-disease associations inferred from the shotgun metagenomic datasets can be validated by external evidence beyond MicroPhenoDB. The balance-based species-disease association inference will accelerate the generation of new microbe-disease association hypotheses in gastrointestinal microecology research and clinical trials. The model repository and the GBDAD database are deployed on the GutBalance server, which supports interactive visualization and systematic interrogation of the disease models, disease-related balances and disease-related species of interest. Fenglong Yang, Quan Zou 0001 |
Briefings Bioinform. | 2 |
| 2021 | Critical downstream analysis steps for single-cell RNA sequencing dataabstractSingle-cell RNA sequencing (scRNA-seq) has enabled us to study biological questions at the single-cell level. Currently, many analysis tools are available to better utilize these relatively noisy data. In this review, we summarize the most widely used methods for critical downstream analysis steps (i.e. clustering, trajectory inference, cell-type annotation and integrating datasets). The advantages and limitations are comprehensively discussed, and we provide suggestions for choosing proper methods in different situations. We hope this paper will be useful for scRNA-seq data analysts and bioinformatics tool developers. Feifei Cui, Chen Lin 0001, Lingling Zhao, Chunyu Wang 0002, Quan Zou 0001 |
Briefings Bioinform. | 6 |
| 2021 | Goals and approaches for each processing step for single-cell RNA sequencing dataabstractSingle-cell RNA sequencing (scRNA-seq) has enabled researchers to study gene expression at the cellular level. However, due to the extremely low levels of transcripts in a single cell and technical losses during reverse transcription, gene expression at a single-cell resolution is usually noisy and highly dimensional; thus, statistical analyses of single-cell data are a challenge. Although many scRNA-seq data analysis tools are currently available, a gold standard pipeline is not available for all datasets. Therefore, a general understanding of bioinformatics and associated computational issues would facilitate the selection of appropriate tools for a given set of data. In this review, we provide an overview of the goals and most popular computational analysis tools for the quality control, normalization, imputation, feature selection and dimension reduction of scRNA-seq data. Feifei Cui, Chunyu Wang 0002, Lingling Zhao, Quan Zou 0001 |
Briefings Bioinform. | 5 |
| 2021 | High-resolution transcription factor binding sites prediction improved performance and interpretability by deep learning methodabstractTranscription factors (TFs) are essential proteins in regulating the spatiotemporal expression of genes. It is crucial to infer the potential transcription factor binding sites (TFBSs) with high resolution to promote biology and realize precision medicine. Recently, deep learning-based models have shown exemplary performance in the prediction of TFBSs at the base-pair level. However, the previous models fail to integrate nucleotide position information and semantic information without noisy responses. Thus, there is still room for improvement. Moreover, both the inner mechanism and prediction results of these models are challenging to interpret. To this end, the Deep Attentive Encoder-Decoder Neural Network (D-AEDNet) is developed to identify the location of TFs-DNA binding sites in DNA sequences. In particular, our model adopts Skip Architecture to leverage the nucleotide position information in the encoder and removes noisy responses in the information fusion process by Attention Gate. Simultaneously, the Transcription Factor Motif Discovery based on Sliding Window (TF-MoDSW), an approach to discover TFs-DNA binding motifs by utilizing the output of neural networks, is proposed to understand the biological meaning of the predicted result. On ChIP-exo datasets, experimental results show that D-AEDNet has better performance than competing methods. Besides, we authenticate that Attention Gate can improve the interpretability of our model by ways of visualization analysis. Furthermore, we confirm that ability of D-AEDNet to learn TFs-DNA binding motifs outperform the state-of-the-art methods and availability of TF-MoDSW to discover biological sequence motifs in TFs-DNA interaction by conducting experiment on ChIP-seq datasets. Yongqing Zhang 0001, Zixuan Wang 0025, Yuanqi Zeng, Jiliu Zhou, Quan Zou 0001 |
Briefings Bioinform. | 5 |
| 2021 | A comprehensive overview and critical evaluation of gene regulatory network inference technologiesabstractGene regulatory network (GRN) is the important mechanism of maintaining life process, controlling biochemical reaction and regulating compound level, which plays an important role in various organisms and systems. Reconstructing GRN can help us to understand the molecular mechanism of organisms and to reveal the essential rules of a large number of biological processes and reactions in organisms. Various outstanding network reconstruction algorithms use specific assumptions that affect prediction accuracy, in order to deal with the uncertainty of processing. In order to study why a certain method is more suitable for specific research problem or experimental data, we conduct research from model-based, information-based and machine learning-based method classifications. There are obviously different types of computational tools that can be generated to distinguish GRNs. Furthermore, we discuss several classical, representative and latest methods in each category to analyze core ideas, general steps, characteristics, etc. We compare the performance of state-of-the-art GRN reconstruction technologies on simulated networks and real networks under different scaling conditions. Through standardized performance metrics and common benchmarks, we quantitatively evaluate the stability of various methods and the sensitivity of the same algorithm applying to different scaling networks. The aim of this study is to explore the most appropriate method for a specific GRN, which helps biologists and medical scientists in discovering potential drug targets and identifying cancer biomarkers. Wenying He, Jijun Tang, Quan Zou 0001, Fei Guo 0001 |
Briefings Bioinform. | 4 |
| 2021 | Minirmd: accurate and fast duplicate removal tool for short reads via multiple minimizersabstractSUMMARY: Removing duplicate and near-duplicate reads, generated by high-throughput sequencing technologies, is able to reduce computational resources in downstream applications. Here we develop minirmd, a de novo tool to remove duplicate reads via multiple rounds of clustering using different length of minimizer. Experiments demonstrate that minirmd removes more near-duplicate reads than existing clustering approaches and is faster than existing multi-core tools. To the best of our knowledge, minirmd is the first tool to remove near-duplicates on reverse-complementary strand. AVAILABILITY AND IMPLEMENTATION: https://github.com/yuansliu/minirmd. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yuansheng Liu, Xiaocai Zhang, Quan Zou 0001, Xiangxiang Zeng |
Bioinform. | 3 |
| 2021 | Identification of sub-Golgi protein localization by use of deep representation learning featuresabstractMOTIVATION: The Golgi apparatus has a key functional role in protein biosynthesis within the eukaryotic cell with malfunction resulting in various neurodegenerative diseases. For a better understanding of the Golgi apparatus, it is essential to identification of sub-Golgi protein localization. Although some machine learning methods have been used to identify sub-Golgi localization proteins by sequence representation fusion, more accurate sub-Golgi protein identification is still challenging by existing methodology. RESULTS: we developed a protein sub-Golgi localization identification protocol using deep representation learning features with 107 dimensions. By this protocol, we demonstrated that instead of multi-type protein sequence feature representation fusion as in previous state-of-the-art sub-Golgi-protein localization classifiers, it is sufficient to exploit only one type of feature representation for more accurately identification of sub-Golgi proteins. Compared with independent testing results for benchmark datasets, our protocol is able to perform generally, reliably and robustly for sub-Golgi protein localization prediction. AVAILABILITYAND IMPLEMENTATION: A use-friendly webserver is freely accessible at http://isGP-DRLF.aibiochem.net and the prediction code is accessible at https://github.com/zhibinlv/isGP-DRLF. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zhibin Lv, Quan Zou 0001, Qinghua Jiang |
Bioinform. | 3 |
| 2021 | BP4RNAseq: a babysitter package for retrospective and newly generated RNA-seq data analyses using both alignment-based and alignment-free quantification methodabstractSUMMARY: Processing raw reads of RNA-sequencing (RNA-seq) data, no matter public or newly sequenced data, involves a lot of specialized tools and technical configurations that are often unfamiliar and time-consuming to learn for non-bioinformatics researchers. Here, we develop the R package BP4RNAseq, which integrates the state-of-art tools from both alignment-based and alignment-free quantification workflows. The BP4RNAseq package is a highly automated tool using an optimized pipeline to improve the sensitivity and accuracy of RNA-seq analyses. It can take only two non-technical parameters and output six formatted gene expression quantification at gene and transcript levels. The package applies to both retrospective and newly generated bulk RNA-seq data analyses and is also applicable for single-cell RNA-seq analyses. It, therefore, greatly facilitates the application of RNA-seq. AVAILABILITY AND IMPLEMENTATION: The BP4RNAseq package for R and its documentation are freely available at https://github.com/sunshanwen/BP4RNAseq. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shanwen Sun, Lei Xu 0047, Quan Zou 0001, Guohua Wang 0001 |
Bioinform. | 3 |
| 2021 | DeepAc4C: a convolutional neural network model with hybrid features composed of physicochemical patterns and distributed representation information for identification of N4-acetylcytidine in mRNAabstractMOTIVATION: N4-acetylcytidine (ac4C) is the only acetylation modification that has been characterized in eukaryotic RNA, and is correlated with various human diseases. Laboratory identification of ac4C is complicated by factors, such as sample hydrolysis and high cost. Unfortunately, existing computational methods to identify ac4C do not achieve satisfactory performance. RESULTS: We developed a novel tool, DeepAc4C, which identifies ac4C using convolutional neural networks (CNNs) using hybrid features composed of physicochemical patterns and a distributed representation of nucleic acids. Our results show that the proposed model achieved better and more balanced performance than existing predictors. Furthermore, we evaluated the effect that specific features had on the model predictions and their interaction effects. Several interesting sequence motifs specific to ac4C were identified. AVAILABILITY AND IMPLEMENTATION: The webserver is freely accessible at https://ac4c.webmalab.cn/, the source code and datasets are accessible at Zenodo with URL https://doi.org/10.5281/zenodo.5138047 and Github with URL https://github.com/wangchao-malab/DeepAc4C. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chao Wang 0043, Ying Ju 0002, Quan Zou 0001, Chen Lin 0001 |
Bioinform. | 3 |
| 2021 | CarSite-II: an integrated classification algorithm for identifying carbonylated sites based on K-means similarity-based undersampling and synthetic minority oversampling techniquesabstractBACKGROUND: Carbonylation is a non-enzymatic irreversible protein post-translational modification, and refers to the side chain of amino acid residues being attacked by reactive oxygen species and finally converted into carbonyl products. Studies have shown that protein carbonylation caused by reactive oxygen species is involved in the etiology and pathophysiological processes of aging, neurodegenerative diseases, inflammation, diabetes, amyotrophic lateral sclerosis, Huntington's disease, and tumor. Current experimental approaches used to predict carbonylation sites are expensive, time-consuming, and limited in protein processing abilities. Computational prediction of the carbonylation residue location in protein post-translational modifications enhances the functional characterization of proteins. RESULTS: In this study, an integrated classifier algorithm, CarSite-II, was developed to identify K, P, R, and T carbonylated sites. The resampling method K-means similarity-based undersampling and the synthetic minority oversampling technique (SMOTE-KSU) were incorporated to balance the proportions of K, P, R, and T carbonylated training samples. Next, the integrated classifier system Rotation Forest uses "support vector machine" subclassifications to divide three types of feature spaces into several subsets. CarSite-II gained Matthew's correlation coefficient (MCC) values of 0.2287/0.3125/0.2787/0.2814, False Positive rate values of 0.2628/0.1084/0.1383/0.1313, False Negative rate values of 0.2252/0.0205/0.0976/0.0608 for K/P/R/T carbonylation sites by tenfold cross-validation, respectively. On our independent test dataset, CarSite-II yield MCC values of 0.6358/0.2910/0.4629/0.3685, False Positive rate values of 0.0165/0.0203/0.0188/0.0094, False Negative rate values of 0.1026/0.1875/0.2037/0.3333 for K/P/R/T carbonylation sites. The results show that CarSite-II achieves remarkably better performance than all currently available prediction tools. CONCLUSION: The related results revealed that CarSite-II achieved better performance than the currently available five programs, and revealed the usefulness of the SMOTE-KSU resampling approach and integration algorithm. For the convenience of experimental scientists, the web tool of CarSite-II is available in http://47.100.136.41:8081/. Yun Zuo 0001, Jianyuan Lin, Xiangxiang Zeng, Quan Zou 0001, Xiangrong Liu |
BMC Bioinform. | 4 |
| 2021 | Using a low correlation high orthogonality feature set and machine learning methods to identify plant pentatricopeptide repeat coding gene/protein
Changli Feng, Quan Zou 0001 |
Neurocomputing | 2 |
| 2021 | A Convolutional Neural Network Using Dinucleotide One-hot Encoder for identifying DNA N6-Methyladenine Sites in the Rice Genome
Zhibin Lv, Hui Ding 0005, Lei Wang 0069, Quan Zou 0001 |
Neurocomputing | 4 |
| 2021 | Prediction of drug-target interactions based on multi-layer network representation learning
Yifan Shang, Lin Gao 0006, Quan Zou 0001, Liang Yu 0002 |
Neurocomputing | 3 |
| 2021 | iPro2L-PSTKNC: A Two-Layer Predictor for Discovering Various Types of Promoters by Position Specific of Nucleotide CompositionabstractPromoters are DNA regulatory elements located proximal to the transcription start site, which are in charge of the initiation of specific gene transcription. In Escherichia coli, promoters can be recognized by σ factors that have multiple families based on distinct function and structure, such as σ24, σ28, σ32, σ38, σ54and σ70. At present, biological methods are mainly used to identify these promoters. However, because it is time-consuming and material-consuming to do biological experiments, computational biology algorithm has emerged as a more effective way to predict the classification. In this study, we develop a novel two-layer seamless predictor called iPro2L-PSTKNC to identify the promoters of the E. coli genome, which based on the feature extraction model we newly proposed that is named as the position specific tendencies of k-mer nucleotide composition (PSTKNC). On the first layer, it is a binary classification predicting whether a sequence is promoter or not. And the second layer is a multiple classification identifying which type the identified promoter belongs to. The ensemble classification SVM performsbest comparing with other algorithms, which gets a promising accuracy and the Matthews correlation coefficient (MCC) at 90.05% and 80.13%. Our data and code are available at https://github.com/lyuyinuo/iPro2L-PSTKNC. Yinuo Lyu, Wenying He, Quan Zou 0001, Fei Guo 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2021 | rBPDL: Predicting RNA-Binding Proteins Using Deep LearningabstractRNA-binding protein (RBP) is a powerful and wide-ranging regulator that plays an important role in cell development, differentiation, metabolism, health and disease. The prediction of RBPs provides valuable guidance for biologists. Although experimental methods have made great progress in predicting RBP, they are time-consuming and not flexible. Therefore, we developed a network model, rBPDL, by combining a convolutional neural network and long short-term memory for multilabel classification of RBPs. Moreover, to achieve better prediction results, we used a voting algorithm for ensemble learning of the model. We compared rBPDL with state-of-the-art methods and found that rBPDL significantly improved identification performance for the RBP68 dataset, with a macro-Area Under Curve (AUC), micro-AUC, and weighted AUC of 0.936, 0.962, and 0.946, respectively. Furthermore, through AUC statistical analysis of the RBP domain, we analyzed the performance of rBPDL and found that the RBP identification performance in the same domain was similar. In addition, we analyzed the performance preferences and physicochemical properties of the binding protein amino acids and explored the characteristics that affect the binding by using the RBP86 dataset. Mengting Niu, Jin Wu 0002, Quan Zou 0001, Lei Xu 0047 |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | GRRFNet: Guided Regularized Random Forest-based Gene Regulatory Network Inference Using Data IntegrationabstractGene regulatory network (GRN) inference based on gene expression data is still a huge challenge in systems biology. Genomic data, including time-series expression data, steady-state data, knockout data, and other biological data, such as Gene Ontology (GO) annotations, provide information on potential gene regulation. However, most existing methods continue to use only a single dataset for GRN inference. To integrate these types of data and improve the accuracy of inference, we propose a new data-integration strategy based on guided regular random forest (GRRF) for GRN inference, dubbed GRRFNet. Specifically, first, time-series data and steady-state data as main datasets are integrated to generate learning samples; simultaneously, other datasets are processed to design penalty coefficients for guiding feature selection; then, a GRRF model is applied to integrate the prior information with a main dataset to learn the transcription function and evaluate the importance of feature; finally, the score of the feature's importance is used as the possibility of the gene regulatory relationships to construct the GRN. To evaluate the performance of GRRFNet, we compare it with GENIE3, dynGENIE3, and GRIEF at the artificial DREAM4 dataset and the real Escherichia coli dataset. Although GRRFNet does not yield the best performance on every network, its competitiveness is still reflected herein. Yongqing Zhang 0001, Qingyuan Chen, Dongrui Gao, Quan Zou 0001 |
BIBM | 4 |
| 2020 | Computational methods for identifying the critical nodes in biological networksabstractA biological network is complex. A group of critical nodes determines the quality and state of such a network. Increasing studies have shown that diseases and biological networks are closely and mutually related and that certain diseases are often caused by errors occurring in certain nodes in biological networks. Thus, studying biological networks and identifying critical nodes can help determine the key targets in treating diseases. The problem is how to find the critical nodes in a network efficiently and with low cost. Existing experimental methods in identifying critical nodes generally require much time, manpower and money. Accordingly, many scientists are attempting to solve this problem by researching efficient and low-cost computing methods. To facilitate calculations, biological networks are often modeled as several common networks. In this review, we classify biological networks according to the network types used by several kinds of common computational methods and introduce the computational methods used by each type of network. Xiangrong Liu, Zengyan Hong, Juan Liu 0003, Alfonso Rodríguez-Patón, Quan Zou 0001, Xiangxiang Zeng |
Briefings Bioinform. | 6 |
| 2020 | HITS-PR-HHblits: protein remote homology detection by combining PageRank and Hyperlink-Induced Topic SearchabstractAs one of the most important fundamental problems in protein sequence analysis, protein remote homology detection is critical for both theoretical research (protein structure and function studies) and real world applications (drug design). Although several computational predictors have been proposed, their detection performance is still limited. In this study, we treat protein remote homology detection as a document retrieval task, where the proteins are considered as documents and its aim is to find the highly related documents with the query documents in a database. A protein similarity network was constructed based on the true labels of proteins in the database, and the query proteins were then connected into the network based on the similarity scores calculated by three ranking methods, including PSI-BLAST, Hmmer and HHblits. The PageRank algorithm and Hyperlink-Induced Topic Search (HITS) algorithm were respectively performed on this network to move the homologous proteins of query proteins to the neighbors of the query proteins in the network. Finally, PageRank and HITS algorithms were combined, and a predictor called HITS-PR-HHblits was proposed to further improve the predictive performance. Tested on the SCOP and SCOPe benchmark datasets, the experimental results showed that the proposed protocols outperformed other state-of-the-art methods. For the convenience of the most experimental scientists, a web server for HITS-PR-HHblits was established at http://bioinformatics.hitsz.edu.cn/HITS-PR-HHblits, by which the users can easily get the results without the need to go through the mathematical details. The HITS-PR-HHblits predictor is a protocol for protein remote homology detection using different sets of programs, which will become a very useful computational tool for proteome analysis. Bin Liu 0014, Shuangyan Jiang, Quan Zou 0001 |
Briefings Bioinform. | 3 |
| 2020 | Clustering and classification methods for single-cell RNA-sequencing dataabstractAppropriate ways to measure the similarity between single-cell RNA-sequencing (scRNA-seq) data are ubiquitous in bioinformatics, but using single clustering or classification methods to process scRNA-seq data is generally difficult. This has led to the emergence of integrated methods and tools that aim to automatically process specific problems associated with scRNA-seq data. These approaches have attracted a lot of interest in bioinformatics and related fields. In this paper, we systematically review the integrated methods and tools, highlighting the pros and cons of each approach. We not only pay particular attention to clustering and classification methods but also discuss methods that have emerged recently as powerful alternatives, including nonlinear and linear methods and descending dimension methods. Finally, we focus on clustering and classification methods for scRNA-seq data, in particular, integrated methods, and provide a comprehensive description of scRNA-seq data and download URLs. Ren Qi, Anjun Ma, Qin Ma 0003, Quan Zou 0001 |
Briefings Bioinform. | 4 |
| 2020 | Critical evaluation of web-based prediction tools for human protein subcellular localizationabstractHuman protein subcellular localization has an important research value in biological processes, also in elucidating protein functions and identifying drug targets. Over the past decade, a number of protein subcellular localization prediction tools have been designed and made freely available online. The purpose of this paper is to summarize the progress of research on the subcellular localization of human proteins in recent years, including commonly used data sets proposed by the predecessors and the performance of all selected prediction tools against the same benchmark data set. We carry out a systematic evaluation of several publicly available subcellular localization prediction methods on various benchmark data sets. Among them, we find that mLASSO-Hum and pLoc-mHum provide a statistically significant improvement in performance, as measured by the value of accuracy, relative to the other methods. Meanwhile, we build a new data set using the latest version of Uniprot database and construct a new GO-based prediction method HumLoc-LBCI in this paper. Then, we test all selected prediction tools on the new data set. Finally, we discuss the possible development directions of human protein subcellular localization. Availability: The codes and data are available from http://www.lbci.cn/syn/. Yinan Shen, Yijie Ding, Jijun Tang, Quan Zou 0001, Fei Guo 0001 |
Briefings Bioinform. | 4 |
| 2020 | Transcription factors-DNA interactions in rice: identification and verificationabstractThe completion of the rice genome sequence paved the way for rice functional genomics research. Additionally, the functional characterization of transcription factors is currently a popular and crucial objective among researchers. Transcription factors are one of the groups of proteins that bind to either enhancer or promoter regions of genes to regulate expression. On the basis of several typical examples of transcription factor analyses, we herein summarize selected research strategies and methods and introduce their advantages and disadvantages. This review may provide some theoretical and technical guidelines for future investigations of transcription factors, which may be helpful to develop new rice varieties with ideal traits. Quan Zou 0001 |
Briefings Bioinform. | 3 |
| 2020 | Empirical comparison and analysis of web-based cell-penetrating peptide prediction toolsabstractCell-penetrating peptides (CPPs) facilitate the delivery of therapeutically relevant molecules, including DNA, proteins and oligonucleotides, into cells both in vitro and in vivo. This unique ability explores the possibility of CPPs as therapeutic delivery and its potential applications in clinical therapy. Over the last few decades, a number of machine learning (ML)-based prediction tools have been developed, and some of them are freely available as web portals. However, the predictions produced by various tools are difficult to quantify and compare. In particular, there is no systematic comparison of the web-based prediction tools in performance, especially in practical applications. In this work, we provide a comprehensive review on the biological importance of CPPs, CPP database and existing ML-based methods for CPP prediction. To evaluate current prediction tools, we conducted a comparative study and analyzed a total of 12 models from 6 publicly available CPP prediction tools on 2 benchmark validation sets of CPPs and non-CPPs. Our benchmarking results demonstrated that a model from the KELM-CPPpred, namely KELM-hybrid-AAC, showed a significant improvement in overall performance, when compared to the other 11 prediction models. Moreover, through a length-dependency analysis, we find that existing prediction tools tend to more accurately predict CPPs and non-CPPs with the length of 20-25 residues long than peptides in other length ranges. Ran Su, Quan Zou 0001, Balachandran Manavalan, Leyi Wei |
Briefings Bioinform. | 3 |
| 2020 | Comparative analysis and prediction of quorum-sensing peptides using feature representation learning and machine learning algorithmsabstractQuorum-sensing peptides (QSPs) are the signal molecules that are closely associated with diverse cellular processes, such as cell-cell communication, and gene expression regulation in Gram-positive bacteria. It is therefore of great importance to identify QSPs for better understanding and in-depth revealing of their functional mechanisms in physiological processes. Machine learning algorithms have been developed for this purpose, showing the great potential for the reliable prediction of QSPs. In this study, several sequence-based feature descriptors for peptide representation and machine learning algorithms are comprehensively reviewed, evaluated and compared. To effectively use existing feature descriptors, we used a feature representation learning strategy that automatically learns the most discriminative features from existing feature descriptors in a supervised way. Our results demonstrate that this strategy is capable of effectively capturing the sequence determinants to represent the characteristics of QSPs, thereby contributing to the improved predictive performance. Furthermore, wrapping this feature representation learning strategy, we developed a powerful predictor named QSPred-FL for the detection of QSPs in large-scale proteomic data. Benchmarking results with 10-fold cross validation showed that QSPred-FL is able to achieve better performance as compared to the state-of-the-art predictors. In addition, we have established a user-friendly webserver that implements QSPred-FL, which is currently available at http://server.malab.cn/QSPred-FL. We expect that this tool will be useful for the high-throughput prediction of QSPs and the discovery of important functional mechanisms of QSPs. Leyi Wei, Fuyi Li, Jiangning Song, Ran Su, Quan Zou 0001 |
Briefings Bioinform. | 6 |
| 2020 | Predicting disease-associated circular RNAs using deep forests combined with positive-unlabeled learning methodsabstractIdentification of disease-associated circular RNAs (circRNAs) is of critical importance, especially with the dramatic increase in the amount of circRNAs. However, the availability of experimentally validated disease-associated circRNAs is limited, which restricts the development of effective computational methods. To our knowledge, systematic approaches for the prediction of disease-associated circRNAs are still lacking. In this study, we propose the use of deep forests combined with positive-unlabeled learning methods to predict potential disease-related circRNAs. In particular, a heterogeneous biological network involving 17 961 circRNAs, 469 miRNAs, and 248 diseases was constructed, and then 24 meta-path-based topological features were extracted. We applied 5-fold cross-validation on 15 disease data sets to benchmark the proposed approach and other competitive methods and used Recall@k and PRAUC@k to evaluate their performance. In general, our method performed better than the other methods. In addition, the performance of all methods improved with the accumulation of known positive labels. Our results provided a new framework to investigate the associations between circRNA and disease and might improve our understanding of its functions. Xiangxiang Zeng, Quan Zou 0001 |
Briefings Bioinform. | 4 |
| 2020 | Sequence clustering in bioinformatics: an empirical studyabstractSequence clustering is a basic bioinformatics task that is attracting renewed attention with the development of metagenomics and microbiomics. The latest sequencing techniques have decreased costs and as a result, massive amounts of DNA/RNA sequences are being produced. The challenge is to cluster the sequence data using stable, quick and accurate methods. For microbiome sequencing data, 16S ribosomal RNA operational taxonomic units are typically used. However, there is often a gap between algorithm developers and bioinformatics users. Different software tools can produce diverse results and users can find them difficult to analyze. Understanding the different clustering mechanisms is crucial to understanding the results that they produce. In this review, we selected several popular clustering tools, briefly explained the key computing principles, analyzed their characters and compared them using two independent benchmark datasets. Our aim is to assist bioinformatics users in employing suitable clustering tools effectively to analyze big sequencing data. Related data, codes and software tools were accessible at the link http://lab.malab.cn/∼lg/clustering/. Quan Zou 0001, Xingpeng Jiang, Xiangrong Liu, Xiangxiang Zeng |
Briefings Bioinform. | 1 |
| 2020 | StackCPPred: a stacking and pairwise energy content-based prediction of cell-penetrating peptides and their uptake efficiencyabstractMOTIVATION: Cell-penetrating peptides (CPPs) are a vehicle for transporting into living cells pharmacologically active molecules, such as short interfering RNAs, nanoparticles, plasmid DNAs and small peptides, thus offering great potential as future therapeutics. Existing experimental techniques for identifying CPPs are time-consuming and expensive. Thus, the prediction of CPPs from peptide sequences by using computational methods can be useful to annotate and guide the experimental process quickly. Many machine learning-based methods have recently emerged for identifying CPPs. Although considerable progress has been made, existing methods still have low feature representation capabilities, thereby limiting further performance improvements. RESULTS: We propose a method called StackCPPred, which proposes three feature methods on the basis of the pairwise energy content of the residue as follows: RECM-composition, PseRECM and RECM-DWT. These features are used to train stacking-based machine learning methods to effectively predict CPPs. On the basis of the CPP924 and CPPsite3 datasets with jackknife validation, StackDPPred achieved 94.5% and 78.3% accuracy, which was 2.9% and 5.8% higher than the state-of-the-art CPP predictors, respectively. StackCPPred can be a powerful tool for predicting CPPs and their uptake efficiency, facilitating hypothesis-driven experimental design and accelerating their applications in clinical therapy. AVAILABILITY AND IMPLEMENTATION: Source code and data can be downloaded from https://github.com/Excelsior511/StackCPPred. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiangzheng Fu, Xiangxiang Zeng, Quan Zou 0001 |
Bioinform. | 4 |
| 2020 | Basic polar and hydrophobic properties are the main characteristics that affect the binding of transcription factors to methylation sitesabstractMOTIVATION: Methylation and transcription factors (TFs) are part of the mechanisms regulating gene expression. However, the numerous mechanisms regulating the interactions between methylation and TFs remain unknown. We employ machine-learning techniques to discover the characteristics of TFs that bind to methylation sites. RESULTS: The classical machine-learning analysis process focuses on improving the performance of the analysis method. Conversely, we focus on the functional properties of the TF sequences. We obtain the principal properties of TFs, namely, the basic polar and hydrophobic Ile amino acids affecting the interaction between TFs and methylated DNA. The recall of the positive instances is 0.878 when their basic polar value is >0.1743. Both basic polar and hydrophobic Ile amino acids distinguish 74% of TFs bound to methylation sites. Therefore, we infer that basic polar amino acids affect the interactions of TFs with methylation sites. Based on our results, the role of the hydrophobic Ile residue is consistent with that described in previous studies, and the basic polar amino acids may also be a key factor modulating the interactions between TFs and methylation. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Quan Zou 0001 |
Bioinform. | 2 |
| 2020 | PPTPP: a novel therapeutic peptide prediction method using physicochemical property encoding and adaptive feature representation learningabstractMOTIVATION: Peptide is a promising candidate for therapeutic and diagnostic development due to its great physiological versatility and structural simplicity. Thus, identifying therapeutic peptides and investigating their properties are fundamentally important. As an inexpensive and fast approach, machine learning-based predictors have shown their strength in therapeutic peptide identification due to excellences in massive data processing. To date, no reported therapeutic peptide predictor can perform high-quality generic prediction and informative physicochemical properties (IPPs) identification simultaneously. RESULTS: In this work, Physicochemical Property-based Therapeutic Peptide Predictor (PPTPP), a Random Forest-based prediction method was presented to address this issue. A novel feature encoding and learning scheme were initiated to produce and rank physicochemical property-related features. Besides being capable of predicting multiple therapeutics peptides with high comparability to established predictors, the presented method is also able to identify peptides' informative IPP. Results presented in this work not only illustrated the soundness of its working capacity but also demonstrated its potential for investigating other therapeutic peptides. AVAILABILITY AND IMPLEMENTATION: https://github.com/YPZ858/PPTPP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yu P. Zhang, Quan Zou 0001 |
Bioinform. | 2 |
| 2020 | Protein Complexes Identification with Family-Wise Error Rate ControlabstractThe detection of protein complexes from protein-protein interaction network is a fundamental issue in bioinformatics and systems biology. To solve this problem, numerous methods have been proposed from different angles in the past decades. However, the study on detecting statistically significant protein complexes still has not received much attention. Although there are a few methods available in the literature for identifying statistically significant protein complexes, none of these methods can provide a more strict control on the error rate of a protein complex in terms of family-wise error rate (FWER). In this paper, we propose a new detection method SSF that is capable of controlling the FWER of each reported protein complex. More precisely, we first present a p-value calculation method based on Fisher's exact test to quantify the association between each protein and a given candidate protein complex. Consequently, we describe the key modules of the SSF algorithm: a seed expansion procedure for significant protein complexes search and a set cover strategy for redundancy elimination. The experimental results on five benchmark data sets show that: (1) our method can achieve the highest precision; (2) it outperforms three competing methods in terms of normalized mutual information (NMI) and F1 score in most cases. Zengyou He, Can Zhao 0007, Bo Xu 0008, Quan Zou 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2020 | DeepAVP: A Dual-Channel Deep Neural Network for Identifying Variable-Length Antiviral PeptidesabstractAntiviral peptides (AVPs) have been experimentally verified to block virus into host cells, which have antiviral activity with decapeptide amide. Therefore, utilization of experimentally validated antiviral peptides is a potential alternative strategy for targeting medically important viruses. In this article, we propose a dual-channel deep neural network ensemble method for analyzing variable-length antiviral peptides. The LSTM channel can capture long-term dependencies for effectively studying original variable-length sequence data. The CONV channel can build dynamic neural network for analyzing the local evolution information. Also, our model can fine-tune the substitution matrix for specifically functional peptides. Applying it to a novel experimentally verified dataset, our AVPs predictor, DeepAVP, demonstrates state-of-the-art performance of [Formula: see text] accuracy and 0.85 MCC, which is far better than existing prediction methods for identifying antiviral peptides. Therefore, DeepAVP, web server for predicting the effective AVPs, would make significantly contributions to peptide-based antiviral research. Jiawei Li 0018, Yuqian Pu, Jijun Tang, Quan Zou 0001, Fei Guo 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2019 | Grouped Correlational Generative Adversarial Networks for Discrete Electronic Health RecordsabstractUsing Generative Adversarial Networks (GANs) to generate synthetic Electronic Health Records (EHR) has attracted increasing attention. However, in existing approaches, the events in EHRs are treated as separate variables which are indiscriminately entered into the model, without taking into account the meaning and grouping of them. Besides, the efficacy of treatment is often neglected. In this paper, we first embed the efficacy information into the disease diagnosis, and then propose Grouped Correlational GAN (GcGAN) to explicitly learn inherent correlations between different groups of variables. We also introduce a dense connection to strengthen the generator capacity in GcGAN. Experimental results on real-world data demonstrate that the generated data from GcGAN are able to simulate real-world data in terms of distribution statistics. The results on multi-label treatment recommendation tasks show that GcGAN can boost the performances by augmenting the training dataset with the generated data and outperforms state-of-the-art approaches. It can also automatically distinguish between disease-specific drugs and adjuvant drugs, which enhances the model interpretability. Fan Yang 0010, Zhongping Yu, Yunfan Liang, Xiaolu Gan, Kaibiao Lin, Quan Zou 0001, Yifeng Zeng |
BIBM | 6 |
| 2019 | A Prototype System using Location-based Twitter Data for Disaster ManagementabstractThe characteristics of Twitter, such as strong realtime, high attention and wide audience, make it an important information channel for people to get social hotspots and emergencies after the news media. Combining Twitter and disaster mitigation work can effectively improve the real-time data security in disaster emergency response. This paper introduces a system for collecting Twitter data posted by people during a disaster and analyzing and presenting these data. The experimental results prove that the prototype system is built successfully and effectively. Quan Zou 0001 |
IGARSS | 1 |
| 2019 | 4mCPred: machine learning methods for DNA N4-methylcytosine sites predictionabstractMOTIVATION: N4-methylcytosine (4mC), an important epigenetic modification formed by the action of specific methyltransferases, plays an essential role in DNA repair, expression and replication. The accurate identification of 4mC sites aids in-depth research to biological functions and mechanisms. Because, experimental identification of 4mC sites is time-consuming and costly, especially given the rapid accumulation of gene sequences. Supplementation with efficient computational methods is urgently needed. RESULTS: In this study, we developed a new tool, 4mCPred, for predicting 4mC sites in Caenorhabditis elegans, Drosophila melanogaster, Arabidopsis thaliana, Escherichia coli, Geoalkalibacter subterraneus and Geobacter pickeringii. 4mCPred consists of two independent models, 4mCPred_I and 4mCPred_II, for each species. The predictive results of independent and cross-species tests demonstrated that the performance of 4mCPred_I is a useful tool. To identify position-specific trinucleotide propensity (PSTNP) and electron-ion interaction potential features, we used the F-score method to construct predictive models and to compare their PSTNP features. Compared with other existing predictors, 4mCPred achieved much higher accuracies in rigorous jackknife and independent tests. We also analyzed the importance of different features in detail. AVAILABILITY AND IMPLEMENTATION: The web-server 4mCPred is accessible at http://server.malab.cn/4mCPred/index.jsp. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Wenying He, Cangzhi Jia, Quan Zou 0001 |
Bioinform. | 3 |
| 2019 | Exploring sequence-based features for the improved prediction of DNA N4-methylcytosine sites in multiple speciesabstractMOTIVATION: As one of important epigenetic modifications, DNA N4-methylcytosine (4mC) is recently shown to play crucial roles in restriction-modification systems. For better understanding of their functional mechanisms, it is fundamentally important to identify 4mC modification. Machine learning methods have recently emerged as an effective and efficient approach for the high-throughput identification of 4mC sites, although high predictive error rates are still challenging for existing methods. Therefore, it is highly desirable to develop a computational method to more accurately identify m4C sites. RESULTS: In this study, we propose a machine learning based predictor, namely 4mcPred-SVM, for the genome-wide detection of DNA 4mC sites. In this predictor, we present a new feature representation algorithm that sufficiently exploits sequence-based information. To improve the feature representation ability, we use a two-step feature optimization strategy, thereby obtaining the most representative features. Using the resulting features and Support Vector Machine (SVM), we adaptively train the optimal models for different species. Comparative results on benchmark datasets from six species indicate that our predictor is able to achieve generally better performance in predicting 4mC sites as compared to the state-of-the-art predictors. Importantly, the sequence-based features can reliably and robust predict 4mC sites, facilitating the discovery of potentially important sequence characteristics for the prediction of 4mC sites. AVAILABILITY AND IMPLEMENTATION: The user-friendly webserver that implements the proposed 4mcPred-SVM is well established, and is freely accessible at http://server.malab.cn/4mcPred-SVM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Leyi Wei, Shasha Luan, Luis Augusto Eijy Nagai, Ran Su, Quan Zou 0001 |
Bioinform. | 5 |
| 2019 | Iterative feature representations improve N4-methylcytosine site predictionabstractMOTIVATION: Accurate identification of N4-methylcytosine (4mC) modifications in a genome wide can provide insights into their biological functions and mechanisms. Machine learning recently have become effective approaches for computational identification of 4mC sites in genome. Unfortunately, existing methods cannot achieve satisfactory performance, owing to the lack of effective DNA feature representations that are capable to capture the characteristics of 4mC modifications. RESULTS: In this work, we developed a new predictor named 4mcPred-IFL, aiming to identify 4mC sites. To represent and capture discriminative features, we proposed an iterative feature representation algorithm that enables to learn informative features from several sequential models in a supervised iterative mode. Our analysis results showed that the feature representations learnt by our algorithm can capture the discriminative distribution characteristics between 4mC sites and non-4mC sites, enlarging the decision margin between the positives and negatives in feature space. Additionally, by evaluating and comparing our predictor with the state-of-the-art predictors on benchmark datasets, we demonstrate that our predictor can identify 4mC sites more accurately. AVAILABILITY AND IMPLEMENTATION: The user-friendly webserver that implements the proposed 4mcPred-IFL is well established, and is freely accessible at http://server.malab.cn/4mcPred-IFL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Leyi Wei, Ran Su, Shasha Luan, Zhijun Liao, Balachandran Manavalan, Quan Zou 0001 |
Bioinform. | 6 |
| 2019 | PEPred-Suite: improved and robust prediction of therapeutic peptides using adaptive feature representation learningabstractMOTIVATION: Prediction of therapeutic peptides is critical for the discovery of novel and efficient peptide-based therapeutics. Computational methods, especially machine learning based methods, have been developed for addressing this need. However, most of existing methods are peptide-specific; currently, there is no generic predictor for multiple peptide types. Moreover, it is still challenging to extract informative feature representations from the perspective of primary sequences. RESULTS: In this study, we have developed PEPred-Suite, a bioinformatics tool for the generic prediction of therapeutic peptides. In PEPred-Suite, we introduce an adaptive feature representation strategy that can learn the most representative features for different peptide types. To be specific, we train diverse sequence-based feature descriptors, integrate the learnt class information into our features, and utilize a two-step feature optimization strategy based on the area under receiver operating characteristic curve to extract the most discriminative features. Using the learnt representative features, we trained eight random forest models for eight different types of functional peptides, respectively. Benchmarking results showed that as compared with existing predictors, PEPred-Suite achieves better and robust performance for different peptides. As far as we know, PEPred-Suite is currently the first tool that is capable of predicting so many peptide types simultaneously. In addition, our work demonstrates that the learnt features can reliably predict different peptides. AVAILABILITY AND IMPLEMENTATION: The user-friendly webserver implementing the proposed PEPred-Suite is freely accessible at http://server.malab.cn/PEPred-Suite. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Leyi Wei, Ran Su, Quan Zou 0001 |
Bioinform. | 4 |
| 2019 | Identifying protein-protein interface via a novel multi-scale local sequence and structural representationabstractAbstract Background Protein-protein interaction plays a key role in a multitude of biological processes, such as signal transduction, de novo drug design, immune responses, and enzymatic activities. Gaining insights of various binding abilities can deepen our understanding of the interaction. It is of great interest to understand how proteins in a complex interact with each other. Many efficient methods have been developed for identifying protein-protein interface. Results In this paper, we obtain the local information on protein-protein interface, through multi-scale local average block and hexagon structure construction. Given a pair of proteins, we use a trained support vector regression (SVR) model to select best configurations. On Benchmark v4.0, our method achieves average Irmsd value of 3.28Å and overall Fnat value of 63%, which improves upon Irmsd of 3.89Å and Fnat of 49% for ZRANK, and Irmsd of 3.99Å and Fnat of 46% for ClusPro. On CAPRI targets, our method achieves average Irmsd value of 3.45Å and overall Fnat value of 46%, which improves upon Irmsd of 4.18Å and Fnat of 40% for ZRANK, and Irmsd of 5.12Å and Fnat of 32% for ClusPro. The success rates by our method, FRODOCK 2.0, InterEvDock and SnapDock on Benchmark v4.0 are 41.5%, 29.0%, 29.4% and 37.0%, respectively. Conclusion Experiments show that our method performs better than some state-of-the-art methods, based on the prediction quality improved in terms of CAPRI evaluation criteria. All these results demonstrate that our method is a valuable technological tool for identifying protein-protein interface. Fei Guo 0001, Quan Zou 0001, Jijun Tang, Junhai Xu |
BMC Bioinform. | 2 |
| 2019 | A novel collaborative filtering model for LncRNA-disease association prediction based on the Naïve Bayesian classifierabstractBACKGROUND: Since the number of known lncRNA-disease associations verified by biological experiments is quite limited, it has been a challenging task to uncover human disease-related lncRNAs in recent years. Moreover, considering the fact that biological experiments are very expensive and time-consuming, it is important to develop efficient computational models to discover potential lncRNA-disease associations. RESULTS: In this manuscript, a novel Collaborative Filtering model called CFNBC for inferring potential lncRNA-disease associations is proposed based on Naïve Bayesian Classifier. In CFNBC, an original lncRNA-miRNA-disease tripartite network is constructed first by integrating known miRNA-lncRNA associations, miRNA-disease associations and lncRNA-disease associations, and then, an updated lncRNA-miRNA-disease tripartite network is further constructed through applying the item-based collaborative filtering algorithm on the original tripartite network. Finally, based on the updated tripartite network, a novel approach based on the Naïve Bayesian Classifier is proposed to predict potential associations between lncRNAs and diseases. The novelty of CFNBC lies in the construction of the updated lncRNA-miRNA-disease tripartite network and the introduction of the item-based collaborative filtering algorithm and Naïve Bayesian Classifier, which guarantee that CFNBC can be applied to predict potential lncRNA-disease associations efficiently without entirely relying on known miRNA-disease associations. Simulation results show that CFNBC can achieve a reliable AUC of 0.8576 in the Leave-One-Out Cross Validation (LOOCV), which is considerably better than previous state-of-the-art results. Moreover, case studies of glioma, colorectal cancer and gastric cancer demonstrate the excellent prediction performance of CFNBC as well. CONCLUSIONS: According to simulation results, due to the satisfactory prediction performance, CFNBC may be an excellent addition to biomedical researches in the future. Jingwen Yu, Zhanwei Xuan, Quan Zou 0001, Lei Wang 0069 |
BMC Bioinform. | 4 |
| 2019 | Integration of deep feature representations and handcrafted features to improve the prediction of N6-methyladenosine sites
Leyi Wei, Ran Su, Xiu-Ting Li, Quan Zou 0001, Xing Gao 0004 |
Neurocomputing | 5 |
| 2019 | Details in the evaluation of circular RNA detection tools: Reply to Chen and ChuangabstractChia-Ying Chen and Trees-Juen Chuang (referred as CYC & TJC below) recently submitted their comment [1] on our previous paper [2].In their paper, they scrutinized the CircBase [3] candidates that we used and pointed out several weak points of our paper.In summary, they suggested that the positive dataset we derived from CircBase required further evaluation.They also indicated that using all of these candidates as our dataset was not appropriate.They further suggested that three main confounding factors may affect our assessment of circRNA detection tools and that their performances should be re-evaluated.Before we begin to discuss their comment, we will briefly introduce the positive dataset we used.First, as stated in our previous paper, the 14,689 candidates detected in HeLa cells were downloaded from CircBase and reported by the study of Salzman et al. [4].These candidates were not identified with the use of find_circ [5] tool.As described in the study of Salzman et al. [4], all UCSC annotated exons in scrambled order were used to construct a custom database and identify circRNA candidates.Second, in our positive dataset, constant coverage of 10× for the intervening sequence and a minimum of two read pairs (paired-end simulated reads) to cross the back-spliced junction sites were generated for each candidate.Now, we will discuss the three confounding factors they listed in their paper.First, they suggested to remove 1046 candidates with unannotated exon boundaries from the positive dataset, especially candidates without canonical splice signals, such as GT-AG, GC-AG, or AT-AC, for the junctions.As mentioned above, CircBase-deposited circRNA candidates that we used were identified by Salzman et al. [4]; the candidates identified by their method should all match the exon boundaries.The discrepancies may be caused by inconsistent gene annotation files used.Salzman et al. [4] used UCSC known genes [6], whereas CYC & TJC used NCBI RefSeq-identified mRNA annotation files.We manually checked several candidates marked with "junctions with unannotated exon boundaries" in CYC & TJC's Supplemental Dataset S1.The junction sites of these candidates were annotated as exon boundaries in UCSC known genes annotation file (http://hgdownload.soe.ucsc.edu/goldenPath/hg19/database/knownGene.txt.gz).Thus, detection of circRNAs with annotated exon boundaries relies on the gene annotation files used, and novel candidates may be missed because of the incompleteness of the current database [7].For example, Szabo et al. [7] reinforced an annotation-based algorithm with a de novo module and discovered a validated circRNA from the not-fully-annotated RMST gene and several U12 cir-cRNAs produced from unannotated boundaries.Such case was also demonstrated by Xiao-Ou Zhang et al. [8].They detected thousands of novel exons (non-RefSeq, non-Ensembl, or non-UCSC known genes) in circRNAs by using an updated CIRCexplorere2 tool, and several of them were confirmed by Northern blot analysis and Sanger sequencing after RT-PCR [8].Other examples were shown by Salzman et al. [4], they found several noncoding RNA genes expressed Xiangxiang Zeng, Maozu Guo 0001, Quan Zou 0001 |
PLoS Comput. Biol. | 4 |
| 2019 | Fast Prediction of Protein Methylation Sites Using a Sequence-Based Feature Selection TechniqueabstractProtein methylation, an important post-translational modification, plays crucial roles in many cellular processes. The accurate prediction of protein methylation sites is fundamentally important for revealing the molecular mechanisms undergoing methylation. In recent years, computational prediction based on machine learning algorithms has emerged as a powerful and robust approach for identifying methylation sites, and much progress has been made in predictive performance improvement. However, the predictive performance of existing methods is not satisfactory in terms of overall accuracy. Motivated by this, we propose a novel random-forest-based predictor called MePred-RF, integrating several discriminative sequence-based feature descriptors and improving feature representation capability using a powerful feature selection technique. Importantly, unlike other methods based on multiple, complex information inputs, our proposed MePred-RF is based on sequence information alone. Comparative studies on benchmark datasets via vigorous jackknife tests indicate that our proposed MePred-RF method remarkably outperforms other state-of-the-art predictors, leading by a 4.5 percent average in terms of overall accuracy. A user-friendly webserver that implements the proposed method has been established for researchers' convenience, and is now freely available for public use through http://server.malab.cn/MePred-RF. We anticipate our research tool to be useful for the large-scale prediction and analysis of protein methylation sites. Leyi Wei, Pengwei Xing, Gaotao Shi, Zhi-Liang Ji, Quan Zou 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2019 | Meta-Path Methods for Prioritizing Candidate Disease miRNAsabstractMicroRNAs (miRNAs) play critical roles in regulating gene expression at post-transcriptional levels. Numerous experimental studies indicate that alterations and dysregulations in miRNAs are associated with important complex diseases, especially cancers. Predicting potential miRNA-disease association is beneficial not only to explore the pathogenesis of diseases, but also to understand biological processes. In this work, we propose two methods that can effectively predict potential miRNA-disease associations using our reconstructed miRNA and disease similarity networks, which are based on the latest experimental data. We reconstruct a miRNA functional similarity network using the following biological information: the miRNA family information, miRNA cluster information, experimentally valid miRNA-target association and disease-miRNA information. We also reconstruct a disease similarity network using disease functional information and disease semantic information. We present Katz with specific weights and Katz with machine learning, on the comprehensive heterogeneous network. These methods, which achieve corresponding AUC values of 0.897 and 0.919, exhibit performance superior to the existing methods. Comprehensive data networks and reasonable considerations guarantee the high performance of our methods. Contrary to several methods, which cannot work in such situations, the proposed methods also predict associations for diseases without any known related miRNAs. A web service for the download and prediction of relationships between diseases and miRNAs is available at http://lab.malab.cn/soft/MDPredict/. Xuan Zhang 0010, Quan Zou 0001, Alfonso Rodríguez-Patón, Xiangxiang Zeng |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2019 | Advanced Machine Learning Techniques for BioinformaticsabstractThe papers in this special section focus on the machine learning methods, and applications of these methods to computational biology. Quan Zou 0001, Qi Liu 0019 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2018 | Towards Cost Effective Privacy Provision for Typed Resources in IoT Environment (S)abstractWe present privacy resources in IoT as data, information, and knowledge.We construct a privacy protection architecture on our previously proposed DIKW graphs: Data Graph, Information Graph, and Knowledge Graph.On this architecture, we search privacy protection target resources both as they appear explicitly in their original types and as they appear implicitly which means that they are expressed not in their original types.For a single privacy protection target, it may have various concrete compositions in various layers of DIKW Graph.It becomes more complex since the implementation of a privacy target might also be intertwined with the implementation of other privacy targets.We propose to protect target resources according to their types by either isolating the elements comprising an implementation, or weakening relationships among elements comprising an implement.To optimize among several choices of implementing a protection in a business environment, we introduced the tradeoff between customers' expectations/investment and privacy providers' expectation.Thereafter we proposed to prioritize implementation according to their ratio of cost/benefit. Yucong Duan, Zhengyang Song, Xiaoxian Yang, Quan Zou 0001, Xiaobing Sun 0001 |
SEKE | 4 |
| 2018 | O-GlcNAcPRED-II: an integrated classification algorithm for identifying O-GlcNAcylation sites based on fuzzy undersampling and a K-means PCA oversampling techniqueabstractMotivation: Protein O-GlcNAcylation (O-GlcNAc) is an important post-translational modification of serine (S)/threonine (T) residues that involves multiple molecular and cellular processes. Recent studies have suggested that abnormal O-G1cNAcylation causes many diseases, such as cancer and various neurodegenerative diseases. With the available protein O-G1cNAcylation sites experimentally verified, it is highly desired to develop automated methods to rapidly and effectively identify O-GlcNAcylation sites. Although some computational methods have been proposed, their performance has been unsatisfactory, particularly in terms of prediction sensitivity. Results: In this study, we developed an ensemble model O-GlcNAcPRED-II to identify potential O-GlcNAcylation sites. A K-means principal component analysis oversampling technique (KPCA) and fuzzy undersampling method (FUS) were first proposed and incorporated to reduce the proportion of the original positive and negative training samples. Then, rotation forest, a type of classifier-integrated system, was adopted to divide the eight types of feature space into several subsets using four sub-classifiers: random forest, k-nearest neighbour, naive Bayesian and support vector machine. We observed that O-GlcNAcPRED-II achieved a sensitivity of 81.05%, specificity of 95.91%, accuracy of 91.43% and Matthew's correlation coefficient of 0.7928 for five-fold cross-validation run 10 times. Additionally, the results obtained by O-GlcNAcPRED-II on two independent datasets also indicated that the proposed predictor outperformed five published prediction tools. Availability and implementation: http://121.42.167.206/OGlcPred/. Supplementary information: Supplementary data are available at Bioinformatics online. Cangzhi Jia, Yun Zuo 0001, Quan Zou 0001 |
Bioinform. | 3 |
| 2018 | Tumor origin detection with tissue-specific miRNA and DNA methylation markersabstractMotivation: A clear identification of the primary site of tumor is of great importance to the next targeted site-specific treatments and could efficiently improve patient's overall survival. Even though many classifiers based on gene expression had been proposed to predict the tumor primary, only a few studies focus on using DNA methylation (DNAm) profiles to develop classifiers, and none of them compares the performance of classifiers based on different profiles. Results: We introduced novel selection strategies to identify highly tissue-specific CpG sites and then used the random forest approach to construct the classifiers to predict the origin of tumors. We also compared the prediction performance by applying similar strategy on miRNA expression profiles. Our analysis indicated that these classifiers had an accuracy of 96.05% (Maximum-Relevance-Maximum-Distance: 90.02-99.99%) or 95.31% (principal component analysis: 79.82-99.91%) on independent DNAm datasets, and an overall accuracy of 91.30% (range 79.33-98.74%) on independent miRNA test sets for predicting tumor origin. This suggests that our feature selection methods are very effective to identify tissue-specific biomarkers and the classifiers we developed can efficiently predict the origin of tumors. We also developed a user-friendly webserver that helps users to predict the tumor origin by uploading miRNA expression or DNAm profile of their interests. Availability and implementation: The webserver, and relative data, code are accessible at http://server.malab.cn/MMCOP/. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Shixiang Wan, Andrew E. Teschendorff, Quan Zou 0001 |
Bioinform. | 5 |
| 2018 | Prediction of potential disease-associated microRNAs using structural perturbation methodabstractMotivation: The identification of disease-related microRNAs (miRNAs) is an essential but challenging task in bioinformatics research. Similarity-based link prediction methods are often used to predict potential associations between miRNAs and diseases. In these methods, all unobserved associations are ranked by their similarity scores. Higher score indicates higher probability of existence. However, most previous studies mainly focus on designing advanced methods to improve the prediction accuracy while neglect to investigate the link predictability of the networks that present the miRNAs and diseases associations. In this work, we construct a bilayer network by integrating the miRNA-disease network, the miRNA similarity network and the disease similarity network. We use structural consistency as an indicator to estimate the link predictability of the related networks. On the basis of the indicator, a derivative algorithm, called structural perturbation method (SPM), is applied to predict potential associations between miRNAs and diseases. Results: The link predictability of bilayer network is higher than that of miRNA-disease network, indicating that the prediction of potential miRNAs-diseases associations on bilayer network can achieve higher accuracy than based merely on the miRNA-disease network. A comparison between the SPM and other algorithms reveals the reliable performance of SPM which performed well in a 5-fold cross-validation. We test fifteen networks. The AUC values of SPM are higher than some well-known methods, indicating that SPM could serve as a useful computational method for improving the identification accuracy of miRNA‒disease associations. Moreover, in a case study on breast neoplasm, 80% of the top-20 predicted miRNAs have been manually confirmed by previous experimental studies. Availability and implementation: https://github.com/lecea/SPM-code.git. Supplementary information: Supplementary data are available at Bioinformatics online. Xiangxiang Zeng, Linyuan Lu, Quan Zou 0001 |
Bioinform. | 4 |
| 2018 | Efficient computation of motif discovery on Intel Many Integrated Core (MIC) ArchitectureabstractBACKGROUND: Novel sequence motifs detection is becoming increasingly essential in computational biology. However, the high computational cost greatly constrains the efficiency of most motif discovery algorithms. RESULTS: In this paper, we accelerate MEME algorithm targeted on Intel Many Integrated Core (MIC) Architecture and present a parallel implementation of MEME called MIC-MEME base on hybrid CPU/MIC computing framework. Our method focuses on parallelizing the starting point searching method and improving iteration updating strategy of the algorithm. MIC-MEME has achieved significant speedups of 26.6 for ZOOPS model and 30.2 for OOPS model on average for the overall runtime when benchmarked on the experimental platform with two Xeon Phi 3120 coprocessors. CONCLUSIONS: Furthermore, MIC-MEME has been compared with state-of-arts methods and it shows good scalability with respect to dataset size and the number of MICs. Source code: https://github.com/hkwkevin28/MIC-MEME . Shaoliang Peng, Minxia Cheng, Yingbo Cui 0001, Runxin Guo, Xiaoyu Zhang 0008, Shunyun Yang, Xiangke Liao, Yutong Lu, Quan Zou 0001, Benyun Shi |
BMC Bioinform. | 11 |
| 2018 | cmFSM: a scalable CPU-MIC coordinated drug-finding tool by frequent subgraph miningabstractBACKGROUND: Frequent subgraphs mining is a significant problem in many practical domains. The solution of this kind of problem can particularly used in some large-scale drug molecular or biological libraries to help us find drugs or core biological structures rapidly and predict toxicity of some unknown compounds. The main challenge is its efficiency, as (i) it is computationally intensive to test for graph isomorphisms, and (ii) the graph collection to be mined and mining results can be very large. Existing solutions often require days to derive mining results from biological networks even with relative low support threshold. Also, the whole mining results always cannot be stored in single node memory. RESULTS: In this paper, we implement a parallel acceleration tool for classical frequent subgraph mining algorithm called cmFSM. The core idea is to employ parallel techniques to parallelize extension tasks, so as to reduce computation time. On the other hand, we employ multi-node strategy to solve the problem of memory constraints. The parallel optimization of cmFSM is carried out on three different levels, including the fine-grained OpenMP parallelization on single node, multi-node multi-process parallel acceleration and CPU-MIC collaborated parallel optimization. CONCLUSIONS: Evaluation results show that cmFSM clearly outperforms the existing state-of-the-art miners even if we only hold a few parallel computing resources. It means that cmFSM provides a practical solution to frequent subgraph mining problem with huge number of mining results. Specifically, our solution is up to one order of magnitude faster than the best CPU-based approach on single node and presents a promising scalability of massive mining tasks in multi-node scenario. More source code are available at:Source Code: https://github.com/ysycloud/cmFSM . Shunyun Yang, Runxin Guo, Xiangke Liao, Quan Zou 0001, Benyun Shi, Shaoliang Peng |
BMC Bioinform. | 5 |
| 2018 | Prediction of human protein subcellular localization using deep learning
Leyi Wei, Yijie Ding, Ran Su, Jijun Tang, Quan Zou 0001 |
J. Parallel Distributed Comput. | 5 |
| 2018 | Processing Optimization of Typed Resources with Synchronized Storage and Computation Adaptation in Fog ComputingabstractWide application of the Internet of Things (IoT) system has been increasingly demanding more hardware facilities for processing various resources including data, information, and knowledge. With the rapid growth of generated resource quantity, it is difficult to adapt to this situation by using traditional cloud computing models. Fog computing enables storage and computing services to perform at the edge of the network to extend cloud computing. However, there are some problems such as restricted computation, limited storage, and expensive network bandwidth in Fog computing applications. It is a challenge to balance the distribution of network resources. We propose a processing optimization mechanism of typed resources with synchronized storage and computation adaptation in Fog computing. In this mechanism, we process typed resources in a wireless‐network‐based three‐tier architecture consisting of Data Graph, Information Graph, and Knowledge Graph. The proposed mechanism aims to minimize processing cost over network, computation, and storage while maximizing the performance of processing in a business value driven manner. Simulation results show that the proposed approach improves the ratio of performance over user investment. Meanwhile, conversions between resource types deliver support for dynamically allocating network resources. Zhengyang Song, Yucong Duan, Shixiang Wan, Xiaobing Sun 0001, Quan Zou 0001, Honghao Gao, Donghai Zhu |
Wirel. Commun. Mob. Comput. | 5 |
| 2017 | Iteratively collective prediction of disease-gene associations through the incomplete networkabstractThe prediction of links between genes and disease is still one of the biggest challenges in the field of human health. Almost all state-of-the-art studies on the prediction of gene-disease links focuson a single pair of links, ignoring the associations and interactions among different types of links. Moreover, the biological information networks are usually incomplete. In this paper, we study the similarity measure to be used on two different types of nodes, based on the metapaths between them (Wsrm). Then an iterative self-updating approach for link prediction using heterogeneous information network is proposed to fit the incompletion of the network (ISL), which is a semi-supervised learning formula. Using the biological integrated network constructed from OMIM and HumanNet dataset (30,896 nodes and 1,200,166 edges) we applied our framework. The area under the receiver operating characteristic is 0.941, indicating that our approach significantly outperforms the state-of-the-art gene-disease link prediction approaches. Moreover, the sensitivity analysis signifies that our approach is robust. Consequently, our proposed framework demonstrates an efficient and accurate approach for link prediction between genes and diseases. In addition, during iteration, the accuracy of the result gradually increases. The example dataset and the implementation of our approach is avaliable at https://github.com/xymeng16/ISL. Xiangyi Meng, Quan Zou 0001, Alfonso Rodríguez-Patón, Xiangxiang Zeng |
BIBM | 2 |
| 2017 | Learning Planning and Recommendation Based on an Adaptive Architecture on Data Graph, Information Graph and Knowledge Graph
Lixu Shao, Yucong Duan, Zhangbing Zhou, Quan Zou 0001, Honghao Gao |
CollaborateCom | 4 |
| 2017 | Wavelet-based single image super-resolution with an overall enhancement procedureabstractIn this paper, we address the problem of generating a super-resolution image based on a dictionary of low- and high-resolution exemplars from a single input image in wavelet domain with a overall enhancement procedure. Most methods extract different kinds of features in low-resolution image and high-resolution images to establish the mapping relation. But in this paper, we implement wavelet-transform to extract the same kind of feature to make the mapping more reasonable. Meanwhile we implement local Lipschitz regularity constraint and structure-keeping constraint to preserve the local singularity and edge in our method. Compared with current state-of-art methods on standard images, our method obtains both visual and PSNR improvement. Zongqing Lu 0001, Quan Zou 0001, Fei Zhou 0001, Qingmin Liao |
ICASSP | 2 |
| 2017 | A Pay as You Use Resource Security Provision Approach Based on Data Graph, Information Graph and Knowledge Graph
Lixu Shao, Yucong Duan, Li-Zhen Cui 0001, Quan Zou 0001, Xiaobing Sun 0001 |
IDEAL | 4 |
| 2017 | Specifying architecture of knowledge graph with data graph, information graph, knowledge graph and wisdom graphabstractKnowledge graphs have been widely adopted, in large part owing to their schema-less nature. It enables knowledge graphs to grow seamlessly and allows for new relationships and entities as needed. Knowledge graph has become a powerful tool to represent knowledge in the form of a labelled directed graph and to give semantics to textual information. A knowledge graph is a graph constructed by representing each item, entity and user as nodes, and linking those nodes that interact with each other via edges. Knowledge graph has abundant natural semantics and can contain various and more complete information. Its expression mechanism is close to natural language. However, we still lack a unified definition and standard expression form of knowledge graph. We propose to clarify the expression of knowledge graph as a whole. We clarify the architecture of knowledge graph from data, information, knowledge, and wisdom aspects respectively. We also propose to specify knowledge graph in a progressive manner as four basic forms including data graph, information graph, knowledge graph and wisdom graph. Yucong Duan, Lixu Shao, Gongzhu Hu, Zhangbing Zhou, Quan Zou 0001, Zhaoxin Lin |
SERA | 5 |
| 2017 | Bidirectional value driven design between economical planning and technical implementation based on data graph, information graph and knowledge graphabstractValue-Driven Design enables rational decisions to be made in terms of the optimum business and technical solution at every level of engineering design by employing economics in decision making. In order to maximize the business profitability, we propose to bridge bidirectional value driven design between economic planning and technology implementation on the basis of the data graph, information graph and knowledge graph. We use data graph, information graph and knowledge graph to analyze problems that have negative impact on activities of software development including requirement analysis, summary design and detail design. We propose to improve system reliability and robustness by managing data and information reuse, redundancy as well as structure. Lixu Shao, Yucong Duan, Xiaobing Sun 0001, Quan Zou 0001, Rongqi Jing, Jiami Lin |
SERA | 4 |
| 2017 | Machine learning and graph analytics in computational biomedicine
Quan Zou 0001, Lei Chen 0007, Tao Huang 0004, Yungang Xu |
Artif. Intell. Medicine | 1 |
| 2017 | CMSA: a heterogeneous CPU/GPU computing system for multiple similar RNA/DNA sequence alignmentabstractThe multiple sequence alignment (MSA) is a classic and powerful technique for sequence analysis in bioinformatics. With the rapid growth of biological datasets, MSA parallelization becomes necessary to keep its running time in an acceptable level. Although there are a lot of work on MSA problems, their approaches are either insufficient or contain some implicit assumptions that limit the generality of usage. First, the information of users’ sequences, including the sizes of datasets and the lengths of sequences, can be of arbitrary values and are generally unknown before submitted, which are unfortunately ignored by previous work. Second, the center star strategy is suited for aligning similar sequences. But its first stage, center sequence selection, is highly time-consuming and requires further optimization. Moreover, given the heterogeneous CPU/GPU platform, prior studies consider the MSA parallelization on GPU devices only, making the CPUs idle during the computation. Co-run computation, however, can maximize the utilization of the computing resources by enabling the workload computation on both CPU and GPU simultaneously. This paper presents CMSA, a robust and efficient MSA system for large-scale datasets on the heterogeneous CPU/GPU platform. It performs and optimizes multiple sequence alignment automatically for users’ submitted sequences without any assumptions. CMSA adopts the co-run computation model so that both CPU and GPU devices are fully utilized. Moreover, CMSA proposes an improved center star strategy that reduces the time complexity of its center sequence selection process from O(mn 2) to O(mn). The experimental results show that CMSA achieves an up to 11× speedup and outperforms the state-of-the-art software. CMSA focuses on the multiple similar RNA/DNA sequence alignment and proposes a novel bitmap based algorithm to improve the center star strategy. We can conclude that harvesting the high performance of modern GPU is a promising approach to accelerate multiple sequence alignment. Besides, adopting the co-run computation model can maximize the entire system utilization significantly. The source code is available at https://github.com/wangvsa/CMSA . Chen Wang 0004, Shanjiang Tang, Ce Yu, Quan Zou 0001 |
BMC Bioinform. | 5 |
| 2017 | Local-DPP: An improved DNA-binding protein prediction method by exploring local evolutionary information
Leyi Wei, Jijun Tang, Quan Zou 0001 |
Inf. Sci. | 3 |
| 2017 | A comprehensive overview and evaluation of circular RNA detection toolsabstractCircular RNA (circRNA) is mainly generated by the splice donor of a downstream exon joining to an upstream splice acceptor, a phenomenon known as backsplicing. It has been reported that circRNA can function as microRNA (miRNA) sponges, transcriptional regulators, or potential biomarkers. The availability of massive non-polyadenylated transcriptomes data has facilitated the genome-wide identification of thousands of circRNAs. Several circRNA detection tools or pipelines have recently been developed, and it is essential to provide useful guidelines on these pipelines for users, including a comprehensive and unbiased comparison. Here, we provide an improved and easy-to-use circRNA read simulator that can produce mimicking backsplicing reads supporting circRNAs deposited in CircBase. Moreover, we compared the performance of 11 circRNA detection tools on both simulated and real datasets. We assessed their performance regarding metrics such as precision, sensitivity, F1 score, and Area under Curve. It is concluded that no single method dominated on all of these metrics. Among all of the state-of-the-art tools, CIRI, CIRCexplorer, and KNIFE, which achieved better balanced performance between their precision and sensitivity, compared favorably to the other methods. Xiangxiang Zeng, Maozu Guo 0001, Quan Zou 0001 |
PLoS Comput. Biol. | 4 |
| 2017 | Inferring MicroRNA-Disease Associations by Random Walk on a Heterogeneous Network with Multiple Data SourcesabstractSince the discovery of the regulatory function of microRNA (miRNA), increased attention has focused on identifying the relationship between miRNA and disease. It has been suggested that computational method are an efficient way to identify potential disease-related miRNAs for further confirmation using biological experiments. In this paper, we first highlighted three limitations commonly associated with previous computational methods. To resolve these limitations, we established disease similarity subnetwork and miRNA similarity subnetwork by integrating multiple data sources, where the disease similarity is composed of disease semantic similarity and disease functional similarity, and the miRNA similarity is calculated using the miRNA-target gene and miRNA-lncRNA (long non-coding RNA) associations. Then, a heterogeneous network was constructed by connecting the disease similarity subnetwork and the miRNA similarity subnetwork using the known miRNA-disease associations. We extended random walk with restart to predict miRNA-disease associations in the heterogeneous network. The leave-one-out cross-validation achieved an average area under the curve (AUC) of 0:8049 across 341 diseases and 476 miRNAs. For five-fold cross-validation, our method achieved an AUC from 0:7970 to 0:9249 for 15 human diseases. Case studies further demonstrated the feasibility of our method to discover potential miRNA-disease associations. An online service for prediction is freely available at http://ifmda.aliapp.com. Yuansheng Liu, Xiangxiang Zeng, Zengyou He, Quan Zou 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2017 | Prediction and Validation of Disease Genes Using HeteSim ScoresabstractDeciphering the gene disease association is an important goal in biomedical research. In this paper, we use a novel relevance measure, called HeteSim, to prioritize candidate disease genes. Two methods based on heterogeneous networks constructed using protein-protein interaction, gene-phenotype associations, and phenotype-phenotype similarity, are presented. In HeteSim_MultiPath (HSMP), HeteSim scores of different paths are combined with a constant that dampens the contributions of longer paths. In HeteSim_SVM (HSSVM), HeteSim scores are combined with a machine learning method. The 3-fold experiments show that our non-machine learning method HSMP performs better than the existing non-machine learning methods, our machine learning method HSSVM obtains similar accuracy with the best existing machine learning method CATAPULT. From the analysis of the top 10 predicted genes for different diseases, we found that HSSVM avoid the disadvantage of the existing machine learning based methods, which always predict similar genes for different diseases. The data sets and Matlab code for the two methods are freely available for download at http://lab.malab.cn/data/HeteSim/index.jsp. Xiangxiang Zeng, Yuanlu Liao, Yuansheng Liu, Quan Zou 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2016 | Latent factor model with heterogeneous similarity regularization for predicting gene-disease associationsabstractThe correct prediction of human genes related to diseases has been a challenge in biological research. Considering extensive gene-disease data verified by biological experiments, we can apply computational methods to perform correct predictions with reduced time and expenses. On the basis of a previously designed latent factorization model (LFM), which performs well in recommender systems, we propose a latent factor model with heterogeneous similarity regularization (LFMHSR) to predict disease-related genes. Various types of data, including those of humans and other related species, are used in this method. First, model I with an average heterogeneous regularization is proposed on the basis of a typical LFM. Second, model II with personal heterogeneous regularization is developed to improve the deficiency of the previous model. Data on other nonhuman species and vector space similarity or Pearson correlation coefficient metrics are also utilized in our method. Results reveal that the performance of LFMHSR is 7% more efficient than that of other existing approaches. Therefore, our proposed approach can be employed to predict novel diseases or genes with no known associations. Xiangxiang Zeng, Ningxiang Ding, Quan Zou 0001 |
BIBM | 3 |
| 2016 | Multiple sequence alignment and reconstructing phylogenetic trees with HadoopabstractMultiple sequence alignment (MSA) is the “Holy Grail” problem in computational biology, but bottlenecks arise in the massive MSA of homologous sequences. Most of the available state-of-the-art software tools cannot address large-scale datasets, or they run rather slowly. The similarity of homologous DNA sequences is often ignored. Lack of parallelization is still a challenge for MSA research. Building the phylogenetic trees for ultra-large sequences is also a time-consuming work. MSA is the previous work for phylogenetic reconstruction. With the development of parallel computation, we employed Hadoop platform to solve the two computational intensive problems. Trie trees and suffix trees were used for accelerating multiple similar DNA sequences alignment. The expected time complexity was decreased to linear time from square time. For the phylogenetic tree reconstruction, clustering and multiple-sequence alignment were executed in parallel, and the basic phylogenetic trees were built using the neighbour-joining model. Experiments on two large datasets, both more than 1 GB, show that our software tool can outperform other common phylogenetic reconstruction tools. Furthermore, data, software codes, and web servers were all opened in http://lab.malab.cn/soft/halign/ and http://lab.malab.cn/soft/HPtree/. Quan Zou 0001 |
BIBM | 1 |
| 2016 | HPTree: Reconstructing phylogenetic trees for ultra-large unaligned DNA sequences via NJ model and HadoopabstractConstructing phylogenetic tree for ultra-large sequences (eg. Files more than 1GB) is quite difficult, especially for the unaligned DNA sequences. It is meaningless and impracticable to do multiple sequence alignment for large diverse DNA sequences. We try to do clustering firstly for the mounts of DNA sequences, and divide them into several clusters. Then each cluster is aligned and phylogenetic analysed in parallel. Hadoop, which is the most popular parallel platform in cloud computing, is employed for this process. Our software tool HPTree can handle the >1GB DNA sequence file or more than 1,000,000 DNA sequences in few hours. Users could try HPTree in the cloud computing platform (eg. Amazon) or their own clusters for the big data phylogenetic tree reconstruction. No super machine or large memory is required. HPTree could benefit the users who focus on population evolution or long common genes (eg. 16s rRNA) evolution. The software tool along with its codes and datasets are accessible at http://lab.malab.cn/soft/HPtree/. Quan Zou 0001, Shixiang Wan, Xiangxiang Zeng |
BIBM | 1 |
| 2016 | BDSCyto: An Automated Approach for Identifying Cytokines Based on Best Dimension Searching
Quan Zou 0001, Shixiang Wan, Bing Han 0003, Zhihui Zhan |
PRICAI | 1 |
| 2016 | Integrative approaches for predicting microRNA function and prioritizing disease-related microRNA using biological interaction networksabstractMicroRNAs (miRNA) play critical roles in regulating gene expressions at the posttranscriptional levels. The prediction of disease-related miRNA is vital to the further investigation of miRNA's involvement in the pathogenesis of disease. In previous years, biological experimentation is the main method used to identify whether miRNA was associated with a given disease. With increasing biological information and the appearance of new miRNAs every year, experimental identification of disease-related miRNAs poses considerable difficulties (e.g. time-consumption and high cost). Because of the limitations of experimental methods in determining the relationship between miRNAs and diseases, computational methods have been proposed. A key to predict potential disease-related miRNA based on networks is the calculation of similarity among diseases and miRNA over the networks. Different strategies lead to different results. In this review, we summarize the existing computational approaches and present the confronted difficulties that help understand the research status. We also discuss the principles, efficiency and differences among these methods. The comprehensive comparison and discussion elucidated in this work provide constructive insights into the matter. Xiangxiang Zeng, Xuan Zhang 0010, Quan Zou 0001 |
Briefings Bioinform. | 3 |
| 2016 | Annotation-retrieval reinforcement by visual cognition modeling on manifold
Quan Zou 0001 |
Neurocomputing | 4 |
| 2016 | Hierarchical support vector machine based structural classification with fused hierarchies
Yahong Han, Quan Zou 0001, Qinghua Hu |
Neurocomputing | 3 |
| 2016 | Advanced learning for large-scale heterogeneous computing
Quan Zou 0001, Wei Liu 0005, Michele Merler, Rongrong Ji |
Neurocomputing | 1 |
| 2016 | A novel features ranking metric with application to scalable visual and bioinformatics data classification
Quan Zou 0001, Jian-Cang Zeng, Liujuan Cao, Rongrong Ji |
Neurocomputing | 1 |
| 2015 | HAlign: Fast multiple similar DNA/RNA sequence alignment based on the centre star strategyabstractAbstract Motivation: Multiple sequence alignment (MSA) is important work, but bottlenecks arise in the massive MSA of homologous DNA or genome sequences. Most of the available state-of-the-art software tools cannot address large-scale datasets, or they run rather slowly. The similarity of homologous DNA sequences is often ignored. Lack of parallelization is still a challenge for MSA research. Results: We developed two software tools to address the DNA MSA problem. The first employed trie trees to accelerate the centre star MSA strategy. The expected time complexity was decreased to linear time from square time. To address large-scale data, parallelism was applied using the hadoop platform. Experiments demonstrated the performance of our proposed methods, including their running time, sum-of-pairs scores and scalability. Moreover, we supplied two massive DNA/RNA MSA datasets for further testing and research. Availability and implementation: The codes, tools and data are accessible free of charge at http://datamining.xmu.edu.cn/software/halign/. Contact: [email protected] or [email protected] Quan Zou 0001, Qinghua Hu, Maozu Guo 0001, Guohua Wang 0001 |
Bioinform. | 1 |
| 2015 | Identification of cytokine via an improved genetic algorithm
Xiangxiang Zeng, Sisi Yuan, Xianxian Huang, Quan Zou 0001 |
Frontiers Comput. Sci. | 4 |
| 2015 | Asynchronous spiking neural P systems with rules on synapses
Tao Song 0001, Quan Zou 0001, Xiangrong Liu, Xiangxiang Zeng |
Neurocomputing | 2 |
| 2014 | Survey of MapReduce frame operation in bioinformaticsabstractBioinformatics is challenged by the fact that traditional analysis tools have difficulty in processing large-scale data from high-throughput sequencing. The open source Apache Hadoop project, which adopts the MapReduce framework and a distributed file system, has recently given bioinformatics researchers an opportunity to achieve scalable, efficient and reliable computing performance on Linux clusters and on cloud computing services. In this article, we present MapReduce frame-based applications that can be employed in the next-generation sequencing and other biological domains. In addition, we discuss the challenges faced by this field as well as the future works on parallel computing in bioinformatics. Quan Zou 0001, Xu-Bin Li, Wen-Rui Jiang, Ziyu Lin, Gui-Lin Li |
Briefings Bioinform. | 1 |
| 2014 | Using distances between Top-n-gram and residue pairs for protein remote homology detectionabstractBACKGROUND: Protein remote homology detection is one of the central problems in bioinformatics, which is important for both basic research and practical application. Currently, discriminative methods based on Support Vector Machines (SVMs) achieve the state-of-the-art performance. Exploring feature vectors incorporating the position information of amino acids or other protein building blocks is a key step to improve the performance of the SVM-based methods. RESULTS: Two new methods for protein remote homology detection were proposed, called SVM-DR and SVM-DT. SVM-DR is a sequence-based method, in which the feature vector representation for protein is based on the distances between residue pairs. SVM-DT is a profile-based method, which considers the distances between Top-n-gram pairs. Top-n-gram can be viewed as a profile-based building block of proteins, which is calculated from the frequency profiles. These two methods are position dependent approaches incorporating the sequence-order information of protein sequences. Various experiments were conducted on a benchmark dataset containing 54 families and 23 superfamilies. Experimental results showed that these two new methods are very promising. Compared with the position independent methods, the performance improvement is obvious. Furthermore, the proposed methods can also provide useful insights for studying the features of protein families. CONCLUSION: The better performance of the proposed methods demonstrates that the position dependant approaches are efficient for protein remote homology detection. Another advantage of our methods arises from the explicit feature space representation, which can be used to analyze the characteristic features of protein families. The source code of SVM-DT and SVM-DR is available at http://bioinformatics.hitsz.edu.cn/DistanceSVM/index.jsp. Bin Liu 0014, Jinghao Xu, Quan Zou 0001, Ruifeng Xu 0001, Xiaolong Wang 0001, Qingcai Chen |
BMC Bioinform. | 3 |
| 2014 | nDNA-prot: identification of DNA-binding proteins based on unbalanced classificationabstractBACKGROUND: DNA-binding proteins are vital for the study of cellular processes. In recent genome engineering studies, the identification of proteins with certain functions has become increasingly important and needs to be performed rapidly and efficiently. In previous years, several approaches have been developed to improve the identification of DNA-binding proteins. However, the currently available resources are insufficient to accurately identify these proteins. Because of this, the previous research has been limited by the relatively unbalanced accuracy rate and the low identification success of the current methods. RESULTS: In this paper, we explored the practicality of modelling DNA binding identification and simultaneously employed an ensemble classifier, and a new predictor (nDNA-Prot) was designed. The presented framework is comprised of two stages: a 188-dimension feature extraction method to obtain the protein structure and an ensemble classifier designated as imDC. Experiments using different datasets showed that our method is more successful than the traditional methods in identifying DNA-binding proteins. The identification was conducted using a feature that selected the minimum Redundancy and Maximum Relevance (mRMR). An accuracy rate of 95.80% and an Area Under the Curve (AUC) value of 0.986 were obtained in a cross validation. A test dataset was tested in our method and resulted in an 86% accuracy, versus a 76% using iDNA-Prot and a 68% accuracy using DNA-Prot. CONCLUSIONS: Our method can help to accurately identify DNA-binding proteins, and the web server is accessible at http://datamining.xmu.edu.cn/~songli/nDNA. In addition, we also predicted possible DNA-binding protein sequences in all of the sequences from the UniProtKB/Swiss-Prot database. Xiangxiang Zeng, Quan Zou 0001 |
BMC Bioinform. | 6 |
| 2014 | LibD3C: Ensemble classifiers with a clustering and dynamic selection strategy
Chen Lin 0001, Sridhar Krishnan 0001, Quan Zou 0001 |
Neurocomputing | 6 |
| 2014 | Improved and Promising Identificationof Human MicroRNAs by Incorporatinga High-Quality Negative SetabstractMicroRNA (miRNA) plays an important role as a regulator in biological processes. Identification of (pre-) miRNAs helps in understanding regulatory processes. Machine learning methods have been designed for pre-miRNA identification. However, most of them cannot provide reliable predictive performances on independent testing data sets. We assumed this is because the training sets, especially the negative training sets, are not sufficiently representative. To generate a representative negative set, we proposed a novel negative sample selection technique, and successfully collected negative samples with improved quality. Two recent classifiers rebuilt with the proposed negative set achieved an improvement of ~6 percent in their predictive performance, which confirmed this assumption. Based on the proposed negative set, we constructed a training set, and developed an online system called miRNApre specifically for human pre-miRNA identification. We showed that miRNApre achieved accuracies on updated human and non-human data sets that were 34.3 and 7.6 percent higher than those achieved by current methods. The results suggest that miRNApre is an effective tool for pre-miRNA identification. Additionally, by integrating miRNApre, we developed a miRNA mining tool, mirnaDetect, which can be applied to find potential miRNAs in genome-scale data. MirnaDetect achieved a comparable mining performance on human chromosome 19 data as other existing methods. Leyi Wei, Minghong Liao, Yue Gao 0002, Rongrong Ji, Zengyou He, Quan Zou 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2012 | Performance Optimization of Analysis Rules in Real-Time Active Data Warehouses
Ziyu Lin, Dongzhan Zhang, Chen Lin 0001, Yongxuan Lai, Quan Zou 0001 |
APWeb | 5 |
| 2012 | PaGeFinder: quantitative identification of spatiotemporal pattern genesabstractUNLABELLED: Pattern Gene Finder (PaGeFinder) is a web-based server for on-line detection of gene expression patterns from serial transcriptomic data generated by high-throughput technologies like microarray or next-generation sequencing. Three particular parameters, the specificity measure, the dispersion measure and the contribution measure, were introduced and implemented in PaGeFinder to help quantitative and interactive identification of pattern genes like housekeeping genes, specific (selective) genes and repressed genes. Besides the on-line computation service, the PaGeFinder also provides downloadable Java programs for local detection of gene expression patterns. AVAILABILITY: http://bioinf.xmu.edu.cn:8080/PaGeFinder/index.jsp Shi-chang Hu, Quan Zou 0001, Zhi-Liang Ji |
Bioinform. | 4 |
| 2012 | Identify content quality in online social networksabstractThe flooding of low-quality user generated contents (UGC) in online social network (OSN) has been a threat to web knowledge management systems. Recently several domain-specific systems have been developed addressing this problem, for example, predict correct answer in QA community; recognise reliable comment in products review forums etc. Major drawback of most research efforts is the lack of a general framework applicable to all OSNs. In this study, the authors start by analysing the effects of distinguishing features on UGC quality in different types of OSNs. Extensive statistical analysis leads to the discovery of existence of diverse patterns of human information sharing activity in dissimilar OSNs. This discovery is employed as prior knowledge in the classification framework, which decompose the original highly imbalanced problem into several balanced sub-problems. Ensemble classifiers are adopted in samples from clusters generated by incompact features. Experiments show the proposed framework is both effective and efficient for several OSNs.Contributions of this study are two-fold: (i) model posting activity in different types of OSNs; (ii) propose novel classification framework to identify UGC quality. Chen Lin 0001, Zhenhua Huang 0001, Fan Yang 0010, Quan Zou 0001 |
IET Commun. | 4 |
| 2011 | Maintaining Internal Consistency of Report for Real-Time OLAP with Layer-Based View
Ziyu Lin, Yongxuan Lai, Chen Lin 0001, Yi Xie 0004, Quan Zou 0001 |
APWeb | 5 |
| 2010 | TiSGeD: a database for tissue-specific genesabstractUNLABELLED: The tissue-specific genes are a group of genes whose function and expression are preferred in one or several tissues/cell types. Identification of these genes helps better understanding of tissue-gene relationship, etiology and discovery of novel tissue-specific drug targets. In this study, a statistical method is introduced to detect tissue-specific genes from more than 123 125 gene expression profiles over 107 human tissues, 67 mouse tissues and 30 rat tissues. As a result, a novel subject-specialized repository, namely the tissue-specific genes database (TiSGeD), is developed to represent the analyzed results. Auxiliary information of tissue-specific genes was also collected from biomedical literatures. AVAILABILITY: http://bioinf.xmu.edu.cn/databases/TiSGeD/index.html. Sheng-Jian Xiao, Quan Zou 0001, Zhi-Liang Ji |
Bioinform. | 3 |
| 2008 | Interest filter vs. interest operator: Face recognition using Fisher linear discriminant based on interest filter representation
Tuo Zhao, Zhizheng Liang, David Zhang 0001, Quan Zou 0001 |
Pattern Recognit. Lett. | 4 |